Papers by Gaurav Kumar

17 papers
Exemplar Encoder-Decoder for Neural Conversation Generation (P18-1)

Copied to clipboard

Challenge: Existing approaches to generate conversational systems suffer from lack of diversity in responses and generation of short, repetitive and uninteresting responses.
Approach: They propose a novel conversation model that uses similar examples from training data to generate responses.
Outcome: The proposed model outperforms state-of-the-art sequence to sequence learning on several evaluation metrics on two large data sets.
Detecting Incongruent News Articles Using Multi-head Attention Dual Summarization (2022.aacl-main)

Copied to clipboard

Challenge: Recent studies on incongruity detection focus on estimating the similarity between the headline and the encoding of the body or its summary but most of these methods fail to handle inconvenient news articles created with embedded noise.
Approach: They propose a method which generates two types of summaries that capture the congruent and incongruent parts in the body separately.
Outcome: The proposed method outperforms the state-of-the-art methods over three publicly available datasets.
Seeing Beyond: Enhancing Visual Question Answering with Multi-Modal Retrieval (2025.coling-industry)

Copied to clipboard

Challenge: Multi-modal Large language models still suffer from model hallucination and lack of specific knowledge when answering challenging questions.
Approach: They propose to use a multi-modal retrieval augmented generation method to integrate knowledge from all modalities into a model to enable alignment between query and knowledge.
Outcome: The proposed method achieves significant performance improvement on the VQA dataset.
Enhancing User Safety: Context-Aware Detection of Offensive Query-Ad Pairs in Multimodal Search Advertising (2026.eacl-industry)

Copied to clipboard

Challenge: Multi-modal online advertisements require robust content moderation to ensure user safety . key challenges include nuanced, multi-modal nature of ads, severe data scarcity and class imbalance due to the rarity of offensive content .
Approach: They propose a framework that detects offensive content only when a user's search query is paired with a specific ad .
Outcome: The proposed framework reduces the serving of offensive query-ad pairs by more than 80% while maintaining the efficiency required for real-time advertising systems.
MM-SOC: Benchmarking Multimodal Large Language Models in Social Media Platforms (2024.findings-acl)

Copied to clipboard

Challenge: Social media platforms are hubs for multimodal information exchange, encompassing text, images, and videos, making it challenging for machines to comprehend the information or emotions associated with interactions in online spaces.
Approach: They propose a benchmark to evaluate MLLMs' understanding of multimodal social media content and a large-scale YouTube tagging dataset to evaluate their performance.
Outcome: The proposed model performs better in a zero-shot setting, suggesting potential improvements.
Cross-Modal Attribute Insertions for Assessing the Robustness of Vision-and-Language Learning (2023.acl-long)

Copied to clipboard

Challenge: Existing approaches to model multimodal data do not leverage cross-modal information . augmenting input text using cross-module attribute insertions results in poor performance .
Approach: They propose a multimodal deep learning approach that adds visual attributes to inputs to enhance model robustness.
Outcome: The proposed approach is modular, controllable, and task-agnostic.
Robust In-Context Selection via Online Learned Position-Corrected Attention (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods to fix this limitation can be classified into two ways: (1) Methods that use the LLM to generate the selection either via logits of item identifiers, or explicit rank permutations often requiring multiple LLM calls or fine-tuning.
Approach: They propose a method that harnesses attention patterns available from a single forward call on the Large Language Model (LLM) the method learns the logic for item selection using a few in-context examples and a simple online position-debiasing mechanism to correct attention distortion.
Outcome: The proposed method improves selection performance over direct generation and prior attention-based methods while remaining robust to prompt variations and item ordering.
Gaining Insights into Unrecognized User Utterances in Task-Oriented Dialog Systems (2022.emnlp-industry)

Copied to clipboard

Challenge: Goal-oriented dialog systems fail to recognize the intent of natural language requests due to system errors, incomplete service coverage, or insufficient training.
Approach: They propose an end-to-end pipeline for processing unrecognized user utterances, deployed in a commercial task-oriented dialog system, including a specifically-tailored clustering algorithm, a novel approach to cluster representative extraction, and cluster naming.
Outcome: The proposed components show that they improve the performance of the proposed system in the analysis of unrecognized user requests.
Complexity-Weighted Loss and Diverse Reranking for Sentence Simplification (N19-1)

Copied to clipboard

Challenge: Recent research has applied sequence-to-sequence (Seq2Sequen) models to text simplification . generic models tend to copy directly from the original sentence, resulting in outputs that are long and complex.
Approach: They propose to incorporate word complexities into the loss function during training and generate a large set of diverse candidate simplifications at test time.
Outcome: The proposed model can perform competitively with state-of-the-art systems while generating simpler sentences.
Curriculum Learning for Domain Adaptation in Neural Machine Translation (N19-1)

Copied to clipboard

Challenge: Neural machine translation (NMT) performance drops when domains do not match and in-domain training data is scarce.
Approach: They propose a curriculum learning approach to adapt generic neural machine translation models to a specific domain.
Outcome: The proposed approach outperforms unadapted and adapted baselines in two domains and two language pairs.
Cross-Modal Projection in Multimodal LLMs Doesn’t Really Project Visual Attributes to Textual Space (2024.acl-short)

Copied to clipboard

Challenge: Existing multimodal large language models are limited to general-purpose multimodal tasks like question-answering on natural images.
Approach: They propose to use cross-modal projection networks and a large language model to model domain-specific visual attributes of MLLMs.
Outcome: The proposed models gain domain-specific visual capabilities when the projection is fine-tuned, but the updates do not extract relevant domain-specific visual attributes.
AMUSED: A Multi-Stream Vector Representation Method for Use in Natural Dialogue (2020.lrec-1)

Copied to clipboard

Challenge: Current architectures only take care of semantic and contextual information for a given query and fail to fully account for syntactic and external knowledge which are crucial for generating responses in a chit-chat system.
Approach: They propose a multi-stream deep learning architecture that learns unified embeddings for query-response pairs by incorporating Graph Convolution Networks over their dependency parse.
Outcome: The proposed architecture improves on the next sentence prediction task and significantly improves existing techniques.
Reinforcement Learning based Curriculum Optimization for Neural Machine Translation (N19-1)

Copied to clipboard

Challenge: a heterogeneous training dataset can vary in characteristics such as domain, translation quality, and degree of difficulty.
Approach: They propose to use reinforcement learning to learn an optimal curriculum for NMT training . they find it can beat uniform baselines and hand-designed, state-of-the-art curricula .
Outcome: The proposed approach beats baselines and hand-designed curricula on English-to-French datasets.
Adversarial Robustness of Prompt-based Few-Shot Learning for Natural Language Understanding (2023.findings-acl)

Copied to clipboard

Challenge: Recent few-shot learning methods focus on improving downstream task performance, but there is limited understanding of the adversarial robustness of such methods.
Approach: They evaluate prompt-based FSL methods against fully fine-tuned models to better understand the impact of various factors towards robustness.
Outcome: The proposed methods show that they are less robust in the face of adversarial perturbations than fully fine-tuned models.
TabPert : An Effective Platform for Tabular Perturbation (2021.emnlp-demo)

Copied to clipboard

Challenge: Current transformers-based models outperform humans on factual evidence evaluation tasks when presented as simple unstructured text.
Approach: TabPert generates counterfactual data to assess model tabular reasoning issues.
Outcome: TabPert analyzes the model's shortcomings methodically and quantitatively.
A Community-Centric Perspective for Characterizing and Detecting Anti-Asian Violence-Provoking Speech (2024.acl-long)

Copied to clipboard

Challenge: Using data from 420k Twitter posts, we characterize anti-Asian violence-provoking speech and collect a community-crowdsourced dataset to facilitate its large-scale detection.
Approach: They develop a codebook to characterize anti-Asian violence-provoking speech and collect a community-crowdsourced dataset to facilitate its large-scale detection.
Outcome: The proposed codebook analyzes 420k tweets over 3 years and compares classifiers with hateful speech classifier classifier to detect hateful content.
Robustness of Fusion-based Multimodal Classifiers to Cross-Modal Content Dilutions (2022.emnlp-main)

Copied to clipboard

Challenge: Existing work has focused on understanding the robustness of vision-and-language models to imperceptible variations on benchmark tasks.
Approach: They develop a model that generates additional dilution text that maintains relevance and topical coherence with the image and existing text, and when added to the original text, leads to misclassification of the multimodal input.
Outcome: The proposed model outperforms fusion-based classifiers on Crisis Humanitarianism and Sentiment Detection tasks by 23.3% and 22.5% in presence of dilutions generated by the model.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations