Papers by Pushpak Bhattacharyya

148 papers
CEASE, a Corpus of Emotion Annotated Suicide notes in English (2020.lrec-1)

Copied to clipboard

Challenge: Existing studies on suicide notes have not explored the topic of emotion detection.
Approach: They develop a fine-grained emotion annotated corpus of suicide notes in English and use it to perform emotion detection on a curated dataset.
Outcome: The proposed model performs emotion detection on a curated dataset of 205 suicide notes in English.
Looks can be Deceptive: Distinguishing Repetition Disfluency from Reduplication (2025.coling-main)

Copied to clipboard

Challenge: Existing research indicates that disfluencies can constitute up to 5.9% of words in spontaneous speech, with repetitions accounting for over half of these disfluency.
Approach: They propose to use a dataset to analyze reduplication and repetition in speech using computational linguistics to evaluate transformer-based models.
Outcome: The proposed models achieve macro F1 scores of up to 85.62% in Hindi, 83.95% in Telugu, and 84.82% in Marathi for reduplication-repetition classification.
A Retrofitting Model for Incorporating Semantic Relations into Word Embeddings (2020.coling-main)

Copied to clipboard

Challenge: Existing word embedding models mix semantic similarity with other types of relatedness.
Approach: They propose a model that leverages relational knowledge available in a knowledge resource to improve word embeddings.
Outcome: The proposed model improves word embeddings on synonymy, antonymy and hypernymy relations in WordNet and significantly improves lexical entailment detection task.
All-in-One: A Deep Attentive Multi-task Learning Framework for Humour, Sarcasm, Offensive, Motivation, and Sentiment on Memes (2020.aacl-main)

Copied to clipboard

Challenge: Empirical results show the efficacy of our proposed multi-task framework over existing state-of-the-art systems.
Approach: They propose a multi-task, multi-modal deep learning framework to solve multiple tasks simultaneously.
Outcome: The proposed framework performs better than existing state-of-the-art systems on a complicated form of information, i.e., memes.
Role of Language Relatedness in Multilingual Fine-tuning of Language Models: A Case Study in Indo-Aryan Languages (2021.emnlp-main)

Copied to clipboard

Challenge: A study of multilingual fine-tuning yields better performance on downstream NLP applications . low resource languages such as Oriya and Punjabi are found to be the largest beneficiaries of multi-lingual fine tuning.
Approach: They propose to leverage the relatedness of languages that belong to the same family in NLP models by multilingual fine-tuning.
Outcome: The proposed approach improves performance on downstream NLP tasks by 15% compared to monolingual fine-tuning.
GRAFT: A Graph-based Flow-aware Agentic Framework for Document-level Machine Translation (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing Document-level machine translation systems struggle to handle discourse-level phenomena such as pronoun resolution, lexical cohesion, and ellipsis.
Approach: They propose a graph-based document-level machine translation framework that leverages Large Language Models to model translation flow and discourse structure.
Outcome: The proposed framework outperforms commercial and closed systems in eight languages and six domains.
IndiFoodVQA: Advancing Visual Question Answering and Reasoning with a Knowledge-Infused Synthetic Data Generation Pipeline (2024.findings-eacl)

Copied to clipboard

Challenge: Large Vision Language Models lack domain-specific data for reasoning on complex problems.
Approach: They propose to use explicit knowledge-infused questions, answers, and reasons to answer and reason upon the questions.
Outcome: The proposed model improves by 25% over the baseline model.
How low is too low? A monolingual take on lemmatisation in Indian languages (2021.naacl-main)

Copied to clipboard

Challenge: Prior work on ML based lemmatization focused on high resource languages, where data sets (word forms) are readily available.
Approach: They propose to use neural methods to relate inflected forms of words to their dictionary form to reduce the sparse data problem.
Outcome: The proposed methods can give competitive accuracy even in low resource setting.
Hi-GEC: Hindi Grammar Error Correction in Low Resource Scenario (2025.coling-main)

Copied to clipboard

Challenge: Automated Grammatical Error Correction (GEC) is a scarcely explored low-resource language . a recent study focused on English, but it focused on Hindi, which presents unique challenges due to its complex syntax and intricate morphology.
Approach: They propose to use a human-edited dataset to generate Hindi GEC data . they also investigate round trip translation using diverse languages for the technique .
Outcome: The proposed method outperforms other methods in Hindi, showing that it is highly efficient.
CaRVE: Critiquing and Refining Visual Elaborations for Figurative Language Illustrations (2026.findings-acl)

Copied to clipboard

Challenge: Existing text-to-image frameworks for figurative illustration rely on proprietary models or human supervision to achieve adequate alignment.
Approach: They propose a critique-driven framework that uses VLM feedback to refine visual elaborations for figurative image generation.
Outcome: The proposed framework outperforms existing figurative image-to-text pipelines on human-supervised visual elaborations.
Looking Beyond the Pixels: Evaluating Visual Metaphor Understanding in VLMs (2025.findings-emnlp)

Copied to clipboard

Challenge: Visual metaphors are a complex vision–language phenomenon that requires both perceptual and conceptual reasoning to understand.
Approach: They introduce a visual metaphor dataset featuring 2177 synthetic and 350 human-annotated images and benchmark several SOTA VLMs on two tasks: Visual Metaphor Captioning (VMC) and Visual Metamorphosis VQA (VM-VQA).
Outcome: The proposed model outperforms standard few-shot baselines on visual metaphors and VM-VQA tasks.
DeepSentiPeer: Harnessing Sentiment in Review Texts to Recommend Peer Review Decisions (P19-1)

Copied to clipboard

Challenge: Existing peer review system is not straightforward and requires domain knowledge, expertise, and intelligence of human reviewers, which is somewhat elusive with the current state of AI.
Approach: They propose to use peer review texts to predict acceptance or rejection of a manuscript based on reviewer sentiment.
Outcome: The proposed deep neural architecture achieves significant performance improvement over baselines (29% error reduction) in a recently released dataset of peer reviews.
Meme-ingful Analysis: Enhanced Understanding of Cyberbullying in Memes Through Multimodal Explanations (2024.eacl-long)

Copied to clipboard

Challenge: Recent laws like “right to explanations” have spurred research in developing interpretable models . a recent study has shown that multimodal explanations improve performance in generating textual justifications .
Approach: They propose to use visual and textual modalities to explain why a given meme is cyberbullying . they use a Contrastive Language-Image Pretraining approach to generate textual justifications .
Outcome: The proposed model improves performance in visual and textual explanations and identifies the visual evidence supporting a decision.
A Sentiment and Emotion Aware Multimodal Multiparty Humor Recognition in Multilingual Conversational Setting (2022.coling-1)

Copied to clipboard

Challenge: Humor is an essential aspect of daily conversation, and people try to provoke humor in their talks.
Approach: They propose a multitask framework that annotates Hindi utterances with sentiment and emotion classes.
Outcome: The proposed framework improves on the recently released Hindi Humor dataset . it takes sentiment and emotion into account to understand humor .
ToxVidLM: A Multimodal Framework for Toxicity Detection in Code-Mixed Videos (2024.findings-acl)

Copied to clipboard

Challenge: Using a dataset of 931 videos with 4021 code-mixed Hindi-English utterances, we find that video content with multiple modalities is more accurate and more accurate than textual content.
Approach: They propose to use a dataset to analyze toxic content in video content in non-English languages by leveraging language models.
Outcome: The proposed framework achieves an Accuracy and Weighted F1 score of 94.29% and 94.35% for the first time in its class.
Eyes Show the Way: Modelling Gaze Behaviour for Hallucination Detection (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for hallucination detection depend on knowledge sources that are explicit such as Wikipedia or knowledge graphs.
Approach: They propose a cognitive approach that leverages gaze signals from humans to detect hallucinations in natural language processing (NLP) they collect and introduce an eye tracking corpus consisting of 500 instances, annotated by five annotators for hallucinism detection.
Outcome: The proposed approach achieves a balanced accuracy of 87.1% on a FactCC dataset.
Understand the Implication: Learning to Think for Pragmatic Understanding (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods rely on annotated labels but overlook the reasoning process humans naturally use to interpret implicit meaning.
Approach: They propose a dataset that includes explicit reasoning for both correct and incorrect interpretations and propose supervised fine-tuning to improve their performance.
Outcome: The proposed dataset improves LLMs' pragmatic understanding by 11.12% across model families and 16.10% over label trained models.
FrameNet-assisted Noun Compound Interpretation (2021.findings-acl)

Copied to clipboard

Challenge: Existing methods for predicting semantic labels for noun compound interpretation are difficult.
Approach: They propose to predict semantic labels in a continuous embedding space using FrameNet data.
Outcome: The proposed method performs well on unseen labels, with 5% and 2% improvement over baselines for frame and FE prediction.
Rethinking Research on Stereotypes: An Analysis through Social Psychological and Computational Perspectives (2026.findings-acl)

Copied to clipboard

Challenge: Existing research on stereotypical biases ignores literature on them and results in resource wastage.
Approach: They argue that stereotypes are social constructs shaping human perception and behavior that can produce harmful outcomes under specific conditions.
Outcome: The proposed models can inherit and amplify stereotypes under certain conditions.
HalluDetect: Detecting, Mitigating, and Benchmarking Hallucinations in Conversational Systems in the Legal Domain (2025.emnlp-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) are widely used in industry but still produce hallucinations, limiting their reliability in critical applications.
Approach: They propose to reduce hallucinations in consumer grievance chatbots by reducing their token accuracy by 0.4159 per turn.
Outcome: The proposed system achieves an F1 score of 68.92% outperforming baseline detectors by 22.47% while maintaining the highest token accuracy.
Retrofitting Light-weight Language Models for Emotions using Supervised Contrastive Learning (2023.emnlp-main)

Copied to clipboard

Challenge: a novel retrofitting method to induce emotion aspects into pre-trained language models is proposed . the models are computationally less expensive and open, but do not capture affective aspects of human communication well.
Approach: They propose a retrofitting method to induce emotion aspects into pre-trained language models . they retrofit text fragments exhibiting similar emotions into pretrained networks .
Outcome: The proposed method produces emotion-aware text representations for sentiment analysis and sarcasm detection tasks.
Disfluency Correction using Unsupervised and Semi-supervised Learning (2021.eacl-main)

Copied to clipboard

Challenge: Disfluencies in conversational speech can affect performance of downstream NLP tasks.
Approach: They propose a disfluency correction model that converts disfluent to fluent text . they use unsupervised encoder-decoder models to generate semi-supervised models .
Outcome: The proposed model achieves a BLEU score of 79.39 on the Switchboard corpus test set and 85.28 with semi-supervision.
IndicIRSuite: Multilingual Dataset and Neural Information Models for Indian Languages (2024.acl-short)

Copied to clipboard

Challenge: IndicIRSuite is the first attempt at building large-scale Neural Information Retrieval resources for a large number of Indian languages.
Approach: They introduce Neural Information Retrieval resources for 11 widely spoken Indian Languages from two major Indian language families.
Outcome: Experiments show that Indic-ColBERT improves on INDIC-MARCO datasets for 11 languages, and that it can be used to improve IR for Indian languages.
Multi-Task Learning Framework for Mining Crowd Intelligence towards Clinical Treatment (N18-2)

Copied to clipboard

Challenge: In recent past, social media has emerged as an active platform in the context of healthcare and medicine.
Approach: They propose to use a novel adversarial learning approach to capture medical sentiments expressed in a medical blog to analyze the user's opinions on health-related issues.
Outcome: The proposed framework can capture the user's opinions on health-related issues at a medical blog level.
Synthesize, if you do not have: Effective Synthetic Dataset Creation Strategies for Self-Supervised Opinion Summarization in E-commerce (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches to generate general and aspect-specific opinion summarization are limited due to their reliance on human-specified aspects and seed words.
Approach: They propose synthetic dataset creation approaches for general and aspect-specific opinion summarization . general opinion summaries struggle to generate faithful to the input reviews, they say . aspect- specific opinion summarisation models are limited due to reliance on human-specified aspects .
Outcome: The proposed approach outperforms existing models on three e-commerce test sets on general and aspect-specific opinion summarization.
Multilingual Unsupervised NMT using Shared Encoder and Language-Specific Decoders (P19-1)

Copied to clipboard

Challenge: Existing approaches to train multiple languages with a shared encoder and multiple decoders are based on denoising autoencoding of each language and back-translating between English and multiple non-English languages.
Approach: They propose a multilingual unsupervised NMT scheme which trains multiple languages with a shared encoder and multiple decoders.
Outcome: The proposed model performs better than the separately trained bilingual models on monolingual corpora and improves by 1.48 BLEU points on WMT test sets.
Indian Language Wordnets and their Linkages with Princeton WordNet (L18-1)

Copied to clipboard

Challenge: Wordnets are rich lexico-semantic resources. Linked wordnets link similar concepts in wordnet of different languages.
Approach: They propose to map 18 Indian wordnets linked with Princeton WordNet . they use expansion approach with Hindi Wordnet as pivot .
Outcome: The proposed mappings of 18 Indian wordnets are based on Princeton WordNet . they show that availability of such resources will have a direct impact on NLP progress .
One Prompt To Rule Them All: LLMs for Opinion Summary Evaluation (2024.acl-long)

Copied to clipboard

Challenge: Existing evaluation methods for opinion summarizations lack adequate opinion summary evaluation datasets.
Approach: They propose a dataset that combines 7 dimensions crucial to opinion summaries . they propose OP-I-PROMPT, a dimension-independent prompt, and OP PROMPTS, .
Outcome: The proposed model achieves a Spearman correlation of 0.70 with human judgments, surpassing prior methods.
Persona or Context? Towards Building Context adaptive Personalized Persuasive Virtual Sales Assistant (2022.aacl-main)

Copied to clipboard

Challenge: Existing task-oriented conversational agents assume that end-users will always have a pre-determined and servable task goal, which results in dialogue failure in hostile scenarios, such as goal unavailability.
Approach: They propose to build an end-to-end multi-modal persuasive dialogue system incorporating a personalized persuasive module aided goal controller and goal persuader.
Outcome: The proposed system achieves user tasks even in goal unavailability scenarios by persuading them towards a similar and servable goal.
Rad-Flamingo: A Multimodal Prompt driven Radiology Report Generation Framework with Patient-Centric Explanations (2026.findings-eacl)

Copied to clipboard

Challenge: Existing reports are labor-intensive and expert-intensive, resulting in inconsistencies and a lack of patient-centered insight.
Approach: They propose a multimodal prompt-driven report generation framework that integrates diverse data modalities to produce comprehensive and context-aware radiology reports.
Outcome: The proposed framework improves report quality, improves understandability and could foster better patient-doctor communication.
Looking at Radiology Report Generation through a Causal Lens: A Survey (2026.acl-long)

Copied to clipboard

Challenge: Existing surveys on RRG emphasize deep learning while overlooking the critical role of causality.
Approach: They propose to analyze biases across the RRG pipeline and formalize it as a causal modeling problem and review representative causal techniques from the literature.
Outcome: The proposed model can mitigate biases and yield fair, reliable systems with clinically meaningful outputs.
EmoInHindi: A Multi-label Emotion and Intensity Annotated Dataset in Hindi for Emotion Recognition in Dialogues (2022.lrec-1)

Copied to clipboard

Challenge: Existing datasets for emotion recognition in dialogues are in English . existing datasets are limited to a few languages like Hindi .
Approach: They propose a large conversational dataset in Hindi for multi-label emotion and intensity recognition in conversations . they use a Wizard-of-Oz manner to annotate dialogues with 16 emotion labels .
Outcome: The proposed dataset contains 1,814 dialogues with 44,247 utterances in Hindi . it is based on a Wizard-of-Oz manner and can detect emotions in conversation .
Reinforcement Replaces Supervision: Query focused Summarization using Deep Reinforcement Learning (2023.emnlp-main)

Copied to clipboard

Challenge: Query-focused Summarization (QfS) is a system that generates summaries from document(s) based on a query.
Approach: They propose a Query-focused Summarization approach that uses a generalization of Reinforcement Learning (RL) for Natural Language Generation and a better semantic similarity reward.
Outcome: The proposed approach improves on the ROUGE-L metric and in a benchmark dataset.
Together We Can: Multilingual Automatic Post-Editing for Low-Resource Languages (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on multilingual automatic post-editing systems for low-resource Indo-Aryan languages have focused on different models for different language pairs.
Approach: They propose to use a multilingual automatic post-editing system to improve machine translations for low-resource Indo-Aryan languages.
Outcome: The proposed model outperforms English-Hindi and English-Marathi models by 2.5 and 2.39 TER points.
Emotion Enriched Retrofitted Word Embeddings (2022.coling-1)

Copied to clipboard

Challenge: Word embeddings that encode lexical-semantic relations do not capture emotion aspects of words.
Approach: They propose a retrofitting method to update the vectors of emotion bearing words . they find that the retrofitted embeddings achieve better distances between clusters .
Outcome: The proposed method achieves better distances between clusters and clusters for words having the same emotions.
Adversarial Training for Low-Resource Disfluency Correction (2023.findings-acl)

Copied to clipboard

Challenge: Disfluencies can be introduced in conversational speech due to the conversational nature of speech and/or speech impairments such as stuttering.
Approach: They propose an adversarial sequence-tagging model for Disfluency Correction . they evaluate it in Bengali, Hindi, and Marathi languages and use it to correct stuttering disfluencies .
Outcome: The proposed technique improves in Bengali, Hindi, and Marathi languages . it also removes stuttering disfluencies in ASR transcripts introduced by speech impairments .
SrcMix: Mixing of Related Source Languages Benefits Extremely Low-resource Machine Translation (2026.findings-eacl)

Copied to clipboard

Challenge: Multilingual models are widely used for machine translation, but their effectiveness for extremely low-resource languages (ELRLs) is dependent on how related languages are incorporated during fine-tuning.
Approach: They propose a source-side mixing strategy that combines related ELRLs during fine-tuning while constraining the decoder to a single target language.
Outcome: The proposed approach improves performance in high-resource to ELRL translations and in mid-resourced to MT translations.
IndiGEC: Multilingual Grammar Error Correction for Low-Resource Indian Languages (2025.emnlp-main)

Copied to clipboard

Challenge: despite growing interest in GEC, most research has focused on English due to the lack of benchmark datasets for low-resource lan-guages.
Approach: They propose a new approach to generate high-quality synthetic data for GEC using monolingual corpora.
Outcome: The proposed framework outperforms other monolingual methods in English, Hindi, Bengali, Marathi, and Tamil.
Towards Sentiment and Emotion aided Multi-modal Speech Act Classification in Twitter (2021.naacl-main)

Copied to clipboard

Challenge: Speech Act Classification determining the communicative intent of an utterance has been investigated widely over the years as a standalone task.
Approach: They propose a multi-modal, emotion-TA dataset called EmoTA from open-source Twitter dataset and a Dyadic Attention Mechanism framework that integrates intra-modal and inter-modal attention to fuse multiple modalities.
Outcome: The proposed framework boosts the performance of the primary task, i.e., TA classification (TAC), by benefitting from the two secondary tasks, namely, Sentiment and Emotion Analysis compared to its uni-modal and single task TAC variants.
Improving Machine Translation with Phrase Pair Injection and Corpus Filtering (2022.emnlp-main)

Copied to clipboard

Challenge: In this paper, we show that the combination of Phrase Pair Injection and Corpus Filtering boosts the performance of Neural Machine Translation systems.
Approach: They propose to combine Phrase Pair Injection and Corpus Filtering to boost performance of Neural Machine Translation systems.
Outcome: The proposed method improves machine translation models on low-resource language pairs . BLEU score improves over models trained with whole pseudo-parallel corpus augmented with parallel corpus.
Challenge Dataset of Cognates and False Friend Pairs from Indian Languages (2020.lrec-1)

Copied to clipboard

Challenge: Cognates are words that have a common etymological origin and can facilitate the Second Language Acquisition (SLA) however, they also pose a challenge to various NLP applications such as Machine Translation and Cross-lingual Sense Disambiguation.
Approach: They create two cognate datasets for twelve Indian languages and use them to generate cognate sets.
Outcome: The proposed datasets are curated using previously available baseline cognate detection approaches and evaluated with the help of lexicographers.
Are Language Models Agnostic to Linguistically Grounded Perturbations? A Case Study of Indic Languages (2025.findings-naacl)

Copied to clipboard

Challenge: Existing studies do not focus on linguistically grounded attacks, but pre-trained models are susceptible to these perturbations.
Approach: They propose to examine whether pre-trained language models are agnostic to linguistically grounded attacks . they find that PLMs are less susceptible to linguistic perturbations than non-linguistic ones .
Outcome: The proposed model is agnostic to linguistically grounded attacks, but is less susceptible to linguist attacks than non-linguistic models.
StereoDetect: Detecting Stereotypes and Anti-stereotypes the Correct Way Using Social Psychological Underpinnings (2025.findings-emnlp)

Copied to clipboard

Challenge: Stereotypes are known to have harmful effects, making their detection critical . current research focuses on detecting and evaluating stereotypical biases .
Approach: They propose a five-tuple definition and provide precise terminologies disentangling stereotypes, antistereotypes, stereotypical bias, and general bias.
Outcome: The proposed framework disentangles stereotypes, antistereotypes, stereotypical bias, and general bias.
A Match Made in Heaven: A Multi-task Framework for Hyperbole and Metaphor Detection (2023.findings-acl)

Copied to clipboard

Challenge: Existing approaches to detect metaphor and hyperbole independently have not explored their relationship computationally.
Approach: They propose a multi-task deep learning framework to detect hyperbole and metaphor simultaneously by annotating two hyperbolic datasets with metaphor labels.
Outcome: The proposed framework improves state-of-the-art hyperbole detection by 12% over existing methods.
Addressing word-order Divergence in Multilingual Neural Machine Translation for extremely Low Resource Languages (N19-1)

Copied to clipboard

Challenge: Existing studies show that transfer learning works best when the languages are related.
Approach: They propose to pre-order assisting language sentences to match the word order of the source language and train the parent model.
Outcome: The proposed model can improve translation quality in low-resource scenarios by pre-ordering the assisting language sentences to match the word order of the source language and training the parent model.
CoSTA: Code-Switched Speech Translation using Aligned Speech-Text Interleaving (2025.coling-main)

Copied to clipboard

Challenge: More than half of the world's population is presumed to be bilingual . spoken translation of code-switched speech has been under-explored .
Approach: They propose an end-to-end model architecture CoSTA that scaffolds on pretrained ASR and MT modules.
Outcome: The proposed model outperforms existing models by 3.5 BLEU points in spoken translation of code-switched speech.
A Morphology-Based Investigation of Positional Encodings (2024.emnlp-main)

Copied to clipboard

Challenge: Contemporary deep learning models handle languages with diverse morphology . morphological complexity of languages is closely linked with positional encodings .
Approach: They propose to use positional encodings to integrate morphological complexity into deep learning models.
Outcome: The proposed model improves on 22 languages and 5 downstream tasks.
ASAP++: Enriching the ASAP Automated Essay Grading Dataset with Essay Attribute Scores (L18-1)

Copied to clipboard

Challenge: Automated essay grading (AEG) is one of the most challenging activities in natural language processing (NLP).
Approach: They propose to annotate the ASAP AEG dataset and use it to score different attributes of the essays.
Outcome: The proposed resource is based on the ASAP++ dataset, which contains scores for different attributes of the essays, such as content, word choice, organization, sentence fluency, etc.
RPTCS: A Reinforced Persona-aware Topic-guiding Conversational System (2023.eacl-main)

Copied to clipboard

Challenge: Existing systems that control concept transitions in a conversation lack a persona-aware topic transition dataset.
Approach: They propose a persona-aware topic-guiding conversational system that leads the conversation to drift to a set of target concepts depending on the persona of the speaker and the context of the conversation.
Outcome: The proposed system produces fluent responses with no useful information and is based on a conversational dataset with a human-in-loop only quality checks.
Main Predicate and Their Arguments as Explanation Signals For Intent Classification (2025.naacl-long)

Copied to clipboard

Challenge: Intent classification is crucial for conversational agents, and deep learning models perform well in this area due to the lack of suitable benchmark data.
Approach: They propose a technique to augment text samples from intent classification datasets with word-level explanations by marking main predicates and their arguments as explanation signals.
Outcome: The proposed method augments text samples from intent classification datasets with word-level explanations.
“Let’s not Quote out of Context”: Unified Vision-Language Pretraining for Context Assisted Image Captioning (2023.acl-industry)

Copied to clipboard

Challenge: Large enterprises have several teams to create their content for the purpose of marketing, campaigning, or even maintaining a brand presence.
Approach: They propose a new unified Vision-Language (VL) model with a focus on context-assisted image captioning where the caption is generated based on both the image and its context.
Outcome: The proposed model achieves state-of-the-art with an improvement of up to 8.34 CIDEr score on the benchmark news image captioning datasets.
A Semi-supervised Approach to Generate the Code-Mixed Text using Pre-trained Encoder and Transfer Learning (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to train neural network-based models for code-mixing are limited due to language specificity of code-mixed text.
Approach: They propose a deep learning approach to generate code-mixed text from English to multiple languages without any parallel data.
Outcome: The proposed approach generates a code-mixed text from English to multiple languages without any parallel data.
Many Hands Make Light Work: Using Essay Traits to Automatically Score Essays (2022.naacl-main)

Copied to clipboard

Challenge: In automatic essay grading, essay traits are important for scoring the essay holistically . a single-task learning system gives the best results for scoring essays holistically and scoring essay traits.
Approach: They propose a way to score essays using a multi-task learning approach . they compare the MTL-based BiLSTM system to a single-task Learning approach based on LSTMs and BiLStms .
Outcome: The proposed system gives better results for scoring essay holistically and scoring essay traits.
“My life is miserable, have to sign 500 autographs everyday”: Exposing Humblebragging, the Brags in Disguise (2025.findings-acl)

Copied to clipboard

Challenge: Humblebragging is a phenomenon in which individuals present self-promotional statements under the guise of modesty or complaints.
Approach: They propose a task of automatically detecting humblebragging in text and propose '4-tuple definition' they also propose machine learning, deep learning, and large language models to perform the task .
Outcome: The proposed model achieves an F1-score of 0.88 and is non-trivial even for humans.
Towards Emotion-aided Multi-modal Dialogue Act Classification (2020.acl-main)

Copied to clipboard

Challenge: Considerable work on Dialogue Act Classification (DAC) has been done on textual inputs.
Approach: They propose to use a multimodal Emotion aware Dialogue Act dataset to explore the role of multi-modality and emotion recognition in DAC.
Outcome: The proposed dataset shows that multi-modality and emotion recognition improves DAC performance compared to uni-modal and single task DAC variants.
Contextual Inter-modal Attention for Multi-modal Sentiment Analysis (D18-1)

Copied to clipboard

Challenge: Existing methods for multi-modal sentiment analysis are limited due to the use of text, visual and acoustic inputs.
Approach: They propose a recurrent neural network based multi-modal attention framework that leverages contextual information for utterance-level sentiment prediction.
Outcome: The proposed framework performs better on two multi-modal sentiment analysis benchmark datasets with accuracies of 82.31% and 79.80% for the MOSI and MOSEI datasets.
Mental Disorder Classification via Temporal Representation of Text (2024.findings-emnlp)

Copied to clipboard

Challenge: Current methods for mental disorder prediction split data into chunks and use limited context length . mental health professionals lack the skills to diagnose and treat mental disorders .
Approach: They propose a framework which compresses chronologically ordered social media posts into a series of numbers and uses this time variant representation for mental disorder classification.
Outcome: The proposed framework outperforms existing models in depression, self-harm and anorexia . it also shows that the proposed framework can be used across domains .
Unveiling the Invisible: Captioning Videos with Metaphors (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent studies have shown that Vision-Language models cannot understand visual metaphors in memes and adverts.
Approach: They propose a task to describe visual metaphors in videos using a manually created dataset and a new metric called Average Concept Distance to automatically evaluate creativity.
Outcome: The proposed system performs comparable to existing video language models on the proposed task and can be used for future research.
“So You Think You’re Funny?”: Rating the Humour Quotient in Standup Comedy (2021.emnlp-main)

Copied to clipboard

Challenge: Existing datasets for humour classification are limited due to the subjectivity of the content and the multiple interpretations of the data.
Approach: They propose to annotate a multi-modal humour-annotated dataset using stand-up comedy clips and compute a humor quotient using the audience's laughter.
Outcome: The proposed scoring mechanism is validated by comparing with manual scoring methods and achieves an accuracy of 0.813 in terms of QWK.
Novelty Goes Deep. A Deep Neural Solution To Document Level Novelty Detection (C18-1)

Copied to clipboard

Challenge: Existing methods for document-level novelty detection are limited and do not require manual feature engineering.
Approach: They propose a deep Convolutional Neural Networks based model to classify a document as novel or redundant on the basis of documents already seen by the system.
Outcome: The proposed model outperforms the state-of-the-art on a document-level novelty detection dataset by a margin of 5% in terms of accuracy.
Eyes are the Windows to the Soul: Predicting the Rating of Text Quality Using Gaze Behaviour (P18-1)

Copied to clipboard

Challenge: Existing methods to predict text quality include estimating subjective aspects of text, like structure, clarity, etc.
Approach: They propose to capture gaze behaviour to help predict text quality by reporting improvements obtained by adding gaze features to traditional textual features for score prediction.
Outcome: The proposed model shows that capturing gaze behaviour improves the accuracy of score prediction when the reader has fully understood the text.
Pretraining Language Models Using Translationese (2024.emnlp-main)

Copied to clipboard

Challenge: a recent study shows that large language models perform well in low-resource languages . a vast majority of languages don't have comparable data as compared to English .
Approach: They propose to use Translationese as synthetic data for pre-training language models for low-resource languages.
Outcome: The proposed method reduces performance of LMs trained on clean data in Indian languages . the proposed model performs better in English than in other languages, but is not comparable to English.
Unsupervised Aspect-Level Sentiment Controllable Style Transfer (2020.aacl-main)

Copied to clipboard

Challenge: Unsupervised style transfer has been explored in text.
Approach: They propose a system where aspect-level sentiments can be controlled at the output . they propose to use unsupervised techniques such as ABSA masked-language-modelling .
Outcome: The proposed system is successful in controlling aspect-level sentiments.
MEISD: A Multimodal Multi-Label Emotion, Intensity and Sentiment Dialogue Dataset for Emotion Recognition and Sentiment Analysis in Conversations (2020.coling-main)

Copied to clipboard

Challenge: Emotion and sentiment classification in dialogues has gained popularity in recent times . a number of datasets are imbalanced in representing different emotions and consist of an only single emotion.
Approach: They propose to use a dataset to analyze emotions and sentiments in dialogues . they use text, audio and video to identify the correct emotions with the appropriate intensity and sentiment in an utterance of a dialogue .
Outcome: The proposed datasets are balanced in representing different emotions and consist of only one emotion.
DISCO: A Large Scale Human Annotated Corpus for Disfluency Correction in Indo-European Languages (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing research on disfluency correction has primarily focused on English due to the unavailability of large-scale open-source datasets.
Approach: They propose to use an annotated human-annotated corpus to analyze disfluency correction in four important Indo-European languages to demonstrate the benefits.
Outcome: The proposed model improves BLEU scores by 5.65 points when used with a state-of-the-art machine translation system.
A Unified Multi-task Adversarial Learning Framework for Pharmacovigilance Mining (P19-1)

Copied to clipboard

Challenge: Existing methods of identifying ADRs are reliable but time-consuming and offer a limited amount of ADR relevant information.
Approach: They propose a neural network-inspired multi-task learning framework that can simultaneously extract ADRs from various sources.
Outcome: The proposed framework achieves state-of-the-art performance on three publicly available real-world benchmark pharmacovigilance datasets, a Twitter dataset from PSB 2016 Social Me- dia Shared Task, CADEC corpus and Medline ADR corpus.
Ordinal and Attribute Aware Response Generation in a Multimodal Dialogue System (P19-1)

Copied to clipboard

Challenge: Existing multimodal dialogue systems are based on unimodal sources, capturing information from text and image.
Approach: They propose a position and attribute aware attention mechanism to learn enhanced image representation conditioned on the user utterance.
Outcome: The proposed model outperforms the state-of-the-art models on text similarity metrics.
DocCGen: Document-based Controlled Code Generation (2024.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) produce state-of-the-art performance on natural language to code generation for resource-rich general-purpose languages like C++, Java, and Python.
Approach: They propose a framework that breaks the NL-to-Code generation task into two steps . they use library documentation to detect the correct libraries and schema rules extracted from the documentation to constrain the decoding .
Outcome: The proposed framework improves different sized language models across all six evaluation metrics, reducing syntactic and semantic errors in structured code.
Divide, Link, and Conquer: Recall-oriented Schema Linking for NL-to-SQL via Question Decomposition (2025.emnlp-industry)

Copied to clipboard

Challenge: Open-source LLMs often depend on large proprietary models, which introduce serious privacy concerns.
Approach: They propose a plug-and-play framework that improves SQL generation for smaller LLMs . they propose to apply question decomposition at the schema linking stage rather than during SQL generation .
Outcome: The proposed framework improves schema linking recall by 25.1% and execution accuracy by 8.2% on the BIRD benchmark.
Treat us like the sequences we are: Prepositional Paraphrasing of Noun Compounds using LSTM (C18-1)

Copied to clipboard

Challenge: Using prepositions, noun compounds are interpreted in two ways: labelling and paraphrasing.
Approach: They propose to paraphrase noun compounds using prepositions by using parallelly aligned sequences of words.
Outcome: The proposed approach performs well on datasets manually annotated with prepositions.
Seeing Is Believing! towards Knowledge-Infused Multi-modal Medical Dialogue Generation (2024.lrec-main)

Copied to clipboard

Challenge: Existing models of disease diagnosis using AI do not use knowledge infusion.
Approach: They propose a transformer-based, knowledge-infused multi-modal medical dialogue generation framework . they propose 'discourse-aware' image identifier that recognizes signs and their severity .
Outcome: The proposed model outperforms state-of-the-art models by 7.84% in the english language.
Looking inside Noun Compounds: Unsupervised Prepositional and Free Paraphrasing (2020.findings-emnlp)

Copied to clipboard

Challenge: Noun compound interpretation is the task of uncovering the semantic relation between the components of a noun compound.
Approach: They propose an unsupervised method for identifying such relations between the components of a noun compound using pre-trained contextualized language models.
Outcome: The proposed method outperforms supervised approaches for free paraphrasing and prepositional paraphrases using pre-trained language models to uncover ‘missing’ words.
Reinforced Multi-task Approach for Multi-hop Question Generation (2020.coling-main)

Copied to clipboard

Challenge: Empirical evaluation shows our model to outperform the single-hop question generation models on both automatic evaluation metrics such as BLEU, METEOR, and ROUGE and human evaluation metrics for quality and coverage of the generated questions.
Approach: They propose a question-aware reward function to maximize the utilization of supporting facts in the context.
Outcome: The proposed model outperforms single-hop neural question generation models on automatic evaluation metrics and human evaluation metrics for quality and coverage of the generated questions.
Refer to the Reference: Reference-focused Synthetic Automatic Post-Editing Data Generation (2025.coling-main)

Copied to clipboard

Challenge: Existing approaches to synthetic APE data generation use source (src) sentences in a parallel corpus to obtain translations (mt) through an MT system and treat corresponding reference (ref) sentences as post-edits (pe).
Approach: They propose a reference-focused synthetic APE data generation technique that uses ‘ref’ instead of src’ sentences to obtain corrupted translations.
Outcome: The proposed technique improves on English-German, English-Russian, English -Marathi, English and Hindi language pairs.
Recommendation Chart of Domains for Cross-Domain Sentiment Analysis: Findings of A 20 Domain Study (2020.lrec-1)

Copied to clipboard

Challenge: Cross-domain sentiment analysis (CDSA) is a well-known problem in text analysis, but sufficient datasets may not be available for a domain to be trained.
Approach: They propose to use 11 similarity metrics to facilitate cross-domain sentiment analysis to identify the best domains for CDSA for a given target domain.
Outcome: The proposed approach performs better on 20 domain pairs and is validated by 11 similarity metrics.
Filtering Back-Translated Data in Unsupervised Neural Machine Translation (2020.coling-main)

Copied to clipboard

Challenge: Current state of the art approaches for unsupervised neural machine translation (NMT) use only monolingual data for training.
Approach: They propose an approach to filter back-translated data as part of the training process of unsupervised neural machine translation (NMT) they propose a weight component based on the quality of pseudo parallel sentence pairs generated in back-translation phase.
Outcome: The proposed approach improves the training performance of unsupervised neural machine translation systems by giving weight to good pseudo parallel sentence pairs in the back-translation phase.
From Recall to Creation: Generating Follow-Up Questions Using Bloom’s Taxonomy and Grice’s Maxims (2025.acl-industry)

Copied to clipboard

Challenge: In-car AI assistants struggle with multi-turn conversations and fail to handle cognitively complex follow-up questions.
Approach: They propose a framework that leverages Bloom's Taxonomy to generate follow-up questions with increasing cognitive complexity and a Gricean-inspired evaluation framework to assess their Logical Consistency, Informativeness, Relevance, and Clarity.
Outcome: The proposed framework validates both LLM-based and human evaluations and identifies the specific cognitive complexity level at which in-car AI assistants begin to falter information.
MMQA: A Multi-domain Multi-lingual Question-Answering Framework for English and Hindi (L18-1)

Copied to clipboard

Challenge: Existing work on multi-domain, multi-lingual question answering is limited to the same language.
Approach: They curate 500 articles in six different domains from the web and create question-answer pairs . they develop a deep learning based model for classifying an input question into coarse and finer categories .
Outcome: The proposed model accuracies 90.12% and 80.30% for coarse and finer classes . the proposed model is the first attempt to create multi-domain, multi-lingual question answering evaluation involving English and Hindi.
Incorporating Politeness across Languages in Customer Care Responses: Towards building a Multi-lingual Empathetic Dialogue Agent (2020.lrec-1)

Copied to clipboard

Challenge: Qualitative and quantitative analysis shows that our proposed model can converse in both the languages and the information shared between the languages helps in improving the performance of the overall system.
Approach: They propose a deep learning framework that can handle different languages and incorporate courteous behaviour in generic customer care responses in a multi-lingual scenario.
Outcome: The proposed model can converse in both languages and the information shared between the languages helps in improving the overall performance of the system.
A Unified Framework for Multilingual and Code-Mixed Visual Question Answering (2020.aacl-main)

Copied to clipboard

Challenge: Existing techniques for visual question answering focus on English questions, but many applications require a multilingual module.
Approach: They propose a deep learning framework for multilingual and code- mixed visual question answering . they create Hindi and Code-mixed VQA datasets by exploiting linguistic properties of these languages .
Outcome: The proposed model is capable of predicting answers from the questions in Hindi, English or Code- mixed (Hindi-English) languages.
ReDepress: A Cognitive Framework for Detecting Depression Relapse from Social Media (2025.emnlp-main)

Copied to clipboard

Challenge: Almost 50% of depression patients face the risk of going into relapse.
Approach: They propose to validate a social media dataset on depression relapse using cognitive theories of depression.
Outcome: The first clinically validated social media dataset focused on depression relapse comprises 204 Reddit users annotated by mental health professionals.
ConCodeEval: Evaluating Large Language Models for Code Constraints in Domain-Specific Languages (2025.acl-industry)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated potential in code generation and natural language understanding, but they struggle with code constraints.
Approach: They propose to use Large Language Models to handle constraints represented in code . they use JSON, YAML, XML, Python, and natural language to test their effectiveness .
Outcome: The proposed benchmark shows that LLMs can handle code constraints better than natural language . the results suggest that conscious choice of representations can lead to optimal use of LLM in enterprise use cases involving code constraints.
Cognition-aware Cognate Detection (2021.eacl-main)

Copied to clipboard

Challenge: Existing approaches to cognate detection use orthographic, phonetic and semantic similarity based features sets.
Approach: They propose a method for enriching feature sets with cognitive features extracted from gaze behaviour data from human readers’ gaze behaviour.
Outcome: The proposed method improves cognate detection performance by 10% and 12% over existing methods.
A Multimodal Corpus for Emotion Recognition in Sarcasm (2022.lrec-1)

Copied to clipboard

Challenge: sarcasm and emotion are often used in conversational systems to generate the right response.
Approach: They use a sarcastic expression dataset pre-annotated with 9 emotions to detect emotion . they identify and correct 343 incorrect emotion labels and label each sarkastic utterance with one of four sarcasm types.
Outcome: The proposed model outperforms state-of-the-art sarcasm detection methods by using a multimodal sarcastic detection dataset.
HiNER: A large Hindi Named Entity Recognition Dataset (2022.lrec-1)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a lowerlevel task that aims to provide class labels like Person, Location, Organisation, Time, and Number to words in free text.
Approach: They propose to use a standard-abiding Hindi NER dataset to analyze the annotations of a class of naming entities in free text.
Outcome: The proposed dataset achieves a weighted F1 score of 88.78 with all the tags and 92.22 when we collapse the tag-set.
Modelling Context Emotions using Multi-task Learning for Emotion Controlled Dialog Generation (2021.eacl-main)

Copied to clipboard

Challenge: Recent research has tackled this task using neural generative methods by augmenting emotion classes with the input sequences.
Approach: They propose to use a self-attention based encoder and a decoder with dot product attention mechanism to generate a viable response with a specified emotion.
Outcome: The proposed model outperforms baselines on automatic evaluation measures such as F1 and BLEU scores, thus resulting in more fluent and adequate responses.
The IIT Bombay English-Hindi Parallel Corpus (L18-1)

Copied to clipboard

Challenge: a corpus of 1.49 million parallel segments is available in the public domain . the corpus is the largest publicly available English-Hindi parallel corpus .
Approach: They present the IIT Bombay English-Hindi Parallel Corpus . they present a compilation of public and private parallel corpora .
Outcome: The corpus contains 1.49 million parallel segments, of which 694k were not previously available in the public domain.
COMMA-DEER: COmmon-sense Aware Multimodal Multitask Approach for Detection of Emotion and Emotional Reasoning in Conversations (2022.coling-1)

Copied to clipboard

Challenge: Mental health is a critical component of the United Nations’ Sustainable Development Goals (SDGs), particularly Goal 3 which aims to provide “good health and well-being”.
Approach: They propose a task of detecting emotional reasoning and accompanying emotions in conversations that is manually annotated at the utterance level.
Outcome: The proposed model achieves 6% accuracy and 4.62% accuracy on the emotion detection task and 3.56% accuracy, and 3.31% F1 on the ER detection task, compared to the existing state-of-the-art model.
Happy Are Those Who Grade without Seeing: A Multi-Task Learning Approach to Grade Essays Using Gaze Behaviour (2020.aacl-main)

Copied to clipboard

Challenge: Using gaze behaviour to solve automatic essay grading tasks is costly in terms of time and money.
Approach: They propose to collect gaze behaviour from 48 essays and learn gaze behaviour for the rest of the essays using a multi-task learning framework.
Outcome: The proposed approach achieves a statistically significant improvement over the state-of-the-art system for the essay sets where gaze data is available.
RG-VQA: Leveraging Retriever-Generator Pipelines for Knowledge Intensive Visual Question Answering (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to improve the reasoning capabilities of VQA systems are limited due to complexity of graph neural networks and end-to-end training.
Approach: They propose a method to integrate Dense Passage Retrievers with Vision Language Models to boost the reasoning capabilities of VQA systems.
Outcome: The proposed method outperforms human accuracy and GPT-4 in the ScienceQA dataset.
Part-of-speech Tagging for Extremely Low-resource Indian Languages (2024.findings-acl)

Copied to clipboard

Challenge: Modern natural language processing systems thrive when given access to large datasets, but a large fraction of the world’s languages are not privy to such benefits due to sparse documentation and inadequate digital representation.
Approach: They propose a parallel part-of-speech evaluation dataset for Angika, Magahi, Bhojpuri and Hindi.
Outcome: The proposed approach improves F1 scores by up to 8% on Angika, Magahi, Bhojpuri and Hindi while ignoring the tokenization challenge.
“A Little is Enough”: Few-Shot Quality Estimation based Corpus Filtering improves Machine Translation (2023.findings-acl)

Copied to clipboard

Challenge: Quality Estimation (QE) is the task of evaluating the quality of a translation when reference translation is unavailable.
Approach: They propose a Quality Estimation based Filtering approach to extract high-quality parallel data from the pseudo-parallel corpus.
Outcome: The proposed approach improves the machine translation system performance by up to 1.8 BLEU points over the baseline model.
HindiMD: A Multi-domain Corpora for Low-resource Sentiment Analysis (2022.lrec-1)

Copied to clipboard

Challenge: Social media platforms such as Twitter and Facebook are a new channel of information dissemination for many negative groups for recruitment.
Approach: They propose to use a social media sentiment analysis corpus annotated with the sentiment classes positive, negative and neutral to investigate the polarity of user-expressed opinions.
Outcome: The proposed model is based on a set of benchmark datasets for sentiment analysis across a range of domains and languages.
STORiCo: Storytelling TTS for Hindi with Character Voice Modulation (2024.eacl-short)

Copied to clipboard

Challenge: Existing datasets for read speech for Hindi lack expressiveness and character voice consistency.
Approach: They propose to use a Hindi text-to-speech (TTS) dataset to train a multi-speaker model on the single-sector data and propose to improve expressiveness and character voice consistency.
Outcome: The proposed model improves expressiveness and character voice consistency compared to the baseline single-speaker model.
KGVL-BART: Knowledge Graph Augmented Visual Language BART for Radiology Report Generation (2023.eacl-main)

Copied to clipboard

Challenge: Timely generation of radiology reports and diagnoses is a challenge worldwide due to the enormous number of cases and shortage of radiologists.
Approach: They propose a Knowledge Graph Augmented Vision Language BART model that takes two chest X-ray images and outputs a report with patient-specific findings.
Outcome: The proposed model outperforms state-of-the-art transformer-based models on scoring metrics.
Sentence Level Temporality Detection using an Implicit Time-sensed Resource (L18-1)

Copied to clipboard

Challenge: Temporal sense detection of any word is an important aspect for detecting temporality at the sentence level.
Approach: They build a temporal resource based on a semi-supervised learning approach . they use past, present, future, neutral and atemporal senses to tag sentences .
Outcome: The proposed resource is based on a semi-supervised learning approach . it is used to tag sentences with past, present and future temporal senses .
Affective Retrofitted Word Embeddings (2022.aacl-main)

Copied to clipboard

Challenge: Word embeddings do not capture affective dimensions of valence, arousal, and dominance . valency, valance, and adolescence are present in words, but are not represented in text .
Approach: They propose a method for updating word embeddings for affective meaning . they use a non-linear transformation function that maps pre-trained embedders to an affective vector space .
Outcome: The proposed method improves inter-cluster and intra-c cluster distances for emotion-bearing words.
MemeGuard: An LLM and VLM-based Framework for Advancing Content Moderation via Meme Intervention (2024.acl-long)

Copied to clipboard

Challenge: Existing studies on content moderation of toxic memes focus on text-based content . current research neglects the widespread influence of multimodal content like memes .
Approach: They propose a framework leveraging Large Language Models and Visual Language Model (VLMs) for meme intervention.
Outcome: The proposed framework enables users to generate relevant and effective responses to toxic memes.
Extraction of Message Sequence Charts from Software Use-Case Descriptions (N19-2)

Copied to clipboard

Challenge: Software Requirement Specification documents provide natural language descriptions of the core functional requirements as a set of use-cases.
Approach: They propose a linguistic knowledge-based approach to extract software requirements from use-cases using a textual representation of the core functional requirements.
Outcome: The proposed method performs better than existing techniques and improves performance.
Harnessing Cross-lingual Features to Improve Cognate Detection for Low-resource Languages (2020.coling-main)

Copied to clipboard

Challenge: a study of 14 Indian languages shows that cognates can be detected by word embeddings . cognates are variants of the same lexical form across languages .
Approach: They propose to use cross-lingual word embeddings to detect cognates among 14 Indian languages . they then evaluate the impact of their method on neural machine translation .
Outcome: The proposed method improves on a dataset of 12 Indian languages . it also improves quality of the extracted cognates by up to 2.76 BLEU .
Hollywood Identity Bias Dataset: A Context Oriented Bias Analysis of Movie Dialogues (2022.lrec-1)

Copied to clipboard

Challenge: Movies reflect society and also hold power to transform opinions.
Approach: They propose to annotate movie scripts for identity bias using a dataset that is annotated for gender, race/ethnicity, religion, age, occupation, LGBTQ, and other .
Outcome: The proposed dataset contains dialogue turns annotated for gender, race/ethnicity, religion, age, occupation, LGBTQ, and other, which contains biases like body shaming, personality bias, etc.
“You are Beautiful, Body Image Stereotypes are Ugly!” BIStereo: A Benchmark to Measure Body Image Stereotypes in Language Models (2025.findings-acl)

Copied to clipboard

Challenge: BIStereo is a suite of language models that uncover body image stereotypes in language models.
Approach: They propose a metric, TriSentBias, that captures the biased preferences of LMs towards a certain body type over others.
Outcome: The proposed metric captures biased preferences of LMs towards a certain body type over others.
Courteously Yours: Inducing courteous behavior in Customer Care responses using Reinforced Pointer Generator Network (N19-1)

Copied to clipboard

Challenge: In order to ensure customer satisfaction and retention, it is imperative for customer care agents and chatbots to be cordial and emphatic to the customer.
Approach: They propose a deep learning framework that automatically transforms neutral customer care responses into courteous replies by stylistic transfer.
Outcome: The proposed model can generate courteous expressions consistent with the emotional state of the customer while preserving the content.
PUB: A Pragmatics Understanding Benchmark for Assessing LLMs’ Pragmatics Capabilities (2024.findings-acl)

Copied to clipboard

Challenge: Pragmatics understanding is not well studied in LLMs, but their understanding of pragmatics is lacking.
Approach: They propose to use a dataset to measure LLMs' understanding of pragmatics to evaluate their models.
Outcome: The proposed dataset includes 14 tasks in four pragmatics phenomena, namely; Implicature, Presupposition, Reference, and Deixis.
Quality Estimation-Assisted Automatic Post-Editing (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing APE and QE combination strategies have not shown significant performance gains in the field of automatic post-editing (APE).
Approach: They propose to train a model on APE and QE tasks to improve the APE performance by using a multi-task learning methodology that treats both tasks as a 'bargaining game' they also investigate various existing combination strategies and show that their approach achieves state-of-the-art performance for a ‘distant’ language pair, viz., English-Marathi.
Outcome: The proposed model improves on two different language pairs, viz., English-Marathi and English-German.
“Knowledge is Power”: Constructing Knowledge Graph of Abdominal Organs and Using Them for Automatic Radiology Report Generation (2023.acl-industry)

Copied to clipboard

Challenge: conventional radiology workflows involve dictating diagnosis to transcriptionists, which is prone to delay and error.
Approach: They propose to generate a set of knowledge graphs from a large collection of free-text radiology reports and use them to generate automatic radiology report generation.
Outcome: The proposed model improves the reported BLEU-3, ROUGE-L, METEOR, and CIDEr scores by 2%, 4%, 2% and 2% respectively.
From Perception to Reasoning: Enhancing Vision-Language Models for Mobile UI Understanding (2025.findings-acl)

Copied to clipboard

Challenge: Accurately grounding visual and textual elements within mobile user interfaces remains a challenge for Vision-Language Models (VLMs).
Approach: They propose a mobile UI understanding model trained on a dataset specifically tailored for mobile screen understanding and grounding.
Outcome: The proposed model achieves significant gains in accuracy across all perception tasks and on reasoning benchmarks.
TAP-DLND 1.0 : A Corpus for Document Level Novelty Detection (L18-1)

Copied to clipboard

Challenge: Detecting novelty of an entire document is an AI frontier problem . present state-of-the-art text matching techniques are unable to process such redundancy.
Approach: They propose a document-level novelty detection resource that can be used to benchmark techniques . they crawl news documents across several domains and use it to find out whether they contain new information .
Outcome: The proposed dataset is compared with a standard system for document novelty detection . the proposed system can detect elements that have not appeared before, or new or original .
Multi-domain Tweet Corpora for Sentiment Analysis: Resource Creation and Evaluation (2020.lrec-1)

Copied to clipboard

Challenge: a huge amount of content is being generated every day due to the pervasiveness of social media.
Approach: They firstly create a multi-domain tweet sentiment corpora and then establish a deep neural network based baseline framework to address the above mentioned issues.
Outcome: The proposed dataset achieves 84.65% accuracy for sentiment analysis using a neural network, long short term memory, and gated recurrent unit (GRU).
IndiBias: A Benchmark Dataset to Measure Social Biases in Language Models for Indian Context (2024.naacl-long)

Copied to clipboard

Challenge: Existing benchmark datasets focus on English language and the Western context, leaving a void for a reliable dataset that encapsulates India’s unique socio-cultural nuances.
Approach: They propose to use CrowS-Pairs to create a benchmark dataset that captures and evaluates social biases in Large Language Models (LLMs).
Outcome: The proposed dataset is available in English and Hindi and leverages LLMs ChatGPT and InstructGPT to augment the existing dataset with diverse societal biases and stereotypes prevalent in India.
Towards a Standardized Dataset for Noun Compound Interpretation (L18-1)

Copied to clipboard

Challenge: Noun compounds are interesting constructs in Natural Language Processing . lack of standardized set of relation inventories and annotated datasets hinders interpretation .
Approach: They propose a dataset that uses FrameNet as its semantic relation inventory to examine noun compounds.
Outcome: The proposed dataset is linguistically grounded and uses FrameNet as its semantic relation inventory.
EM-PERSONA: EMotion-assisted Deep Neural Framework for PERSONAlity Subtyping from Suicide Notes (2022.coling-1)

Copied to clipboard

Challenge: Suicide continues to be one of the significant causes of death worldwide . EMotion-assisted personality subtyping is a novel approach to identify personality traits from suicide notes .
Approach: They propose to use a PERSONAlity Detection Framework to identify personality traits from suicide notes and annotate them using a benchmark dataset.
Outcome: The proposed method outperforms baselines on comprehensive evaluation using multiple state-of-the-art systems.
ETF: An Entity Tracing Framework for Hallucination Detection in Code Summaries (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in large language models have significantly enhanced their ability to understand both natural language and code, but are prone to hallucinations.
Approach: They propose a first-of-its-kind dataset, CodeSumEval, with 10K samples, curated specifically for hallucination detection in code summarisation.
Outcome: The proposed framework has a 73% F1 score and is curated specifically for detection of hallucinations in code summarisation.
Identification of Alias Links among Participants in Narratives (P18-2)

Copied to clipboard

Challenge: Identifying distinct and independent participants in a narrative is crucial for many NLP applications.
Approach: They propose an approach based on linguistic knowledge for identification of aliases mentioned using proper nouns, pronouns or noun phrases with common noun headword.
Outcome: The proposed approach performs better than the state-of-the-art approach on four diverse history narratives of varying complexity.
Predict and Use: Harnessing Predicted Gaze to Improve Multimodal Sarcasm Detection (2023.emnlp-main)

Copied to clipboard

Challenge: sarcasm detection depends on content spoken, tonality, facial expressions, context, and personal traits like language proficiency and cognitive capabilities.
Approach: They propose to use synthetic gaze data to improve sarcasm detection in conversational context . they collect gaze features for 20% of data instances and use them to predict gaze features .
Outcome: The proposed model improves performance on a conversational dataset using gaze features . it achieves a gain of 6.6% points on the complete dataset with only predicted gaze features.
Judicious Selection of Training Data in Assisting Language for Multilingual Neural NER (P18-2)

Copied to clipboard

Challenge: Existing approaches to improve NER performance add training data from one or more assisting languages to the primary language.
Approach: They propose a metric based on symmetric KL divergence to filter out highly divergent training instances in the assisting language.
Outcome: The proposed method improves NER performance in many languages, including those with limited training data.
Product Description and QA Assisted Self-Supervised Opinion Summarization (2024.findings-naacl)

Copied to clipboard

Challenge: Existing methods to generate opinion summarization without supervised training data are limited due to the lack of additional sources.
Approach: They propose a synthetic dataset creation strategy that leverages reviews and additional sources to generate a pseudo-summary.
Outcome: The proposed approach achieves 14.5% improvement in ROUGE-1 F1 over existing models.
In-context Mixing (ICM): Code-mixed Prompts for Multilingual LLMs (2024.acl-long)

Copied to clipboard

Challenge: In-context mixing is a prompting technique for effective in-contact learning with multilingual large language models.
Approach: They propose a prompting technique called in-context mixing for effective in-constext learning with multilingual large language models.
Outcome: The proposed prompts perform better with multilingual large language models.
With Prejudice to None: A Few-Shot, Multilingual Transfer Learning Approach to Detect Social Bias in Low Resource Languages (2023.findings-acl)

Copied to clipboard

Challenge: Currently, the majority of social bias datasets available are in English and this inhibits progress on social bias detection in low-resource languages.
Approach: They propose a dataset for social bias detection in Hindi and investigate multilingual transfer learning using publicly available English, Italian, and Korean datasets.
Outcome: The proposed dataset is compared with a dataset available in English, Italian, and Korean using multilingual models.
Solving Data Sparsity for Aspect Based Sentiment Analysis Using Cross-Linguality and Multi-Linguality (N18-1)

Copied to clipboard

Challenge: Efficient word representations play an important role in solving various problems related to NLP, data mining, text mining etc.
Approach: They propose to leverage bilingual word embeddings learned through a parallel corpus to minimize the effect of data sparsity.
Outcome: The proposed model is tested against state-of-the-art methods in two experimental setups.
A Multi-task Learning Framework for Quality Estimation (2023.findings-acl)

Copied to clipboard

Challenge: Conventional approaches to QE involve training separate models at different levels of granularity viz., word-level, sentence-level and document-level .
Approach: They propose to train a single model for sentence-level and word-level QE tasks in a multi-task learning framework and compare them to baseline models.
Outcome: The proposed model improves on the single-pair, multi-patch, and zero-shot settings.
Identifying Transferable Information Across Domains for Cross-domain Sentiment Classification (P18-1)

Copied to clipboard

Challenge: Cross-domain sentiment classification is challenging due to polarity orientation and significance differences . supervised learning algorithms have to be re-trained on every new domain .
Approach: They propose that words that do not change their polarity and significance represent transferable information across domains for cross-domain sentiment classification.
Outcome: The proposed method improves cross-domain sentiment classification performance by identifying polarity-preserving significant words across domains.
Sarcasm Target Identification: Dataset and An Introductory Approach (L18-1)

Copied to clipboard

Challenge: Past work on sarcasm detection has focused on identifying the sarcasm target of ridicule in a sarkastic text.
Approach: They propose a task of extracting the sarcastic target of ridicule from a sarcastical text using a manually annotated dataset and an automatic approach.
Outcome: The proposed approach establishes the viability of sarcasm target identification and will serve as a baseline for future work.
Knowledge Graph - Deep Learning: A Case Study in Question Answering in Aviation Safety Domain (2022.lrec-1)

Copied to clipboard

Challenge: Existing Question Answering systems for commercial aviation use a large number of documents . a Knowledge Graph (KG) guided Deep Learning (DL) based system can be used to query the documents based on accident reports .
Approach: They propose a Knowledge Graph (KG) guided Deep Learning (DL) based Question Answering system to cater to these requirements.
Outcome: The proposed system achieves 7% and 40% increase in accuracy over existing systems.
Why We Feel What We Feel: Joint Detection of Emotions and Their Opinion Triggers in E-commerce (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing research has not explored the joint task of emotion detection and explanatory span identification in e-commerce reviews.
Approach: They propose a joint task unifying Emotion detection and Opinion Trigger extraction (EOT) which explicitly models the relationship between causal text spans (opinion triggers) and affective dimensions (emotion categories).
Outcome: The proposed framework surpasses zero-shot and chain-of-thought techniques across e-commerce domains.
Sentiment and Emotion help Sarcasm? A Multi-task Learning Framework for Multi-Modal Sarcasm, Sentiment and Emotion Analysis (2020.acl-main)

Copied to clipboard

Challenge: Existing systems for sarcasm detection are limited by the use of sarcasm . sarasm is often used to convey thinly veiled disapproval humorously.
Approach: They propose a multi-task deep learning framework to solve sarcasm problems simultaneously . they manually annotate a sarcsm dataset with sentiment and emotion classes .
Outcome: The proposed framework is able to solve sarcasm, sentiment and emotion problems in a multi-modal conversational scenario.
Replace and Report: NLP Assisted Radiology Report Generation (2023.findings-acl)

Copied to clipboard

Challenge: Clinical practice frequently uses medical imaging for diagnosis and treatment.
Approach: They propose a template-based approach to generate radiology reports from radiographs . they use multilabel image classifiers to generate tags, pathological descriptions from tags .
Outcome: The proposed method improves on the most popular radiology report datasets.
Zero-shot Disfluency Detection for Indian Languages (2022.coling-1)

Copied to clipboard

Challenge: Disfluency correction models can help alleviate this problem, but the unavailability of labeled data in low-resource languages impairs progress.
Approach: They propose to use a pretrained multilingual model to detect zero-shot disfluency in Indian languages.
Outcome: The proposed model achieves F1 scores of 75 and higher on five disfluency types across four languages.
A Deep Neural Network based Approach for Entity Extraction in Code-Mixed Indian Social Media Text (L18-1)

Copied to clipboard

Challenge: a huge number of people use social media to express and exchange information in their own languages.
Approach: They propose to use a code-mixed environment to extract higher level features from text . they use 'gadget' algorithm that automatically discovers higher level feature from text.
Outcome: The proposed approach is generic and does not make use of handcrafted features or rules.
Context-aware Interactive Attention for Multi-modal Sentiment and Emotion Analysis (D19-1)

Copied to clipboard

Challenge: Multi-modal analysis is a field emerging in the fields of natural language processing, computer vision and speech processing . multimodal analysis uses a variety of information from multiple sources to build efficient systems . acoustic and visual information can provide better information for classification decisions .
Approach: They propose a recurrent neural network based approach for multi-modal sentiment and emotion analysis . they employ a context-aware attention module to exploit the correspondence among neighboring utterances .
Outcome: The proposed model learns inter-modal interaction among participating modalities through auto-encoder mechanism . it is compared with existing state-of-the-art models on five standard multi-modal affect analysis datasets .
Medical Sentiment Analysis using Social Media: Towards building a Patient Assisted System (L18-1)

Copied to clipboard

Challenge: a study conducted by the pew Internet & American Life Project 1 shows that almost 80 percent of Internet users have explored health-related topic online.
Approach: They propose to crawl medical forums with opinions about medical condition self narrated by users.
Outcome: The proposed system is based on opinions about medical condition self-narrated by users on medical forums.
GenEx: A Commonsense-aware Unified Generative Framework for Explainable Cyberbullying Detection (2023.emnlp-main)

Copied to clipboard

Challenge: a significant gap exists in understanding code-mixed languages and the need for explainability in this context.
Approach: They propose to annotate posts with four labels to identify bullies in code-mixed languages . they propose to use a generative framework to reimagine the multitask problem as a text-to-text generation task.
Outcome: The proposed model outperforms baseline models and state-of-the-art models on the BullyExplain dataset.
Morphology Injection for English-Malayalam Statistical Machine Translation (L18-1)

Copied to clipboard

Challenge: Statistical Machine Translation fails to handle the rich morphology when translating into morphologically rich language.
Approach: They propose a method to generate unseen morphological forms from the parallel corpus . they propose morphology injection method to enrich the corpus with generated morphologies .
Outcome: The proposed method improves the quality of the translation in English-Malayalam.
ScholarlyRead: A New Dataset for Scientific Article Reading Comprehension (2020.lrec-1)

Copied to clipboard

Challenge: Existing studies on MRC on scholarly articles have focused on general domain datasets of news articles and elementary school-level storybooks.
Approach: They propose to generate automatic questions from span-of-word-based scholarly articles’ Reading Comprehension dataset with approximately 10K manually checked passage-question-answer instances.
Outcome: The proposed model yields the F1 score of 37.31% and is useful for building Question-Answering (QA) systems on scientific articles.
A Platform for Event Extraction in Hindi (2020.lrec-1)

Copied to clipboard

Challenge: Event Extraction is an important task in the widespread field of NLP, but there is no benchmark setup in Hindi.
Approach: They propose an Event Extraction framework for Hindi language and develop deep learning based models to set as the baselines.
Outcome: The proposed framework crawls more than seventeen hundred disaster related Hindi news articles from various news sources.
Giving the Old a Fresh Spin: Quality Estimation-Assisted Constrained Decoding for Automatic Post-Editing (2025.naacl-short)

Copied to clipboard

Challenge: Existing methods to improve automatic post-editing (APE) systems struggle with over-correction, despite the principle of minimal editing.
Approach: They propose a method that incorporates word-level Quality Estimation (QE) information during the decoding process.
Outcome: The proposed method improves on English-German, English-Hindi, and English-Marathi language pairs, with TER gains of 0.65, 1.86, and 1.44 points, respectively.
Analysing cross-lingual transfer in lemmatisation for Indian languages (2020.coling-main)

Copied to clipboard

Challenge: Inference-based scripts such as Abjad are difficult for cross-lingual models to learn in extremely low resource scenarios.
Approach: They evaluate cross-lingual approaches for low resource languages and compare their performance against other models using different linguistic factors.
Outcome: The proposed model on six low resource languages from two different families is compared with monolingual models on morphologically rich Indian languages.
Stereotype Detection as a Catalyst for Enhanced Bias Detection: A Multi-Task Learning Approach (2025.findings-acl)

Copied to clipboard

Challenge: a new study addresses bias and stereotypes in language models by exploring how learning them together improves performance.
Approach: They propose a dataset for bias and stereotype detection that integrates religion, gender, socio-economic status, race, profession, and others.
Outcome: The proposed dataset compares encoder-only models and fine-tuned decoder- only models . the results show that learning stereotypes together improves bias detection .
PoliSe: Reinforcing Politeness Using User Sentiment for Customer Care Response Generation (2022.coling-1)

Copied to clipboard

Challenge: Human-machine interactions have increased rapidly assisting humans in their everyday lives.
Approach: They propose to automatically identify the sentiment of the user and transform the neutral responses into polite responses conforming to the sentiment and the conversational history.
Outcome: The proposed approach achieves superior performance compared to baseline models.
Multi-task Learning for Multi-modal Emotion Recognition and Sentiment Analysis (N19-1)

Copied to clipboard

Challenge: Existing frameworks for sentiment and emotion analysis are not efficient for inter-task learning.
Approach: They propose a multi-task learning framework that performs sentiment and emotion analysis together.
Outcome: The proposed framework improves on a CMU-MOSEI dataset for sentiment and emotion analysis.
A Shoulder to Cry on: Towards A Motivational Virtual Assistant for Assuaging Mental Agony (2022.naacl-main)

Copied to clipboard

Challenge: Mental health disorders are one of the primary causes of disability worldwide . lack of qualified and competent mental health professionals is a major problem . we propose a virtual assistant that can act as the first point of contact and comfort for mental health patients.
Approach: They propose a virtual assistant that can act as the first point of contact and comfort for mental health patients.
Outcome: The proposed system outperforms baselines in the evaluation of 7k dyadic conversations from a peer-to-peer support platform.
How effective are VLMs in assisting humans in inferring the quality of mental models from Multimodal short answers? (2026.eacl-long)

Copied to clipboard

Challenge: Mental models are "basic units of coherently structured knowledge" but stu-dents do not always construct coherent mental models, which can limit conceptual understanding.
Approach: They propose an approach that infers the quality of students’ mental models from their multimodal responses using concept graphs as an analytical framework.
Outcome: The proposed model infers the quality of students’ mental models from their multimodal responses using concept graphs as an analytical framework.
Angel: Enterprise Search System for the Non-Profit Industry (2023.emnlp-industry)

Copied to clipboard

Challenge: Non-profit industry needs a system for accurately matching fund-seekers with fund-givers aligned in cause and target beneficiary group.
Approach: They propose a search system that takes a fund-giver’s mission description as input and returns a ranked list of fund-seekers as output.
Outcome: The proposed system improves on the non-profit evaluation dataset and the state-of-the-art model.
Addressing Bias and Hallucination in Large Language Models (2024.lrec-tutorials)

Copied to clipboard

Challenge: This tutorial provides a comprehensive overview of two critical aspects of Large Language Models: bias and hallucination.
Approach: This tutorial provides an overview of two critical aspects of Large Language Models: bias and hallucination.
Outcome: This tutorial delves into the complex dimensions of Large Language Models (LLMs) it outlines ethical considerations pertinent to their development and discusses hallucination, a prevalent issue in generative AI systems such as LLMs.
Fine-Grained Temporal Orientation and its Relationship with Psycho-Demographic Correlates (N18-1)

Copied to clipboard

Challenge: Temporal orientation refers to an individual’s tendency to connect to the psychological concepts of past, present or future and affects personality, motivation, emotion, decision making and stress coping processes.
Approach: They propose to use a minimally supervised method to classify tweets in one of three temporal categories, past, present, and future, and a deep bi-directional long-term memory (BLSTM) to measure correlation between sentiment view of temporal orientation and different psycho-demographic factors.
Outcome: The proposed method achieves 78.27% accuracy on a manually created test set.
Can Taxonomy Help? Improving Semantic Question Matching using Question Taxonomy (C18-1)

Copied to clipboard

Challenge: Existing QA systems that answer factual questions with short answers are rare in practice.
Approach: They propose a proposed two-layered taxonomy technique for semantic question matching . they augment state-of-the-art deep learning models with question classes from a deep learning based question classifier .
Outcome: The proposed technique achieves state-of-the-art on an open-domain dataset.
Meta-Learning based Deferred Optimisation for Sentiment and Emotion aware Multi-modal Dialogue Act Classification (2022.aacl-main)

Copied to clipboard

Challenge: Empirically, we show that the optimisation of multi-modal DAC, SA and ER tasks produces better results compared to its different counterparts.
Approach: They propose a dual attention mechanism that integrates sentiment tags into a multi-modal conversational framework that integrate modal attentions and multiple loss optimization.
Outcome: The proposed framework integrates sentiment tags for each utterance and learns generalized features across multiple tasks.
Recon, Answer, Verify: Agents in Search of Truth (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing benchmark datasets suffer from leakage or evidence incompleteness, limiting the realism of current evaluations.
Approach: They propose an agentic framework that iteratively generates and answers sub-questions to verify different aspects of the claim before finally generating the label.
Outcome: The proposed system outperforms existing methods by 57.5% on Politi-Fact-Only and 3.05% on the widely used HOVER datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations