Papers by Rada Mihalcea

98 papers
PAIR: Prompt-Aware margIn Ranking for Counselor Reflection Scoring in Motivational Interviewing (2022.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to provide constructive feedback to counselors are limited by the time and cost involved.
Approach: They propose a system that takes as input a client prompt and a counselor response and outputs a score indicating the level of reflection in the counselor response.
Outcome: The proposed model outperforms baselines on different metrics and can be used to provide useful feedback to counseling trainees.
Causally Testing Gender Bias in LLMs: A Case Study on Occupational Bias (2025.findings-naacl)

Copied to clipboard

Challenge: Existing studies have shown that large language models can cause harmful, human-like biases against various demographics.
Approach: They propose a causal formulation for bias measurement in generative language models based on a list of desiderata for designing robust bias benchmarks and a bias-measuring procedure to investigate occupational gender bias.
Outcome: The proposed framework is generalizable and can be extended to include other datasets.
WhyAct: Identifying Action Reasons in Lifestyle Vlogs (2021.emnlp-main)

Copied to clipboard

Challenge: Existing systems for action recognition rely on pattern memorization and do not understand the action.
Approach: They propose a multimodal model that leverages visual and textual information to automatically infer the reasons corresponding to an action presented in the video.
Outcome: The proposed model leverages visual and textual information to automatically infer the reasons corresponding to an action presented in the video.
Towards Algorithmic Fidelity: Mental Health Representation across Demographics in Synthetic vs. Human-generated Data (2024.lrec-main)

Copied to clipboard

Challenge: Synthetic data generation has the potential to impact domains with scarce data, but we need to understand how different demographics are represented in it.
Approach: They develop a procedure to generate depression data using GPT-3 and analyze it to uncover the types of stressors it assigns to demographic groups.
Outcome: The proposed procedure produces depression data using GPT-3, and compares it to a human-generated dataset.
Towards Extracting Medical Family History from Natural Language Interactions: A New Dataset and Baselines (D19-1)

Copied to clipboard

Challenge: Using dialog agents, we can collect family history data from in-person consultations and crowdsource it to a genetic counselor.
Approach: They propose to use natural language interactions annotated with medical family histories to collect information from a genetic counselor and crowdsourcing.
Outcome: The proposed system averages 0.87 on complex sentences on the targeted relations.
Persuasion at Play: Understanding Misinformation Dynamics in Demographic-Aware Human-LLM Interactions (2026.eacl-long)

Copied to clipboard

Challenge: Existing challenges in misinformation exposure and susceptibility vary across demographics.
Approach: They propose a framework that investigates the bidirectional persuasion dynamics between LLMs and humans when exposed to misinformation.
Outcome: The proposed framework analyzes the spread of misinformation under persuasion among demographic-oriented LLM agents.
MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations (P19-1)

Copied to clipboard

Challenge: Emotion recognition in conversations has gained popularity due to its potential applications. Until now, a large multimodal multi-party emotional conversational database containing more than two speakers per dialogue was missing.
Approach: They propose to extend and enhance EmotionLines by combining 13,000 utterances from Friends dialogues with emotion and sentiment labels.
Outcome: The proposed dataset contains about 13,000 utterances from 1,433 dialogues from the TV-series Friends.
Building Location Embeddings from Physical Trajectories and Textual Representations (2020.aacl-main)

Copied to clipboard

Challenge: Using a dataset consisting of the location trajectories of 729 students over a seven month period, we investigate whether embeddings can represent aspects such as location presence or location functionality.
Approach: They propose to use location embeddings to generate embeddables of sequences of locations a student has visited to identify surface properties captured in the representations.
Outcome: The proposed models can be used to predict depression levels and area of study, and can be applied to complex tasks such as predicting area of studies and depression levels.
Compositional Demographic Word Embeddings (2020.emnlp-main)

Copied to clipboard

Challenge: Word embeddings are usually derived from corpora containing text from many individuals . however, they cannot account for user-specific word preferences, such as using the same word in different ways across contexts.
Approach: They propose a new form of personalized word embeddings that use demographic-specific word representations derived compositionally from full or partial demographic information for a user.
Outcome: The proposed representations outperform generic representations on two English language tasks.
Women’s Syntactic Resilience and Men’s Grammatical Luck: Gender-Bias in Part-of-Speech Tagging and Dependency Parsing (P19-1)

Copied to clipboard

Challenge: linguistic studies have shown the prevalence of various lexical and grammatical patterns in texts authored by a person of a particular gender, but models for part-of-speech tagging and dependency parsing have not adapted to account for these differences.
Approach: They annotate the Wall Street Journal part of the Penn Treebank with the gender information of the articles’ authors and build taggers and parsers trained on this data.
Outcome: The proposed model can account for gendered differences in syntactic tasks and highlight future venues for developing more accurate taggers and parsers.
Revisiting LLM Value Probing Strategies: Are They Robust and Expressive? (2025.emnlp-main)

Copied to clipboard

Challenge: Existing value probing methods that capture in-context information and predict models’ real-world actions are limited and lack systematic comparisons.
Approach: They compare three widely used value probing methods: token likelihood, sequence perplexity, and text generation.
Outcome: The proposed methods exhibit large variances under non-semantic perturbations in prompts and option formats, with sequence perplexity being the most robust overall.
Knowledge Enhanced Reflection Generation for Counseling Dialogues (2022.acl-long)

Copied to clipboard

Challenge: Using retrieval and generative methods, we generate responses using commonsense and domain knowledge.
Approach: They propose a pipeline that collects domain knowledge through web mining and a model that incorporates knowledge generated by COMET using soft positional encoding and masked self-attention.
Outcome: The proposed pipeline collects domain knowledge through web mining and incorporates knowledge generated by COMET using soft positional encoding and masked self-attention.
Benchmarking and Improving LLM Robustness for Personalized Generation (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluations focus on whether a model’s responses align with a user’s preferences, but factuality is an important yet overlooked dimension.
Approach: They propose a scalable framework for evaluating robustness of large language models in personalization and a new dataset, PERGData.
Outcome: The proposed framework improves robustness by 25% across models.
Acoustic Individual Identification of White-Faced Capuchin Monkeys Using Joint Multi-Species Embeddings (2025.acl-short)

Copied to clipboard

Challenge: acoustic identification of animals is an essential task for conservation and wildlife monitoring . but, many methods for automatic identification are hindered by lack of data .
Approach: They explore cross-species pre-training to address the task of individual classification in white-faced capuchin monkeys.
Outcome: The proposed methods can be used to identify calls from individual monkeys using acoustic embeddings from birds and humans.
FIBER: Fill-in-the-Blanks as a Challenging Video Understanding Evaluation Framework (2022.acl-long)

Copied to clipboard

Challenge: Existing video understanding evaluation frameworks that use fill-in-the-blanks do not reflect real-world tasks.
Approach: They propose to use fill-in-the-blanks as a video understanding evaluation framework and introduce a novel dataset that collects multiple perspectives on the same video.
Outcome: The proposed framework does not share the weaknesses of the current state-of-the-art language-informed video understanding tasks, namely: (1) video question answering using multiple-choice questions, where models perform relatively well because they exploit linguistic biases in the task formulation; (2) video captioning, which relies on an open-ended evaluation framework that is often inaccurate because system answers may be perceived as incorrect if they differ in form from the ground truth.
COSMIC: COmmonSense knowledge for eMotion Identification in Conversations (2020.findings-emnlp)

Copied to clipboard

Challenge: Current methods for emotion recognition in conversations often face difficulties in context propagation, emotion shift detection, and differentiating between related emotion classes.
Approach: They propose a framework that incorporates mental states, events, and causal relations to learn interactions between interlocutors participating in a conversation.
Outcome: The proposed framework improves on four conversational benchmark datasets.
Bridging the Digital Divide: Performance Variation across Socio-Economic Factors in Vision-Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Among the minority groups under-represented in AI, data from low-income households are often overlooked in data collection and model evaluation.
Approach: They evaluate the performance of a vision-language model on a geo-diverse dataset . they highlight insights that can help mitigate these issues and propose actionable steps for economic-level inclusive AI development.
Outcome: The proposed model performs lower for the poorer groups than the wealthier groups across topics and countries.
Democratic or Authoritarian? Probing a New Dimension of Political Biases in Large Language Models (2026.eacl-long)

Copied to clipboard

Challenge: Prior work on LLM biases focused on socio-demographic and left–right political dimensions, but little attention has been paid to how they align with broader geopolitical value systems.
Approach: They propose a method to assess how LLMs align with broader geopolitical value systems, particularly the democracy–authoritarianism spectrum.
Outcome: The proposed method combines the F-scale, FavScore and role-model probing to assess which figures are cited as general role models by LLMs.
Exploring the Role of Context in Utterance-level Emotion, Act and Intent Classification in Conversations: An Empirical Study (2021.findings-acl)

Copied to clipboard

Challenge: utterance-level dialogue understanding tasks are often performed at utterrance level and are often conjoined together under the umbrella of utterence-level dialog understanding.
Approach: They propose to use a contextual utterance-level dialogue understanding baseline as a strong framework for six dialogue-understanding tasks.
Outcome: The proposed framework can be easily adapted for other tasks for similar purposes.
Identifying Visible Actions in Lifestyle Vlogs (P19-1)

Copied to clipboard

Challenge: Existing methods for identifying human actions in videos are limited by the number of visual depictions in the videos.
Approach: They propose a multimodal algorithm that leverages visual and linguistic clues to automatically infer which actions are visible in a video.
Outcome: The proposed algorithm can identify actions visible in video while verbally describing them.
MuSE: a Multimodal Dataset of Stressed Emotion (2020.lrec-1)

Copied to clipboard

Challenge: Existing studies on the effects of stress and emotion on the production and perception of emotion are understudied.
Approach: They propose to use a multimodal stressed emotion dataset to study the interplay between the presence of stress and expressions of affect.
Outcome: The proposed dataset combines emotion and stress classification with annotations for the emotional content of the recordings.
Language Guided Visual Question Answering: Elevate Your Multimodal Language Model Using Knowledge-Enriched Prompts (2023.findings-emnlp)

Copied to clipboard

Challenge: Visual question answering (VQA) is a task that requires an understanding of both the image and the question to provide a natural language answer.
Approach: They propose a multimodal framework that leverages language guidance to answer questions more accurately.
Outcome: The proposed framework improves on the multi-choice question-answering task using CLIP and BLIP models.
Annotations on a Budget: Leveraging Geo-Data Similarity to Balance Model Performance and Annotation Cost (2024.lrec-main)

Copied to clipboard

Challenge: Current foundation models have shown impressive performance across various tasks, but they are not effective for everyone due to the imbalanced geographical and economic representation of the data used in the training process.
Approach: They propose to identify the data to be annotated to balance model performance and annotation costs by finding countries with visual similarity for the topics.
Outcome: The proposed methods improve model performance and reduce annotation costs by using data from countries with higher visual similarity for these topics.
Exploring the Value of Personalized Word Embeddings (2020.coling-main)

Copied to clipboard

Challenge: a subset of words belonging to specific psycholinguistic categories vary more in their representations across users . combining generic and personalized word embeddings yields the best performance .
Approach: They propose personalized word embeddings and compare their performance to generic ones . they show that personalized word representations can be leveraged for improved performance .
Outcome: The proposed model can be used for authorship attribution.
Factors Influencing the Surprising Instability of Word Embeddings (N18-1)

Copied to clipboard

Challenge: Word embeddings are low-dimensional, dense vector representations that capture semantic properties of words.
Approach: They examine the stability of word embeddings by examining their properties and analyzing their effects on downstream tasks.
Outcome: The results show that even high frequency words exhibit substantial instability, which can have implications for downstream tasks.
The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning (2025.naacl-long)

Copied to clipboard

Challenge: Large Multimodal Models exhibit impressive performance across multimodal tasks . effectiveness in cross-cultural contexts limited due to predominantly Western-centric nature of data and models . multi-agent models have shown significant capability in solving complex tasks despite limitations in crosscultural context .
Approach: They propose to use a multi-agent framework to enhance cross-cultural image captioning using LMMs with distinct cultural personas to evaluate cultural information within image captions.
Outcome: The proposed model outperforms single-agent models across different metrics and offers valuable insights for future research.
Scalable Performance Analysis for Vision-Language Models (2023.starsem-1)

Copied to clipboard

Challenge: a new method to probe vision-language models is proposed that does not require data annotation and makes use of existing datasets.
Approach: They propose a method that extracts features from a vision-language benchmark and measures their correlation with the output of the target model.
Outcome: The proposed method is scalable and does not require data annotation . it can be used with other models and benchmarks, and is available at https://github.com/MichiganNLP/Scalable-VLM-Probing.
MIME: MIMicking Emotions for Empathetic Response Generation (2020.emnlp-main)

Copied to clipboard

Challenge: Empathy is a fundamental human trait that reflects our ability to understand and reflect the thoughts and feelings of the people we interact with.
Approach: They propose to use polarity-based emotion clusters to generate empathetic responses . they also introduce stochasticity into the emotion mixture that yields emotionally more varied responses compared to the previous work .
Outcome: The proposed methods improve empathy and contextual relevance of the response, and introduce stochasticity into the emotion mixture that yields emotionally more varied responses than the previous work.
ICON: Interactive Conversational Memory Network for Multimodal Emotion Detection (D18-1)

Copied to clipboard

Challenge: Existing studies do not explicitly consider inter-personal influences that thrive in the emotional dynamics of dialogues.
Approach: They propose a multimodal emotion detection framework that extracts multimodal features from conversational videos and hierarchically models the self- and inter-speaker emotional influences into global memories.
Outcome: The proposed model outperforms state-of-the-art networks on multiple classification and regression tasks in two benchmark datasets.
Towards Multimodal Sarcasm Detection (An _Obviously_ Perfect Paper) (P19-1)

Copied to clipboard

Challenge: sarcasm is often expressed through multiple verbal and non-verbal cues, such as a change of tone, overemphasis, drawn-out syllables, or a straight looking face.
Approach: They propose to use multimodal cues to improve sarcasm detection using audiovisual utterances annotated with sarcasm labels to improve the accuracy.
Outcome: The proposed dataset reduces the error rate of sarcasm detection by 12.9% . it is based on audiovisual utterances annotated with sarcasm labels .
Mind the (Belief) Gap: Group Identity in the World of LLMs (2025.findings-acl)

Copied to clipboard

Challenge: Social biases and belief-driven behaviors can significantly impact Large Language Models’ (LLMs) decisions on several tasks.
Approach: They propose a multi-agent framework that simulates belief congruence, a group psychology theory that plays a crucial role in shaping societal interactions and preferences.
Outcome: The proposed framework reduces misinformation dissemination and improves learning by 11% while reducing misinformation dissemination by up to 37%.
Cross-Cultural Analysis of Human Values, Morals, and Biases in Folk Tales (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies on folk tales focus on European tales, ignoring large swaths of the world's diverse cultures.
Approach: They compile a corpus of over 1,900 folk tales originating from 27 diverse cultures across six continents and employ lexicon-based correlation analyses to examine human values, morals, and gender biases.
Outcome: The results show that folk tales are influenced by cultural norms and cultural values and are well-known for their morals and values.
Exploring Self-Identified Counseling Expertise in Online Support Forums (2021.findings-acl)

Copied to clipboard

Challenge: Increasing number of people engage in online health forums, making it important to understand the quality of the advice they receive.
Approach: They examine the role of expertise in responses to help-seeking posts . they find that a classifier can distinguish between peer and self-identified mental health professionals' interactions .
Outcome: The findings show that experts' language use differs between groups, and that their comments engage the support-seeker further.
“Judge me by my size (noun), do you?” YodaLib: A Demographic-Aware Humor Generation Framework (2020.coling-main)

Copied to clipboard

Challenge: Humor is subjective and can be interpreted in different ways by different people.
Approach: They propose an automatic method for filling the blanks in Mad Libs stories . they build upon the BERT platform to predict location-biased word fillings in incomplete sentences .
Outcome: The proposed framework outperforms a semi-automated approach for filling the blanks in Mad Libs stories while accounting for the demographic backgrounds of the desired audience.
How Good Is NLP? A Sober Look at NLP Tasks through the Lens of Social Impact (2021.findings-acl)

Copied to clipboard

Challenge: Recent years have seen many breakthroughs in natural language processing (NLP), transitioning it from a mostly theoretical field to one with many real-world applications.
Approach: They propose a moral philosophy definition of social good and a framework to evaluate the direct and indirect real-world impact of NLP tasks.
Outcome: The proposed framework evaluates the direct and indirect real-world impact of NLP tasks and adopts the methodology of global priorities research to identify priority causes for NLP research.
MoMentS: A Comprehensive Multimodal Benchmark for Theory of Mind (2025.findings-emnlp)

Copied to clipboard

Challenge: MoMentS is a benchmark designed to assess the ToM capabilities of multimodal large language models (LLMs) in short films.
Approach: They introduce a benchmark to assess the ToM capabilities of multimodal large language models (LLMs) through realistic, narrative-rich scenarios presented in short films.
Outcome: The proposed benchmark features long video context windows and realistic social interactions that provide deeper insight into characters’ mental states.
Hitting your MARQ: Multimodal ARgument Quality Assessment in Long Debate Video (2021.emnlp-main)

Copied to clipboard

Challenge: Current literature mostly considers textual content while assessing argument quality, and it is limited to datasets containing short text sequences (18-48 words).
Approach: They propose a set of interpretable debate centric features that are inspired by theories of argument quality and propose MARQ model that summarizes the multimodal signals on long debate videos.
Outcome: The proposed model outperforms baseline models with an error rate reduction of 22.7% on the argument quality prediction task and achieves 81.91% accuracy.
Task-Adaptive Tokenization: Enhancing Long-Form Text Generation Efficacy in Mental Health and Beyond (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods to generate text in mental health are limiting, but they are effective for many tasks.
Approach: They propose a task-adaptive tokenizer that allows for the integration of task-specific tokens into the pre-trained model's tokenization step.
Outcome: The proposed tokenization approach improves generation performance on psychological question-answering tasks in Chinese and English while using 60% fewer tokens.
Are Language Models Consequentialist or Deontological Moral Reasoners? (2025.emnlp-main)

Copied to clipboard

Challenge: Recent studies have focused on the moral judgments in large language models rather than their underlying moral reasoning process.
Approach: They propose a taxonomy of moral rationales to classify reasoning traces according to consequentialism and deontology . they use trolley problems to analyze moral reasoning tracing in large language models .
Outcome: The proposed taxonomy of moral rationales sheds light on consequentialism and deontology . it systematically classifies reasoning traces according to two main ethical theories .
Logical Fallacy Detection (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing language models perform poorly on logical fallacy detection . fallacious arguments can lead to disagreements, conflicts, endless debates, and a lack of consensus .
Approach: They propose a task of logical fallacy detection and propose LogicClimate to detect fallacies in text.
Outcome: The proposed task outperforms the best language model on Logic and LogicClimate . human reasoning is marred by logical fallacies, and some exacerbate misinformation .
Speaker Naming in Movies (N18-1)

Copied to clipboard

Challenge: Identifying speakers and their names in movies is a primary task for many video analysis problems, such as automatic subtitle labeling.
Approach: They propose a model that leverages visual, textual, and acoustic modalities in an unified optimization framework for speaker naming in movies.
Outcome: The proposed model outperforms baseline models on the MovieQA 2017 challenge for speaker naming in movies and TV shows on visual, textual, and acoustic modalities.
Improving Low Compute Language Modeling with In-Domain Embedding Initialisation (2020.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to train language models on in-domain data are limited.
Approach: They propose to initialise and freeze in-domain embeddings to provide a useful representation of rare words in English . they find that the standard configuration is not optimal when rare words are present .
Outcome: The proposed approach improves language modeling by providing a useful representation of rare words in English.
Inferring Social Media Users’ Mental Health Status from Multimodal Information (2020.lrec-1)

Copied to clipboard

Challenge: In the United States alone, one in every four adults suffers from a mental health condition, making mental health a pressing concern.
Approach: They propose to use multimodal cues present in social media posts to predict mental health status by analyzing language, visual, and metadata cue data.
Outcome: The proposed approach improves the performance of the classification task compared to using one modality at a time and can provide important cues into a user’s mental status.
Empathy Identification Systems are not Accurately Accounting for Context (2023.eacl-main)

Copied to clipboard

Challenge: Empathy is a fundamental phenomenon that allows us to better communicate and relate with others.
Approach: They propose a simple model that checks if an input utterance is similar to a small set of empathetic examples, but does not consider dialogue context.
Outcome: The proposed model outperforms state-of-the-art models on benchmarks and empathetic rationale extraction benchmarks.
Micromodels for Efficient, Explainable, and Reusable Systems: A Case Study on Mental Health (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing statistical models are not explainable, struggle in low-resource scenarios and cannot be reused for multiple tasks.
Approach: They propose a micromodel architecture that embeds domain knowledge and provides explanations throughout the model’s decision process.
Outcome: The proposed model is validated on depression classification, PTSD classification, and suicidal risk assessment tasks.
Towards Understanding the Relation between Gestures and Language (2022.coling-1)

Copied to clipboard

Challenge: a new study explores the relationship between gestures and language . we use contrastive learning to learn gesture embeddings .
Approach: They adapt a semi-supervised multimodal model to learn gesture embeddings using Ted talks . they show gestures are predictive of the native language of the speaker .
Outcome: The proposed model learns gesture embeddings from a multimodal dataset . it shows that gesture embeds are predictive of the native language of the speaker .
Unsupervised Discrete Representations of American Sign Language (2024.emnlp-main)

Copied to clipboard

Challenge: Modern NLP models use discrete tokens to represent continuous signals, such as videos, audio, or gestures . modalities that are continuous are difficult to use with discrete models, such a LLM .
Approach: They propose a method that discretizes sequences of fingerspelling signs into tokens . they also propose 'loss function' to improve interpretability of the tokens.
Outcome: The proposed method improves the performance of the tokenizer on downstream tasks.
Understanding the Capabilities and Limitations of Large Language Models for Cultural Commonsense (2024.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated substantial commonsense understanding through numerous benchmark evaluations.
Approach: They conduct a comprehensive examination of the capabilities and limitations of several state-of-the-art LLMs in the context of cultural commonsense tasks.
Outcome: The language used to query the LLMs can impact their performance on cultural-related tasks.
Eeyore: Realistic Depression Simulation via Expert-in-the-Loop Supervised and Preference Optimization (2025.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have been explored for mental healthcare training and therapy client simulation, but they fail to authentically capture diverse client traits and psychological conditions.
Approach: They propose an 8B model optimized for realistic depression simulation with expert input at every stage.
Outcome: The model outperforms GPT-4o in linguistic authenticity and profile adherence.
Chumor 2.0: Towards Better Benchmarking Chinese Humor Understanding from (Ruo Zhi Ba) (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies on humor in non-English languages lack culturally nuanced humor in other languages.
Approach: They construct a Chinese humor explanation dataset using a reddit-like platform . they test ten LLMs and find they are significantly better than existing LLM models .
Outcome: The proposed dataset is the first and largest Chinese humor explanation dataset.
The Generation Gap: Exploring Age Bias in the Value Systems of Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Using the World Value Survey, we find a general inclination of LLM values towards younger demographics, especially when compared to the US population.
Approach: They use data from the World Value Survey to examine the alignment of LLM values with specific age groups.
Outcome: The proposed model can be used to predict the value of a large language model and to assess its performance on 13 categories.
CICERO: A Dataset for Contextualized Commonsense Inference in Dialogues (2022.acl-long)

Copied to clipboard

Challenge: Fig. 1a shows an example where commonsense knowledge is crucial in sifting relevant information from the context.
Approach: They curate a dataset of dyadic conversations with five types of utterance-level reasoning-based inferences: cause, subsequent event, prerequisite, motivation, and emotional reaction.
Outcome: The dataset contains 53,105 of such inferences from 5,672 dialogues.
MUSER: MUltimodal Stress detection using Emotion Recognition as an Auxiliary Task (2021.naacl-main)

Copied to clipboard

Challenge: Existing methods to detect stress have not explored the inter-dependence between emotion and stress.
Approach: They propose a transformer-based model architecture and a novel multi-task learning algorithm with speed-based dynamic sampling strategy to improve stress detection.
Outcome: The proposed model is effective with internal and external auxiliary tasks and achieves state-of-the-art results.
CliniDial: A Naturally Occurring Multimodal Dialogue Dataset for Team Reflection in Action During Clinical Operation (2025.findings-acl)

Copied to clipboard

Challenge: Prior studies have shown that sufficient collaboration is the key factor that determines the outcome of an operation.
Approach: They propose to model the communication between team members during an operation using audio data and physiology signals from two camera angles.
Outcome: The proposed model is based on existing frameworks and invites future effort on developing methods that can deal with real-world clinical data.
VERVE: Template-based ReflectiVE Rewriting for MotiVational IntErviewing (2023.findings-emnlp)

Copied to clipboard

Challenge: During the Covid-19 pandemic, the number of people living with anxiety and depression rose more than four times . counselor training is difficult to speed up due to several factors, such as the need for expert supervision and the laborious and time-extensive process needed to provide evaluative feedback.
Approach: They propose a template-based rewriting system that transforms non-reflective statements into reflective responses using paraphrase-augmented training and adaptive template updating.
Outcome: The proposed model transforms non-reflective statements into more reflective responses while achieving a good content preservation-reflection style trade-off.
World Knowledge for Abstract Meaning Representation Parsing (L18-1)

Copied to clipboard

Challenge: Abstract Meaning Representation (AMR) parsers are based on annotated graphs, but there is still room for improvement .
Approach: They examine the role played by world knowledge in parsing errors in a state-of-the-art parser . they examine the effects of different types of world knowledge on parsers .
Outcome: The proposed model improves on multiple fine-grained metrics, including a 6% increase in named entity F-score, and provides insight into the potential of world knowledge for future work in Abstract Meaning Representation parsing.
Dynamic Reward Adjustment in Multi-Reward Reinforcement Learning for Counselor Reflection Generation (2024.lrec-main)

Copied to clipboard

Challenge: In this paper, we analyze the problem of multi-reward reinforcement learning to optimize for multiple text qualities for natural language generation.
Approach: They propose to use multi-reward reinforcement learning to optimize for multiple text qualities for natural language generation by using bandits.
Outcome: The proposed techniques outperform existing naive and bandit baselines, showcasing their potential for enhancing language models.
EmoBench: Evaluating the Emotional Intelligence of Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for Emotional Intelligence (EI) focus on emotion recognition, neglecting essential EI capabilities.
Approach: They propose a benchmark that proposes a comprehensive definition for machine EI . they propose 400 hand-crafted questions in English and Chinese to evaluate EI.
Outcome: The proposed benchmarks focus on emotion recognition, neglecting EI capabilities . they are constructed from existing datasets, which include frequent patterns and errors . the proposed benchmark includes questions in English and Chinese that require thorough reasoning and understanding .
STaCK: Sentence Ordering with Temporal Commonsense Knowledge (2021.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to sentence order prediction ignore the importance of document level global information, i.e., while predicting relative order of two sentences (s i , s j) other sentences sk from the same document do not play any role.
Approach: They propose a framework based on graph neural networks and temporal commonsense knowledge to model global information and predict relative order of sentences.
Outcome: The proposed method is naturally suitable for order prediction on five different datasets and has potential applications in the evaluation of the quality of machinegenerated documents.
SafetyALFRED: Evaluating Safety-Conscious Planning of Vision Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Existing safety evaluations focus on hazard recognition through disembodied question answering (QA) settings, but lack a critical gap in evaluating an agent.
Approach: They evaluate multimodal large language models with six categories of kitchen hazards . they propose a safety-based approach that prioritizes multi-step corrective actions .
Outcome: The proposed model can recognize hazards in QA settings, but average mitigation success rates are low . the proposed model is based on the embodied agent benchmark ALFRED .
Predicting Human Activities from User-Generated Content (P19-1)

Copied to clipboard

Challenge: Several studies have applied computational approaches to the understanding and modeling of human behavior at scale and in real time.
Approach: They propose a sentence embedding framework tailored to recognize the semantics of human activities and perform automatic clustering of these activities.
Outcome: The proposed framework can make predictions based on the text of user-generated content and self-description.
Tables as Texts or Images: Evaluating the Table Reasoning Ability of LLMs and MLLMs (2024.findings-acl)

Copied to clipboard

Challenge: Recent years have witnessed an explosion of Large Language Models (LLMs), with impressive performance on various NLP tasks.
Approach: They propose to use image-based representations to compare LLMs' performance on table-related tasks such as question-answering and fact-checking to determine their effectiveness.
Outcome: The proposed model performs better on image-based representations than on text-based models.
A Comparative Multidimensional Analysis of Empathetic Systems (2024.eacl-long)

Copied to clipboard

Challenge: Empathetic dialogue systems have received significant attention, but no systematic review has verified these limitations.
Approach: They analyze 21 empathetic dialogue systems using automated methods to examine their progress.
Outcome: The results show that empathetic dialogue systems lack specificity, reflection levels, diversity . the results also offer guidance for developing future systems .
When Do Language Models Endorse Limitations on Human Rights Principles? (2026.findings-eacl)

Copied to clipboard

Challenge: a recent study evaluated how large language models navigate trade-offs involving the Universal Declaration of Human Rights.
Approach: They evaluate how large language models navigate trade-offs involving the Universal Declaration of Human Rights (UDHR) they use 1,152 synthetically generated scenarios across 24 rights articles and eight languages .
Outcome: The proposed models accept limiting economic, social, and cultural rights more often than political and civil rights, the authors show . their models show significant cross-linguistic variation with elevated endorsement rates of rights-limiting actions in Chinese and Hindi compared to English or Romanian .
Evaluating LLMs’ Mathematical and Coding Competency through Ontology-guided Interventions (2025.findings-acl)

Copied to clipboard

Challenge: Current large language models have shown impressive performance on logical reasoning benchmarks . however, the true depth of their competencies and robustness in reasoning tasks remains an open question .
Approach: They propose a general ontology of perturbations and a semi-automatic method to apply perturbations to arithmetic reasoning and code generation datasets to test their LLMs' capabilities.
Outcome: The proposed model outperforms existing models on arithmetic reasoning and code generation tasks.
Towards Dog Bark Decoding: Leveraging Human Speech Processing for Automated Bark Classification (2024.lrec-main)

Copied to clipboard

Challenge: Similar to humans, animals make extensive use of verbal and non-verbal forms of communication, including audio signals.
Approach: They propose to use self-supervised speech representation models pre-trained on human speech to address dog bark classification tasks.
Outcome: The proposed model improves dog recognition, breed identification, gender classification, and context grounding tasks.
Uplifting Lower-Income Data: Strategies for Socioeconomic Perspective Shifts in Large Multi-modal Models (2025.naacl-long)

Copied to clipboard

Challenge: Unequal representation of cultures and socioeconomic groups in training data leads to biased Large Multi-modal (LMM) models.
Approach: They propose and evaluate several prompting strategies that use non-English, geographic, and socioeconomic attributes to improve LMM model performance on underrepresented data.
Outcome: The proposed prompts favor retrieving topic appearances from low-income data on lower-income datasets.
Mining the Cause of Political Decision-Making from Social Media: A Case Study of COVID-19 Policies across the US States (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on political responsiveness focus on long-term policies collected over decades . recent COVID-19 pandemic has given rise to a new political phenomenon, where political leaders make frequent short-term decisions on the same controlled topic.
Approach: They propose to use Twitter data to classify the sentiments toward governors of each state and conduct controlled studies and comparisons.
Outcome: The proposed model focuses on the COVID-19 pandemic, where political leaders make frequent short-term decisions on the same controlled topic.
Analyzing Occupational Distribution Representation in Japanese Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Recent advances in large language models have enabled users to generate fluent and seemingly convincing text, but they have uneven performance in different languages, which is associated with undesirable societal biases toward marginalized populations.
Approach: They develop three Japanese language prompts to probe LLMs’ understanding of Japanese names and their association between gender and occupations.
Outcome: The proposed models can associate Japanese names with correct gendered occupations when using constrained decoding, but with sampling or greedy decoding they prefer a small set of stereotypically genderes.
Beyond Good Intentions: Reporting the Research Landscape of NLP for Social Good (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in natural language processing (NLP) have created a vast number of applications that are aimed at social good applications.
Approach: They propose a dataset with three tasks that can help identify NLP4SG papers and characterize the NLP landscape by: (1) identifying the papers that address a social problem, (2) mapping them to the corresponding UN Sustainable Development Goals, and (3) identifying their methods.
Outcome: The proposed dataset can help identify NLP4SG papers and characterize the NLP landscape by: (1) identifying the papers that address a social problem, (2) mapping them to the corresponding UN Sustainable Development Goals (SDGs), and (3) identifying their methods.
In-the-Wild Video Question Answering (2022.coling-1)

Copied to clipboard

Challenge: Existing video understanding datasets focus on human interactions with little attention being paid to the “in the wild” settings.
Approach: They propose a video understanding dataset of videos recorded outdoors . they propose identifying visual support for a given question and answer .
Outcome: The proposed dataset examines the ability of models to understand videos, including video question answering, video captioning, and fill-inthe-blank tasks.
Implicit Personalization in Language Models: A Systematic Study (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing studies have focused on the implicit personalization problem, but no unified framework exists to study it.
Approach: They propose a mathematical formulation and a moral reasoning framework to study the phenomenon of Implicit Personalization (IP) they propose 'direct intervention' to estimate causal effect of mediator variable that cannot be directly intervened upon.
Outcome: The proposed method estimates the causal effect of a mediator variable that cannot be directly intervened upon.
Analyzing the Surprising Variability in Word Embedding Stability Across Languages (2021.emnlp-main)

Copied to clipboard

Challenge: Word embeddings are powerful representations that form the foundation of many natural language processing architectures.
Approach: They explore word embedding stability in a wide range of languages to gain insight into their stability.
Outcome: The proposed results provide insights into word embedding stability in English and other languages.
Examining Spanish Counseling with MIDAS: a Motivational Interviewing Dataset in Spanish (2025.naacl-short)

Copied to clipboard

Challenge: Cultural and language factors influence counseling, but research has not explored whether this applies to other languages.
Approach: They introduce a Spanish-language counseling dataset that contains expert annotations for counseling reflections and questions.
Outcome: The proposed dataset explores language-based differences in counselor behavior in English and Spanish and develops classifiers in monolingual and multilingual settings.
What Makes a Good Counselor? Learning to Distinguish between High-quality and Low-quality Counseling Conversations (P19-1)

Copied to clipboard

Challenge: Qualitative counseling relies on active collaboration between clients and counselors .
Approach: They propose to use linguistic features to capture differences between high- and low-quality counseling conversations to build automatic classifiers that can predict counseling quality with accuracies of up to 88%.
Outcome: The proposed model can predict counseling quality with accuracies of up to 88%.
Towards Implicit Bias Detection and Mitigation in Multi-Agent LLM Interactions (2024.findings-emnlp)

Copied to clipboard

Challenge: a recent study shows that large language models are susceptible to societal biases due to their exposure to human-generated data.
Approach: They propose two strategies to mitigate implicit gender biases in large language models . they create scenarios where implicit gender is present and develop a metric to assess the presence of biase .
Outcome: The proposed methods mitigate implicit biases with self-reflection and fine-tuning.
Copyright Detective: A Forensic System to Evidence LLMs Flickering Copyright Leakage Risks (2026.acl-demo)

Copied to clipboard

Challenge: **Copyright Detective** is the first interactive forensic system for detecting, analyzing, and visualizing potential copyright risks in LLM outputs.
Approach: They propose a system that detects copyright infringements and visualizes them . they use content recall testing, paraphrase-level similarity analysis and persuasive jailbreak probing .
Outcome: The proposed system detects, analyzes, and visualizes potential copyright risks in LLM outputs.
Modality-specific Learning Rates for Effective Multimodal Additive Late-fusion (2022.findings-acl)

Copied to clipboard

Challenge: Multimodal machine learning uses additive late-fusion to combine feature representations from different modalities into a joint representation.
Approach: They propose a Modality-Specific Learning Rate method to build late-fusion multimodal models from fine-tuned unimodal models.
Outcome: The proposed method outperforms global learning rates on multiple tasks and settings and enables the models to effectively learn each modality.
KinGDOM: Knowledge-Guided DOMain Adaptation for Sentiment Analysis (2020.acl-main)

Copied to clipboard

Challenge: Existing approaches to cross-domain sentiment analysis cannot be reliably deployed due to the distributional mismatch between training and evaluation domains.
Approach: They propose a framework that uses ConceptNet to enrich semantics of documents by providing domain-specific and domain-general background concepts.
Outcome: The proposed framework improves on a domain-adversarial baseline method and can be used in domain adaptation.
LifeQA: A Real-life Dataset for Video Question Answering (2020.lrec-1)

Copied to clipboard

Challenge: Existing video question answering datasets consist of movies and TV shows, but they are not representative of our day-to-day lives.
Approach: They propose a benchmark dataset for video question answering that focuses on day-to-day situations.
Outcome: The proposed dataset analyzes the challenging but realistic aspects of LifeQA . it consists of video clips and over 2.3k multiple-choice questions .
Using Paraphrases to Study Properties of Contextual Embeddings (2022.naacl-main)

Copied to clipboard

Challenge: Previously, paraphrases have been used to probe whether compositionality is accurately captured by BERT, but we believe they can be used to explore many other questions.
Approach: They propose to use paraphrases as a unique source of data to analyze contextualized embeddings, with a particular focus on BERT.
Outcome: The proposed analysis of paraphrases and paraphrase representations using the Paraphrase Database shows that BERT handles polysemous words, but different representations in many cases.
Two is Better than Many? Binary Classification as an Effective Approach to Multi-Choice Question Answering (2022.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to multi-choice question answering are based on binary classifications instead of scoring each answer as a single class.
Approach: They propose a simple refactoring of multi-choice question answering tasks as a series of binary classifications and propose re-framing to make them more efficient.
Outcome: The proposed approach is significantly more effective across different tasks and models.
Demographic-Aware Language Model Fine-tuning as a Bias Mitigation Technique (2022.aacl-short)

Copied to clipboard

Challenge: In this paper, we analyze the variations in gender and racial biases in BERT-like language models when exposed to different demographic groups.
Approach: They analyze gender and racial biases in BERT-like language models when exposed to different demographic groups.
Outcome: The proposed model can mitigate biases in text authored by disadvantaged demographic groups compared to advantaged groups . the proposed model is agnostic to the language of the speakers behind the language .
Revealing Hidden Mechanisms of Cross-Country Content Moderation with Natural Language Processing (2025.findings-acl)

Copied to clipboard

Challenge: Existing knowledge on how and why NLP methods make content moderation decisions is limited . authors examine how and when to use LLMs in content modeation .
Approach: They use Shapley values and LLM-guided explanations to reverse-engineer content moderation decisions across countries.
Outcome: The proposed methods show that they reverse-engineer content moderation decisions across countries and over time.
Leveraging Similar Users for Personalized Language Modeling with Limited Data (2022.acl-long)

Copied to clipboard

Challenge: Recent work suggests that personalized models are more accurate for individual users than one-size-fits-all solutions.
Approach: They propose a model trained on users that are similar to a new user to find similarity between new and existing users.
Outcome: The proposed model can predict what a user will write when they join a platform and not enough text is available.
Small Town or Metropolis? Analyzing the Relationship between Population Size and Language (2020.lrec-1)

Copied to clipboard

Challenge: Prior studies have examined how location affects the type of language that people use . recent electoral results in the united states exemplify a divide in the political opinions of those living in densely populated areas .
Approach: They analyze tweets from different Twitter users to determine whether they are from an urban or rural area.
Outcome: The proposed model trains predictive models to predict whether a user is from an urban or rural area.
Biased TextRank: Unsupervised Graph-Based Content Extraction (2020.coling-main)

Copied to clipboard

Challenge: TextRank ranks text spans according to their importance for language processing tasks and their relevance to an input “focus.”
Approach: They propose a graph-based content extraction method inspired by TextRank that ranks text spans according to their importance for language processing tasks and according to relevance to an input “focus.”
Outcome: The proposed method improves on two different datasets by significant ROUGE-N score margins.
Analyzing the Quality of Counseling Conversations: the Tell-Tale Signs of High-quality Counseling (L18-1)

Copied to clipboard

Challenge: Behavioral and mental health disorders are the most costly and prevalent conditions worldwide.
Approach: They propose to use a dataset to analyze counseling interactions by using aspects such as mirroring, empathy, and reflective listening to build text-based classifiers.
Outcome: The proposed dataset can be used to build text-based classifiers able to predict the overall quality of a counseling conversation and provide insights into the linguistic differences between low-quality and high-quality counseling.
Has It All Been Solved? Open NLP Research Questions Not Solved by Large Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Recent advances in large language models have led to misleading public discourse that “it’s all been solved.”
Approach: They identify 14 research areas encompassing 45 research directions that require new research and are not directly solvable by LLMs.
Outcome: The research areas identified are 45 research directions that require new research and are not directly solvable by LLMs.
CASCADE: Contextual Sarcasm Detection in Online Discussion Forums (C18-1)

Copied to clipboard

Challenge: Existing studies on sarcasm detection focus on lexical, syntactic and semantic cues, but sarcasm can be expressed implicitly without such cue.
Approach: They propose a ContextuAl SarCasm DEtector which extracts contextual information from the discourse of a discussion thread.
Outcome: The proposed model improves on a large Reddit corpus.
Rethinking Table Instruction Tuning (2025.findings-acl)

Copied to clipboard

Challenge: Existing studies have overlooked the impact of hyperparameters on table understanding abilities . authors show that smaller learning rates and fewer training instances can enhance table understanding while preserving general capabilities.
Approach: They propose a hyperparameter-based instruction-tuned model for table-related tasks that improves out-of-domain table understanding ability and general capabilities.
Outcome: The proposed model outperforms existing models on table-related tasks while maintaining strong out-of-domain generalization and general capabilities.
MAiDE-up: Multilingual Deception Detection of AI-generated Hotel Reviews (2025.findings-naacl)

Copied to clipboard

Challenge: Deceptive reviews are becoming more common, especially given the increase in performance and the prevalence of LLMs.
Approach: They compile and make publicly available a dataset of 10,000 real and 10,000 AI-generated fake hotel reviews in ten languages.
Outcome: The proposed model can detect real reviews and fake reviews in 10 languages.
You Are What You Annotate: Towards Better Models through Annotator Representations (2023.findings-emnlp)

Copied to clipboard

Challenge: Annotator disagreement is ubiquitous in natural language processing tasks.
Approach: They propose to model annotators' idiosyncrasies and account for their idioms by creating representations for each annotator and their annotations.
Outcome: The proposed model improves on an existing dataset with eight annotators with inherent disagreements while increasing model size by 1%.
Automatic Detection of Fake News (C18-1)

Copied to clipboard

Challenge: a growing number of fake news detection tools are needed to identify trustworthy news sources.
Approach: They propose to use two novel datasets to automate the identification of fake news . they propose learning experiments to build accurate fake news detectors .
Outcome: The proposed algorithms achieve accuracies of up to 76% and compare them with other tools . the proposed algorithms are based on satirical news sources and fact-checking websites .
Do LLMs Think Fast and Slow? A Causal Study on Sentiment Analysis (2024.findings-emnlp)

Copied to clipboard

Challenge: Sentiment analysis aims to identify the sentiment expressed in a piece of text, often in the form of a review.
Approach: They propose a causal discovery task that distinguishes whether a review "primes" the sentiment and a traditional prediction task to model the sentiment using the review as input.
Outcome: The proposed model improves by 32.13 F1 points on a zero-shot five-class SA.
Hi-ToM: A Benchmark for Evaluating Higher-Order Theory of Mind Reasoning in Large Language Models (2023.findings-emnlp)

Copied to clipboard

Challenge: Theory of Mind (ToM) is the ability to reason about one's own and others' mental states.
Approach: They propose a higher-order theory of mind benchmark and introduce a new deception mechanism to evaluate ToM reasoning.
Outcome: The proposed benchmarks show that the LLMs are not performing well on higher-order tasks.
Box of Lies: Multimodal Deception Detection in Dialogues (N19-1)

Copied to clipboard

Challenge: Deception occurs during everyday conversations, but this setting has received little attention from the research community.
Approach: They propose to analyze multimodal deceptive dialogues in a box of lies game . they use facial and linguistic annotations to identify deceptives and truthful behaviors .
Outcome: The proposed model outperforms both a random and a human baseline and achieves up to 69% accuracy in distinguishing deceptive and truthful behaviors.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations