Multi-dimensional Evaluation of Empathetic Dialogue Responses (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Prior efforts to measure conversational empathy focus on expressed communicative intents, but ignore the fact that conversation is also a collaboration involving both speakers and listeners. |
| Approach: | They propose a multi-dimensional empathy evaluation framework to measure both expressed intents from the speaker’s perspective and perceived empathy from the listener’s viewpoint. |
| Outcome: | The proposed framework measures both expressed intents from the speaker’s perspective and perceived empathy from the listener’s viewpoint. |
Similar Papers
A Comparative Multidimensional Analysis of Empathetic Systems (2024.eacl-long)
Copied to clipboard
| Challenge: | Empathetic dialogue systems have received significant attention, but no systematic review has verified these limitations. |
| Approach: | They analyze 21 empathetic dialogue systems using automated methods to examine their progress. |
| Outcome: | The results show that empathetic dialogue systems lack specificity, reflection levels, diversity . the results also offer guidance for developing future systems . |
Towards Empathetic Open-domain Conversation Models: A New Benchmark and Dataset (P19-1)
Copied to clipboard
| Challenge: | EmpatheticDialogues dataset provides a benchmark for empathetic dialogue generation . human evaluators perceive dialogue models as more epathetic . |
| Approach: | They propose a benchmark for empathetic dialogue generation from a dataset of 25k conversations grounded in emotional situations. |
| Outcome: | The proposed benchmarks show that existing models are perceived to be more empathetic by human evaluators compared to models trained on large-scale Internet conversations. |
EMPATH: An Ensemble Method for Automatic Fine-Grained Turn-Level Dialogue Empathy Evaluation with a Novel Emotional Distance Metric (2026.findings-acl)
Copied to clipboard
| Challenge: | Empathy evaluation metrics are lacking in the competitions, and classical dialogue evaluation metrics require further investigation. |
| Approach: | They propose a framework which combines fine-tuned models, large language models, classical dialogue evaluation metrics, and a novel metric. |
| Outcome: | The proposed framework improves on the WASSA 2024 benchmark and shows a statistically significant 8% improvement on the EX dataset. |
Empathy Identification Systems are not Accurately Accounting for Context (2023.eacl-main)
Copied to clipboard
| Challenge: | Empathy is a fundamental phenomenon that allows us to better communicate and relate with others. |
| Approach: | They propose a simple model that checks if an input utterance is similar to a small set of empathetic examples, but does not consider dialogue context. |
| Outcome: | The proposed model outperforms state-of-the-art models on benchmarks and empathetic rationale extraction benchmarks. |
Harnessing the Power of Large Language Models for Empathetic Response Generation: Empirical Investigations and Improvements (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Empathetic dialogue is an essential part of building harmonious social relationships and contributes to the development of a helpful AI. |
| Approach: | They propose three methods to improve the performance of large language models (LLMs) they propose semantically similar in-context learning, two-stage interactive generation and combination with the knowledge base. |
| Outcome: | The proposed methods achieve state-of-the-art in automatic and human evaluations and the possibility of GPT-4 simulating human evaluators. |
A Taxonomy of Empathetic Response Intents in Human Social Conversations (2020.coling-main)
Copied to clipboard
| Challenge: | Open-domain conversational agents or chatbots are becoming increasingly popular in the natural language processing community. |
| Approach: | They aim to combine dialogue act/intent modelling and neural response generation to produce a large-scale taxonomy for empathetic response intents. |
| Outcome: | The proposed method improves the response quality of chatbots and makes them more controllable and interpretable. |
Modeling Empathetic Alignment in Conversation (2024.naacl-long)
Copied to clipboard
| Challenge: | Empathy requires perspective-taking and is not explicitly modelled in NLP . |
| Approach: | They propose a new approach to recognizing alignment in empathetic speech, grounded in Appraisal Theory, and use reddit to study emotional conversations to examine alignment. |
| Outcome: | The proposed approach can recognize appraisals and alignments in empathetic speech, and mental health professionals engage with substantially more emotional alignment. |
A Computational Approach to Understanding Empathy Expressed in Text-Based Mental Health Support (2020.emnlp-main)
Copied to clipboard
| Challenge: | Empathy measurement has predominantly occurred in synchronous, face-to-face settings, and may not translate to asynchronous, text-based contexts. |
| Approach: | They propose a computational approach to understanding how empathy is expressed in online mental health platforms. |
| Outcome: | The proposed model can identify empathic conversations and extract rationales from them. |
The Pursuit of Empathy: Evaluating Small Language Models for PTSD Dialogue Support (2025.emnlp-main)
Copied to clipboard
Suhas Bn, Yash Mahajan, Dominik O. Mattioli, Andrew M. Sherrill, Rosa I. Arriaga, Christopher Wiese, Saeed Abdullah
| Challenge: | Claude Sonnet 3.5 consistently outperforms all models, but smaller models often approach human-rated empathy levels. |
| Approach: | They introduce a dataset comprising 10,000 two-turn conversations across 500 diverse, clinically-grounded PTSD personas. |
| Outcome: | The proposed model outperforms all models but has a "knowledge transfer ceiling" older adults prefer validation responses while graduate-educated users prefer emotionally layered responses . |
Can Machines Resonate with Humans? Evaluating the Emotional and Empathic Comprehension of LMs (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Empathy plays a pivotal role in fostering prosocial behavior, often triggered by the sharing of personal experiences through narratives. |
| Approach: | They propose to use contrastive learning with masked LMs and supervised fine-tuning with large language models to improve empathy understanding in NLP models. |
| Outcome: | The proposed methods show that there is low agreement among annotators and that cultural differences are a factor in their interpretation of empathy. |