Challenge: Prior efforts to measure conversational empathy focus on expressed communicative intents, but ignore the fact that conversation is also a collaboration involving both speakers and listeners.
Approach: They propose a multi-dimensional empathy evaluation framework to measure both expressed intents from the speaker’s perspective and perceived empathy from the listener’s viewpoint.
Outcome: The proposed framework measures both expressed intents from the speaker’s perspective and perceived empathy from the listener’s viewpoint.

Similar Papers

A Comparative Multidimensional Analysis of Empathetic Systems (2024.eacl-long)

Copied to clipboard

Challenge: Empathetic dialogue systems have received significant attention, but no systematic review has verified these limitations.
Approach: They analyze 21 empathetic dialogue systems using automated methods to examine their progress.
Outcome: The results show that empathetic dialogue systems lack specificity, reflection levels, diversity . the results also offer guidance for developing future systems .
Towards Empathetic Open-domain Conversation Models: A New Benchmark and Dataset (P19-1)

Copied to clipboard

Challenge: EmpatheticDialogues dataset provides a benchmark for empathetic dialogue generation . human evaluators perceive dialogue models as more epathetic .
Approach: They propose a benchmark for empathetic dialogue generation from a dataset of 25k conversations grounded in emotional situations.
Outcome: The proposed benchmarks show that existing models are perceived to be more empathetic by human evaluators compared to models trained on large-scale Internet conversations.
EMPATH: An Ensemble Method for Automatic Fine-Grained Turn-Level Dialogue Empathy Evaluation with a Novel Emotional Distance Metric (2026.findings-acl)

Copied to clipboard

Challenge: Empathy evaluation metrics are lacking in the competitions, and classical dialogue evaluation metrics require further investigation.
Approach: They propose a framework which combines fine-tuned models, large language models, classical dialogue evaluation metrics, and a novel metric.
Outcome: The proposed framework improves on the WASSA 2024 benchmark and shows a statistically significant 8% improvement on the EX dataset.
Empathy Identification Systems are not Accurately Accounting for Context (2023.eacl-main)

Copied to clipboard

Challenge: Empathy is a fundamental phenomenon that allows us to better communicate and relate with others.
Approach: They propose a simple model that checks if an input utterance is similar to a small set of empathetic examples, but does not consider dialogue context.
Outcome: The proposed model outperforms state-of-the-art models on benchmarks and empathetic rationale extraction benchmarks.
Harnessing the Power of Large Language Models for Empathetic Response Generation: Empirical Investigations and Improvements (2023.findings-emnlp)

Copied to clipboard

Challenge: Empathetic dialogue is an essential part of building harmonious social relationships and contributes to the development of a helpful AI.
Approach: They propose three methods to improve the performance of large language models (LLMs) they propose semantically similar in-context learning, two-stage interactive generation and combination with the knowledge base.
Outcome: The proposed methods achieve state-of-the-art in automatic and human evaluations and the possibility of GPT-4 simulating human evaluators.
A Taxonomy of Empathetic Response Intents in Human Social Conversations (2020.coling-main)

Copied to clipboard

Challenge: Open-domain conversational agents or chatbots are becoming increasingly popular in the natural language processing community.
Approach: They aim to combine dialogue act/intent modelling and neural response generation to produce a large-scale taxonomy for empathetic response intents.
Outcome: The proposed method improves the response quality of chatbots and makes them more controllable and interpretable.
Modeling Empathetic Alignment in Conversation (2024.naacl-long)

Copied to clipboard

Challenge: Empathy requires perspective-taking and is not explicitly modelled in NLP .
Approach: They propose a new approach to recognizing alignment in empathetic speech, grounded in Appraisal Theory, and use reddit to study emotional conversations to examine alignment.
Outcome: The proposed approach can recognize appraisals and alignments in empathetic speech, and mental health professionals engage with substantially more emotional alignment.
A Computational Approach to Understanding Empathy Expressed in Text-Based Mental Health Support (2020.emnlp-main)

Copied to clipboard

Challenge: Empathy measurement has predominantly occurred in synchronous, face-to-face settings, and may not translate to asynchronous, text-based contexts.
Approach: They propose a computational approach to understanding how empathy is expressed in online mental health platforms.
Outcome: The proposed model can identify empathic conversations and extract rationales from them.
The Pursuit of Empathy: Evaluating Small Language Models for PTSD Dialogue Support (2025.emnlp-main)

Copied to clipboard

Challenge: Claude Sonnet 3.5 consistently outperforms all models, but smaller models often approach human-rated empathy levels.
Approach: They introduce a dataset comprising 10,000 two-turn conversations across 500 diverse, clinically-grounded PTSD personas.
Outcome: The proposed model outperforms all models but has a "knowledge transfer ceiling" older adults prefer validation responses while graduate-educated users prefer emotionally layered responses .
Can Machines Resonate with Humans? Evaluating the Emotional and Empathic Comprehension of LMs (2024.findings-emnlp)

Copied to clipboard

Challenge: Empathy plays a pivotal role in fostering prosocial behavior, often triggered by the sharing of personal experiences through narratives.
Approach: They propose to use contrastive learning with masked LMs and supervised fine-tuning with large language models to improve empathy understanding in NLP models.
Outcome: The proposed methods show that there is low agreement among annotators and that cultural differences are a factor in their interpretation of empathy.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations