Challenge: Empathetic dialogue systems have received significant attention, but no systematic review has verified these limitations.
Approach: They analyze 21 empathetic dialogue systems using automated methods to examine their progress.
Outcome: The results show that empathetic dialogue systems lack specificity, reflection levels, diversity . the results also offer guidance for developing future systems .

Similar Papers

Empathy Identification Systems are not Accurately Accounting for Context (2023.eacl-main)

Copied to clipboard

Challenge: Empathy is a fundamental phenomenon that allows us to better communicate and relate with others.
Approach: They propose a simple model that checks if an input utterance is similar to a small set of empathetic examples, but does not consider dialogue context.
Outcome: The proposed model outperforms state-of-the-art models on benchmarks and empathetic rationale extraction benchmarks.
Multi-dimensional Evaluation of Empathetic Dialogue Responses (2024.findings-emnlp)

Copied to clipboard

Challenge: Prior efforts to measure conversational empathy focus on expressed communicative intents, but ignore the fact that conversation is also a collaboration involving both speakers and listeners.
Approach: They propose a multi-dimensional empathy evaluation framework to measure both expressed intents from the speaker’s perspective and perceived empathy from the listener’s viewpoint.
Outcome: The proposed framework measures both expressed intents from the speaker’s perspective and perceived empathy from the listener’s viewpoint.
From Traits to Empathy: Personality-Aware Multimodal Empathetic Response Generation (2025.coling-main)

Copied to clipboard

Challenge: Existing approaches focus on acquiring affective and cognitive knowledge from text, but neglect the unique personality traits of individuals and the inherently multimodal nature of human face-to-face conversation.
Approach: They propose a multimodal dialogue system that generates empathetic responses from a perspective that considers the personality traits of users.
Outcome: The proposed system generates empathetic responses from a multimodal perspective and analyzes multimodal data to understand the user’s emotional state and situation.
A Critical Reflection and Forward Perspective on Empathy and Natural Language Processing (2022.findings-emnlp)

Copied to clipboard

Challenge: Empathy recognition and empathetic response generation tasks are well-established research directions, but there is little clarity on what empathy is and how it is being operationalized.
Approach: They argue that current directions will benefit from a clear conceptualization that includes operationalizing cognitive empathy components.
Outcome: The proposed framework will help to define and operationalize empathy in natural language processing.
Multi-Party Empathetic Dialogue Generation: A New Task for Dialog Systems (2022.acl-long)

Copied to clipboard

Challenge: Existing work on empathetic dialogues focused on the two-party scenario, but multi-party dialogues are pervasive in reality.
Approach: They propose a multi-party empathetic dialogue generation task that uses a static-dynamic model to explore emotion and sensibility.
Outcome: The proposed task is based on a model with static sensibility and dynamic emotion . it achieves state-of-the-art performance in multi-party empathetic dialogue learning .
Towards Empathetic Open-domain Conversation Models: A New Benchmark and Dataset (P19-1)

Copied to clipboard

Challenge: EmpatheticDialogues dataset provides a benchmark for empathetic dialogue generation . human evaluators perceive dialogue models as more epathetic .
Approach: They propose a benchmark for empathetic dialogue generation from a dataset of 25k conversations grounded in emotional situations.
Outcome: The proposed benchmarks show that existing models are perceived to be more empathetic by human evaluators compared to models trained on large-scale Internet conversations.
DialSummEval: Revisiting Summarization Evaluation for Dialogues (2022.naacl-main)

Copied to clipboard

Challenge: Current models for dialogue summarization have flaws that may not be well exposed by frequently used metrics such as ROUGE.
Approach: They propose to re-evaluate 18 categories of metrics in terms of four dimensions: coherence, consistency, fluency and relevance, as well as a unified human evaluation of various models for the first time.
Outcome: The proposed dataset will be used to evaluate 18 categories of metrics in terms of coherence, consistency, fluency and relevance, and a unified human evaluation of various models for the first time.
EMPATH: An Ensemble Method for Automatic Fine-Grained Turn-Level Dialogue Empathy Evaluation with a Novel Emotional Distance Metric (2026.findings-acl)

Copied to clipboard

Challenge: Empathy evaluation metrics are lacking in the competitions, and classical dialogue evaluation metrics require further investigation.
Approach: They propose a framework which combines fine-tuned models, large language models, classical dialogue evaluation metrics, and a novel metric.
Outcome: The proposed framework improves on the WASSA 2024 benchmark and shows a statistically significant 8% improvement on the EX dataset.
Medical Dialogue System: A Survey of Categories, Methods, Evaluation and Challenges (2024.findings-acl)

Copied to clipboard

Challenge: Existing medical dialogue systems have significant potential to simplify diagnostic procedure and reduce the cost of collecting information from patients.
Approach: They analyze 325 papers from well-known computer science, natural language processing conferences and journals to find out the major challenges of medical dialog systems.
Outcome: The proposed systems have been surveyed in the medical community but have not been evaluated from a technical perspective.
EmpDG: Multi-resolution Interactive Empathetic Dialogue Generation (2020.coling-main)

Copied to clipboard

Challenge: Existing work on empathetic dialogue generation fails to capture the nuances of human emotion and consider the potential of user feedback.
Approach: They propose a multi-resolution adversarial model - EmpDG - to generate more empathetic responses by exploiting both coarse-grained dialogue-level and fine-grounded token-level emotions.
Outcome: The proposed model outperforms the state-of-the-art models in both content quality and emotion perceptivity.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations