Challenge: Claude Sonnet 3.5 consistently outperforms all models, but smaller models often approach human-rated empathy levels.
Approach: They introduce a dataset comprising 10,000 two-turn conversations across 500 diverse, clinically-grounded PTSD personas.
Outcome: The proposed model outperforms all models but has a "knowledge transfer ceiling" older adults prefer validation responses while graduate-educated users prefer emotionally layered responses .

Similar Papers

Harnessing the Power of Large Language Models for Empathetic Response Generation: Empirical Investigations and Improvements (2023.findings-emnlp)

Copied to clipboard

Challenge: Empathetic dialogue is an essential part of building harmonious social relationships and contributes to the development of a helpful AI.
Approach: They propose three methods to improve the performance of large language models (LLMs) they propose semantically similar in-context learning, two-stage interactive generation and combination with the knowledge base.
Outcome: The proposed methods achieve state-of-the-art in automatic and human evaluations and the possibility of GPT-4 simulating human evaluators.
Towards Empathetic Open-domain Conversation Models: A New Benchmark and Dataset (P19-1)

Copied to clipboard

Challenge: EmpatheticDialogues dataset provides a benchmark for empathetic dialogue generation . human evaluators perceive dialogue models as more epathetic .
Approach: They propose a benchmark for empathetic dialogue generation from a dataset of 25k conversations grounded in emotional situations.
Outcome: The proposed benchmarks show that existing models are perceived to be more empathetic by human evaluators compared to models trained on large-scale Internet conversations.
EMPATH: An Ensemble Method for Automatic Fine-Grained Turn-Level Dialogue Empathy Evaluation with a Novel Emotional Distance Metric (2026.findings-acl)

Copied to clipboard

Challenge: Empathy evaluation metrics are lacking in the competitions, and classical dialogue evaluation metrics require further investigation.
Approach: They propose a framework which combines fine-tuned models, large language models, classical dialogue evaluation metrics, and a novel metric.
Outcome: The proposed framework improves on the WASSA 2024 benchmark and shows a statistically significant 8% improvement on the EX dataset.
A Computational Approach to Understanding Empathy Expressed in Text-Based Mental Health Support (2020.emnlp-main)

Copied to clipboard

Challenge: Empathy measurement has predominantly occurred in synchronous, face-to-face settings, and may not translate to asynchronous, text-based contexts.
Approach: They propose a computational approach to understanding how empathy is expressed in online mental health platforms.
Outcome: The proposed model can identify empathic conversations and extract rationales from them.
SoulChat: Improving LLMs’ Empathy, Listening, and Comfort Abilities through Fine-tuning with Multi-turn Empathy Conversations (2023.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) are used in psychological counseling to provide universal advice.
Approach: They constructed a multi-turn empathetic conversation dataset with 2 million samples . they found that the model's empathy ability is enhanced when finetuning .
Outcome: Experiments show that large language models can be finetuned to provide empathy . but, when applied to mental health or emotional support conversation, there are three main issues .
Tailored Emotional LLM-Supporter: Enhancing Cultural Sensitivity (2026.eacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have shown growing potential in offering emotional support, but their ability to deliver culturally sensitive support remains underexplored due to a lack of resources.
Approach: They propose a large language model dataset that includes 1,729 distress messages, 1,523 cultural signals and 1,041 support strategies with fine-grained emotional and cultural annotations.
Outcome: The proposed models outperform peer-reviewed models and lack cultural sensitivity.
Can AI Relate: Testing Large Language Model Response for Mental Health Support (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) are already being piloted for clinical use in hospitals . recent failures of the Tessa chatbot have led to doubts about their reliability in high-stakes settings.
Approach: They propose safety guidelines for the potential deployment of large language models for mental health response.
Outcome: The proposed framework measures equity in empathy and adherence of LLM responses to motivational interviewing theory.
A Taxonomy of Empathetic Response Intents in Human Social Conversations (2020.coling-main)

Copied to clipboard

Challenge: Open-domain conversational agents or chatbots are becoming increasingly popular in the natural language processing community.
Approach: They aim to combine dialogue act/intent modelling and neural response generation to produce a large-scale taxonomy for empathetic response intents.
Outcome: The proposed method improves the response quality of chatbots and makes them more controllable and interpretable.
Can Machines Resonate with Humans? Evaluating the Emotional and Empathic Comprehension of LMs (2024.findings-emnlp)

Copied to clipboard

Challenge: Empathy plays a pivotal role in fostering prosocial behavior, often triggered by the sharing of personal experiences through narratives.
Approach: They propose to use contrastive learning with masked LMs and supervised fine-tuning with large language models to improve empathy understanding in NLP models.
Outcome: The proposed methods show that there is low agreement among annotators and that cultural differences are a factor in their interpretation of empathy.
I Don’t Need Solution. I Need Emotional Support : Empathetic LLMs based on Emotional Validation (2026.findings-acl)

Copied to clipboard

Challenge: Existing large language models (LLMs) struggle to generate emotional support response, despite observing and reflecting on the help-seeker’s situation . Empathy drives the formation of constructive interpersonal and supportive relationships, including counseling for mental health care .
Approach: They propose to use a two-stage training process to enhance empathetic response generation through empathy acquisition and emotional validation alignment.
Outcome: The proposed method significantly improves empathetic response generation, achieving superior performance in both automatic and human evaluations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations