The Pursuit of Empathy: Evaluating Small Language Models for PTSD Dialogue Support (2025.emnlp-main)
Copied to clipboard
Suhas Bn, Yash Mahajan, Dominik O. Mattioli, Andrew M. Sherrill, Rosa I. Arriaga, Christopher Wiese, Saeed Abdullah
| Challenge: | Claude Sonnet 3.5 consistently outperforms all models, but smaller models often approach human-rated empathy levels. |
| Approach: | They introduce a dataset comprising 10,000 two-turn conversations across 500 diverse, clinically-grounded PTSD personas. |
| Outcome: | The proposed model outperforms all models but has a "knowledge transfer ceiling" older adults prefer validation responses while graduate-educated users prefer emotionally layered responses . |
Similar Papers
Harnessing the Power of Large Language Models for Empathetic Response Generation: Empirical Investigations and Improvements (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Empathetic dialogue is an essential part of building harmonious social relationships and contributes to the development of a helpful AI. |
| Approach: | They propose three methods to improve the performance of large language models (LLMs) they propose semantically similar in-context learning, two-stage interactive generation and combination with the knowledge base. |
| Outcome: | The proposed methods achieve state-of-the-art in automatic and human evaluations and the possibility of GPT-4 simulating human evaluators. |
Towards Empathetic Open-domain Conversation Models: A New Benchmark and Dataset (P19-1)
Copied to clipboard
| Challenge: | EmpatheticDialogues dataset provides a benchmark for empathetic dialogue generation . human evaluators perceive dialogue models as more epathetic . |
| Approach: | They propose a benchmark for empathetic dialogue generation from a dataset of 25k conversations grounded in emotional situations. |
| Outcome: | The proposed benchmarks show that existing models are perceived to be more empathetic by human evaluators compared to models trained on large-scale Internet conversations. |
EMPATH: An Ensemble Method for Automatic Fine-Grained Turn-Level Dialogue Empathy Evaluation with a Novel Emotional Distance Metric (2026.findings-acl)
Copied to clipboard
| Challenge: | Empathy evaluation metrics are lacking in the competitions, and classical dialogue evaluation metrics require further investigation. |
| Approach: | They propose a framework which combines fine-tuned models, large language models, classical dialogue evaluation metrics, and a novel metric. |
| Outcome: | The proposed framework improves on the WASSA 2024 benchmark and shows a statistically significant 8% improvement on the EX dataset. |
A Computational Approach to Understanding Empathy Expressed in Text-Based Mental Health Support (2020.emnlp-main)
Copied to clipboard
| Challenge: | Empathy measurement has predominantly occurred in synchronous, face-to-face settings, and may not translate to asynchronous, text-based contexts. |
| Approach: | They propose a computational approach to understanding how empathy is expressed in online mental health platforms. |
| Outcome: | The proposed model can identify empathic conversations and extract rationales from them. |
SoulChat: Improving LLMs’ Empathy, Listening, and Comfort Abilities through Fine-tuning with Multi-turn Empathy Conversations (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models (LLMs) are used in psychological counseling to provide universal advice. |
| Approach: | They constructed a multi-turn empathetic conversation dataset with 2 million samples . they found that the model's empathy ability is enhanced when finetuning . |
| Outcome: | Experiments show that large language models can be finetuned to provide empathy . but, when applied to mental health or emotional support conversation, there are three main issues . |
Tailored Emotional LLM-Supporter: Enhancing Cultural Sensitivity (2026.eacl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) have shown growing potential in offering emotional support, but their ability to deliver culturally sensitive support remains underexplored due to a lack of resources. |
| Approach: | They propose a large language model dataset that includes 1,729 distress messages, 1,523 cultural signals and 1,041 support strategies with fine-grained emotional and cultural annotations. |
| Outcome: | The proposed models outperform peer-reviewed models and lack cultural sensitivity. |
Can AI Relate: Testing Large Language Model Response for Mental Health Support (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models (LLMs) are already being piloted for clinical use in hospitals . recent failures of the Tessa chatbot have led to doubts about their reliability in high-stakes settings. |
| Approach: | They propose safety guidelines for the potential deployment of large language models for mental health response. |
| Outcome: | The proposed framework measures equity in empathy and adherence of LLM responses to motivational interviewing theory. |
A Taxonomy of Empathetic Response Intents in Human Social Conversations (2020.coling-main)
Copied to clipboard
| Challenge: | Open-domain conversational agents or chatbots are becoming increasingly popular in the natural language processing community. |
| Approach: | They aim to combine dialogue act/intent modelling and neural response generation to produce a large-scale taxonomy for empathetic response intents. |
| Outcome: | The proposed method improves the response quality of chatbots and makes them more controllable and interpretable. |
Can Machines Resonate with Humans? Evaluating the Emotional and Empathic Comprehension of LMs (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Empathy plays a pivotal role in fostering prosocial behavior, often triggered by the sharing of personal experiences through narratives. |
| Approach: | They propose to use contrastive learning with masked LMs and supervised fine-tuning with large language models to improve empathy understanding in NLP models. |
| Outcome: | The proposed methods show that there is low agreement among annotators and that cultural differences are a factor in their interpretation of empathy. |
I Don’t Need Solution. I Need Emotional Support : Empathetic LLMs based on Emotional Validation (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing large language models (LLMs) struggle to generate emotional support response, despite observing and reflecting on the help-seeker’s situation . Empathy drives the formation of constructive interpersonal and supportive relationships, including counseling for mental health care . |
| Approach: | They propose to use a two-stage training process to enhance empathetic response generation through empathy acquisition and emotional validation alignment. |
| Outcome: | The proposed method significantly improves empathetic response generation, achieving superior performance in both automatic and human evaluations. |