Challenge: Existing evaluation frameworks focus on diagnostic accuracy and win-rates and often overlook alignment with patient-specific goals, values, and personalities required for meaningful conversations.
Approach: They propose a framework for synthetically generating realistic, multi-turn mental health sensemaking conversations and a dataset to examine their models in healthcare settings.
Outcome: The proposed framework synthesizes a dataset comprising over 2,200 patient–LLM conversations and evaluates them using human-centric criteria.

Similar Papers

SoulChat: Improving LLMs’ Empathy, Listening, and Comfort Abilities through Fine-tuning with Multi-turn Empathy Conversations (2023.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) are used in psychological counseling to provide universal advice.
Approach: They constructed a multi-turn empathetic conversation dataset with 2 million samples . they found that the model's empathy ability is enhanced when finetuning .
Outcome: Experiments show that large language models can be finetuned to provide empathy . but, when applied to mental health or emotional support conversation, there are three main issues .
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation (2026.eacl-long)

Copied to clipboard

Challenge: Existing benchmarks for large language models are limited in scale, authenticity, and reliability due to the emotionally complex nature of therapeutic dialogue.
Approach: They propose two benchmarks that provide a framework for evaluating large language models for mental health support.
Outcome: The proposed framework provides a framework for generation and evaluation of large-scale authentic dialogue datasets and judge-reliability assessments.
Can AI Relate: Testing Large Language Model Response for Mental Health Support (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) are already being piloted for clinical use in hospitals . recent failures of the Tessa chatbot have led to doubts about their reliability in high-stakes settings.
Approach: They propose safety guidelines for the potential deployment of large language models for mental health response.
Outcome: The proposed framework measures equity in empathy and adherence of LLM responses to motivational interviewing theory.
BotChat: Evaluating LLMs’ Capabilities of Having Multi-Turn Dialogues (2024.findings-naacl)

Copied to clipboard

Challenge: Modern Large Language Models (LLMs) facilitate high-quality, multi-turn dialogues with humans, but human-based evaluation of such a capability requires substantial manual effort.
Approach: They propose to evaluate LLMs' ability to emulate human-like, multi-turn conversations using an LLM-centric approach.
Outcome: The proposed model emulates human-like, multi-turn conversations using an LLM-centric approach.
Interactive Evaluation for Medical LLMs via Task-oriented Dialogue System (2025.coling-main)

Copied to clipboard

Challenge: In typical medical scenarios, doctors often ask a set of questions to gain a comprehensive understanding of patients’ conditions.
Approach: They propose to use multi-turn medical dialogue evaluation to evaluate proactive communication and diagnostic capabilities of medical Large Language Models (LLMs) .
Outcome: The proposed model outperforms existing models on multi-turn question-answering datasets and is therefore cost-effective.
Towards Interpretable Mental Health Analysis with Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies on large language models lack adequate evaluations and prompting strategies for explainability.
Approach: They evaluate the mental health analysis and emotional reasoning ability of large language models (LLMs) using 11 datasets across 5 tasks.
Outcome: The proposed model shows strong in-context learning ability but still has a significant gap with advanced task-specific methods.
Integrating Visual Modalities with Large Language Models for Mental Health Support (2025.coling-main)

Copied to clipboard

Challenge: Existing work of mental health support primarily utilizes unimodal textual data and fails to understand and respond to users’ emotional states comprehensively.
Approach: They propose a framework that integrates multimodal inputs and counseling strategies to enhance the performance of Large Language Models (LLMs) This approach allows LLMs to generate more nuanced and supportive responses.
Outcome: The proposed framework outperforms existing models and delivers more empathetic, coherent, and contextually relevant mental health support responses.
Format Inertia: A Failure Mechanism of LLMs in Medical Pre-Consultation (2025.emnlp-industry)

Copied to clipboard

Challenge: Recent advances in Large Language Models have brought significant improvements to various service domains, including chatbots and medical pre-consultation applications.
Approach: They propose a method that rebalances the turn-count distribution of training data to mitigate Format Inertia in medical pre-consultation tasks.
Outcome: The proposed method significantly alleviates Format Inertia in medical pre-consultation tasks.
Deciphering Cognitive Distortions in Patient-Doctor Mental Health Conversations: A Multimodal LLM-Based Detection and Reasoning Framework (2024.emnlp-main)

Copied to clipboard

Challenge: Cognitive distortion research sheds light on pervasive errors in thinking patterns . authors present method for detecting and reasoning about cognitive distortions .
Approach: They propose a method for detecting and reasoning about cognitive distortions using Large Language Models.
Outcome: The proposed method improves accuracy and depth of detection and reasoning tasks in a zero-shot manner.
Do Large Language Models Align with Core Mental Health Counseling Competencies? (2025.findings-naacl)

Copied to clipboard

Challenge: Large language models are promising for mental health, but their alignment with core counseling competencies remains underexplored.
Approach: They propose a benchmark to evaluate 22 general-purpose and medical-finetuned LLMs across five key competencies.
Outcome: The proposed model outperforms generalist models in Intake, Assessment & Diagnosis but struggles with core counseling attributes and professional practice & ethics.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations