Reasoning Is Not All You Need: Examining LLMs for Multi-Turn Mental Health Conversations (2026.acl-long)
Copied to clipboard
| Challenge: | Existing evaluation frameworks focus on diagnostic accuracy and win-rates and often overlook alignment with patient-specific goals, values, and personalities required for meaningful conversations. |
| Approach: | They propose a framework for synthetically generating realistic, multi-turn mental health sensemaking conversations and a dataset to examine their models in healthcare settings. |
| Outcome: | The proposed framework synthesizes a dataset comprising over 2,200 patient–LLM conversations and evaluates them using human-centric criteria. |
Similar Papers
SoulChat: Improving LLMs’ Empathy, Listening, and Comfort Abilities through Fine-tuning with Multi-turn Empathy Conversations (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models (LLMs) are used in psychological counseling to provide universal advice. |
| Approach: | They constructed a multi-turn empathetic conversation dataset with 2 million samples . they found that the model's empathy ability is enhanced when finetuning . |
| Outcome: | Experiments show that large language models can be finetuned to provide empathy . but, when applied to mental health or emotional support conversation, there are three main issues . |
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation (2026.eacl-long)
Copied to clipboard
Abeer Badawi, Elahe Rahimi, Md Tahmid Rahman Laskar, Sheri Grach, Lindsay Bertrand, Lames Danok, Prathiba Dhanesh, Jimmy Huang, Frank Rudzicz, Elham Dolatabadi
| Challenge: | Existing benchmarks for large language models are limited in scale, authenticity, and reliability due to the emotionally complex nature of therapeutic dialogue. |
| Approach: | They propose two benchmarks that provide a framework for evaluating large language models for mental health support. |
| Outcome: | The proposed framework provides a framework for generation and evaluation of large-scale authentic dialogue datasets and judge-reliability assessments. |
Can AI Relate: Testing Large Language Model Response for Mental Health Support (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models (LLMs) are already being piloted for clinical use in hospitals . recent failures of the Tessa chatbot have led to doubts about their reliability in high-stakes settings. |
| Approach: | They propose safety guidelines for the potential deployment of large language models for mental health response. |
| Outcome: | The proposed framework measures equity in empathy and adherence of LLM responses to motivational interviewing theory. |
BotChat: Evaluating LLMs’ Capabilities of Having Multi-Turn Dialogues (2024.findings-naacl)
Copied to clipboard
Haodong Duan, Jueqi Wei, Chonghua Wang, Hongwei Liu, Yixiao Fang, Songyang Zhang, Dahua Lin, Kai Chen
| Challenge: | Modern Large Language Models (LLMs) facilitate high-quality, multi-turn dialogues with humans, but human-based evaluation of such a capability requires substantial manual effort. |
| Approach: | They propose to evaluate LLMs' ability to emulate human-like, multi-turn conversations using an LLM-centric approach. |
| Outcome: | The proposed model emulates human-like, multi-turn conversations using an LLM-centric approach. |
Interactive Evaluation for Medical LLMs via Task-oriented Dialogue System (2025.coling-main)
Copied to clipboard
| Challenge: | In typical medical scenarios, doctors often ask a set of questions to gain a comprehensive understanding of patients’ conditions. |
| Approach: | They propose to use multi-turn medical dialogue evaluation to evaluate proactive communication and diagnostic capabilities of medical Large Language Models (LLMs) . |
| Outcome: | The proposed model outperforms existing models on multi-turn question-answering datasets and is therefore cost-effective. |
Towards Interpretable Mental Health Analysis with Large Language Models (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies on large language models lack adequate evaluations and prompting strategies for explainability. |
| Approach: | They evaluate the mental health analysis and emotional reasoning ability of large language models (LLMs) using 11 datasets across 5 tasks. |
| Outcome: | The proposed model shows strong in-context learning ability but still has a significant gap with advanced task-specific methods. |
Integrating Visual Modalities with Large Language Models for Mental Health Support (2025.coling-main)
Copied to clipboard
| Challenge: | Existing work of mental health support primarily utilizes unimodal textual data and fails to understand and respond to users’ emotional states comprehensively. |
| Approach: | They propose a framework that integrates multimodal inputs and counseling strategies to enhance the performance of Large Language Models (LLMs) This approach allows LLMs to generate more nuanced and supportive responses. |
| Outcome: | The proposed framework outperforms existing models and delivers more empathetic, coherent, and contextually relevant mental health support responses. |
Format Inertia: A Failure Mechanism of LLMs in Medical Pre-Consultation (2025.emnlp-industry)
Copied to clipboard
| Challenge: | Recent advances in Large Language Models have brought significant improvements to various service domains, including chatbots and medical pre-consultation applications. |
| Approach: | They propose a method that rebalances the turn-count distribution of training data to mitigate Format Inertia in medical pre-consultation tasks. |
| Outcome: | The proposed method significantly alleviates Format Inertia in medical pre-consultation tasks. |
Deciphering Cognitive Distortions in Patient-Doctor Mental Health Conversations: A Multimodal LLM-Based Detection and Reasoning Framework (2024.emnlp-main)
Copied to clipboard
| Challenge: | Cognitive distortion research sheds light on pervasive errors in thinking patterns . authors present method for detecting and reasoning about cognitive distortions . |
| Approach: | They propose a method for detecting and reasoning about cognitive distortions using Large Language Models. |
| Outcome: | The proposed method improves accuracy and depth of detection and reasoning tasks in a zero-shot manner. |
Do Large Language Models Align with Core Mental Health Counseling Competencies? (2025.findings-naacl)
Copied to clipboard
Viet Cuong Nguyen, Mohammad Taher, Dongwan Hong, Vinicius Konkolics Possobom, Vibha Thirunellayi Gopalakrishnan, Ekta Raj, Zihang Li, Heather J. Soled, Michael L. Birnbaum, Srijan Kumar, Munmun De Choudhury
| Challenge: | Large language models are promising for mental health, but their alignment with core counseling competencies remains underexplored. |
| Approach: | They propose a benchmark to evaluate 22 general-purpose and medical-finetuned LLMs across five key competencies. |
| Outcome: | The proposed model outperforms generalist models in Intake, Assessment & Diagnosis but struggles with core counseling attributes and professional practice & ethics. |