Can AI Relate: Testing Large Language Model Response for Mental Health Support (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Large language models (LLMs) are already being piloted for clinical use in hospitals . recent failures of the Tessa chatbot have led to doubts about their reliability in high-stakes settings. |
| Approach: | They propose safety guidelines for the potential deployment of large language models for mental health response. |
| Outcome: | The proposed framework measures equity in empathy and adherence of LLM responses to motivational interviewing theory. |
Similar Papers
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation (2026.eacl-long)
Copied to clipboard
Abeer Badawi, Elahe Rahimi, Md Tahmid Rahman Laskar, Sheri Grach, Lindsay Bertrand, Lames Danok, Prathiba Dhanesh, Jimmy Huang, Frank Rudzicz, Elham Dolatabadi
| Challenge: | Existing benchmarks for large language models are limited in scale, authenticity, and reliability due to the emotionally complex nature of therapeutic dialogue. |
| Approach: | They propose two benchmarks that provide a framework for evaluating large language models for mental health support. |
| Outcome: | The proposed framework provides a framework for generation and evaluation of large-scale authentic dialogue datasets and judge-reliability assessments. |
CBT-LLM: A Chinese Large Language Model for Cognitive Behavioral Therapy-based Mental Health Question Answering (2024.lrec-main)
Copied to clipboard
| Challenge: | Recent advances in artificial intelligence highlight the potential of language models in psychological health support. |
| Approach: | They propose a method to enhance the precision and efficacy of psychological support through large language models. |
| Outcome: | The proposed model generates professional and structured responses in Chinese psychological health Q&A tasks, showcasing its practicality and quality. |
Like a Therapist, But Not: Reddit Narratives of AI in Mental Health Contexts (2026.findings-acl)
Copied to clipboard
| Challenge: | Large language models are increasingly used for emotional support and mental health–related interactions outside clinical settings. |
| Approach: | They analyze 5,126 Reddit posts describing use of AI for emotional support or therapy . positive sentiment is most strongly associated with task and goal alignment, they say . |
| Outcome: | The proposed framework analyzes language, adoption-related attitudes, and relational alignment at scale. positive sentiment is most strongly associated with task and goal alignment. |
CBT-Bench: Evaluating Large Language Models on Assisting Cognitive Behavior Therapy (2025.naacl-long)
Copied to clipboard
Mian Zhang, Xianjun Yang, Xinlu Zhang, Travis Labrum, Jamie C. Chiu, Shaun M. Eack, Fei Fang, William Yang Wang, Zhiyu Chen
| Challenge: | Existing research has explored mental health condition classifications, empathetic conversations, and chatbots designed for simple discourse structures. |
| Approach: | They propose a benchmark for systematic evaluation of cognitive behavioral therapy assistance using Large Language Models (LLMs). |
| Outcome: | The proposed benchmark includes three levels of tasks covering key aspects of cognitive behavioral therapy that could be enhanced through AI assistance. |
Language Model Council: Democratically Benchmarking Foundation Models on Highly Subjective Tasks (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing evaluations of Large Language Models (LLMs) rely on a single large model to score outputs from other LLMs, but this is prone to intra-model bias and many tasks may be too subjective for a one model to judge fairly. |
| Approach: | They propose a language model council where a group of LLMs collaborate to create tests, respond to them, and evaluate each other’s responses to produce a ranking in a democratic fashion. |
| Outcome: | The proposed model produces rankings that are more separable and robust than any individual LLM judge. |
Towards Interpretable Mental Health Analysis with Large Language Models (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies on large language models lack adequate evaluations and prompting strategies for explainability. |
| Approach: | They evaluate the mental health analysis and emotional reasoning ability of large language models (LLMs) using 11 datasets across 5 tasks. |
| Outcome: | The proposed model shows strong in-context learning ability but still has a significant gap with advanced task-specific methods. |
A Survey of Large Language Models in Psychotherapy: Current Landscape and Future Directions (2025.findings-acl)
Copied to clipboard
Hongbin Na, Yining Hua, Zimu Wang, Tao Shen, Beibei Yu, Lilin Wang, Wei Wang, John Torous, Ling Chen
| Challenge: | Large language models (LLMs) can handle extensive context and multi-turn reasoning. |
| Approach: | They propose a taxonomy dividing psychotherapy into stages of assessment, diagnosis, and treatment to examine LLM advancements and challenges. |
| Outcome: | The proposed taxonomy reveals imbalances in current research, such as a focus on common disorders, linguistic biases, fragmented methods, and limited theoretical integration. |
Auditing LLM Responses to Harmful Stereotypes Targeting Mental Health Groups (2026.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) can exhibit imbalanced biases against vulnerable groups, but how they rationalize stereotypes and rights restrictions targeting mental health entities remains underexplored. |
| Approach: | They audit a suite of open-weight LLMs on stereotype-justification prompts tied to mental health identities. |
| Outcome: | The proposed models endorse harmful stereotypes when explicitly asked to justify them, with endorsement varying across model families, versions, and mental health conditions. |
Do Large Language Models Align with Core Mental Health Counseling Competencies? (2025.findings-naacl)
Copied to clipboard
Viet Cuong Nguyen, Mohammad Taher, Dongwan Hong, Vinicius Konkolics Possobom, Vibha Thirunellayi Gopalakrishnan, Ekta Raj, Zihang Li, Heather J. Soled, Michael L. Birnbaum, Srijan Kumar, Munmun De Choudhury
| Challenge: | Large language models are promising for mental health, but their alignment with core counseling competencies remains underexplored. |
| Approach: | They propose a benchmark to evaluate 22 general-purpose and medical-finetuned LLMs across five key competencies. |
| Outcome: | The proposed model outperforms generalist models in Intake, Assessment & Diagnosis but struggles with core counseling attributes and professional practice & ethics. |
Preference Learning Unlocks LLMs’ Psycho-Counseling Skills (2026.findings-acl)
Copied to clipboard
| Challenge: | Current LLMs struggle to consistently provide effective responses to client speeches due to the lack of supervision from high-quality real psycho-counseling data. |
| Approach: | They propose to use a dataset to evaluate therapists' responses to client speeches using a set of professional and comprehensive principles to evaluate their responses. |
| Outcome: | The proposed model achieves an impressive win rate of 87% against GPT-4o. |