PsychePass: Calibrating LLM Therapeutic Competence via Trajectory-Anchored Tournaments (2026.findings-acl)
Copied to clipboard
| Challenge: | evaluating therapeutic competence of large language models remains challenging due to unstructured and longitudinal nature of counseling. |
| Approach: | They propose a framework that calibrates the therapeutic competence of LLMs via trajectory-anchored tournaments. |
| Outcome: | The proposed framework calibrates the therapeutic competence of LLMs via trajectory-anchored tournaments. |
Similar Papers
DialogGuard: Multi-Agent Psychosocial Safety Evaluation Interface of Sensitive LLM Responses (2026.acl-demo)
Copied to clipboard
| Challenge: | Existing tools do not surface subtler psychosocial harms, nor provide explainable rationales that practitioners need. |
| Approach: | They propose an open-source system that lets practitioners inspect, stress-test, and create audit trails for prompted LLM agents across five psychosocial safety dimensions. |
| Outcome: | The open-source DialogGuard system lets practitioners inspect, stress-test, and create audit trails for prompted LLM agents across five psychosocial safety dimensions. |
Preference Learning Unlocks LLMs’ Psycho-Counseling Skills (2026.findings-acl)
Copied to clipboard
| Challenge: | Current LLMs struggle to consistently provide effective responses to client speeches due to the lack of supervision from high-quality real psycho-counseling data. |
| Approach: | They propose to use a dataset to evaluate therapists' responses to client speeches using a set of professional and comprehensive principles to evaluate their responses. |
| Outcome: | The proposed model achieves an impressive win rate of 87% against GPT-4o. |
Feel the Difference? A Comparative Analysis of Emotional Arcs in Real and LLM-Generated CBT Sessions (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Synthetic therapy dialogues generated by large language models (LLMs) lack the nuanced emotional dynamics of real therapy. |
| Approach: | They introduce a dataset of authentic cognitive behavioral therapy dialogues and analyze emotional arcs between real and LLM-generated CBT sessions. |
| Outcome: | The proposed dataset is a comparative analysis of emotional arcs between real and LLM-generated CBT sessions. |
ESC-Judge: A Framework for Comparing Emotional Support Conversational Agents (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) increasingly power mental-health chatbots . yet the field lacks a scalable, theory-grounded way to decide which model is more effective to deploy. |
| Approach: | They propose a framework that grounds head-to-head comparisons of Emotional-Support LLMs in Hill’s Exploration–Insight–Action counselling model. |
| Outcome: | The proposed framework matches PhD-level annotators in 85% of Exploration, 83% of Insight, and 86% of Action decisions, demonstrating human-level reliability at a fraction of the cost. |
Can You Share Your Story? Modeling Clients’ Metacognition and Openness for LLM Therapist Evaluation (2025.findings-acl)
Copied to clipboard
Minju Kim, Dongje Yoo, Yeonjun Hwang, Minseok Kang, Namyoung Kim, Minju Gwak, Beong-woo Kwak, Hyungjoo Chae, Harim Kim, Yunjoong Lee, Min Hee Kim, Dayi Jung, Kyong-Mee Chung, Jinyoung Yeo
| Challenge: | Existing evaluation methods for psychological counseling rely on client simulators that clearly disclose internal states to the therapist, making it difficult to determine whether an LLM therapist can uncover unexpressed perspectives. |
| Approach: | They propose a new evaluation framework featuring a controllable and realistic client simulator which dynamically adapts itself based on the ongoing counseling session. |
| Outcome: | The proposed evaluation framework features a realistic and controllable client simulator which dynamically adapts itself based on the ongoing counseling session, offering a more realistic and challenging evaluation environment. |
Are LLMs effective psychological assessors? Leveraging adaptive RAG for interpretable mental health screening through psychometric practice (2025.acl-long)
Copied to clipboard
| Challenge: | standardized questionnaires are essential tools for mental health screening, but computational approaches bypass these tools in favor of black-box classification. |
| Approach: | They propose a questionnaire-guided screening framework that bridges psychological practice and computational methods through adaptive Retrieval-Augmented Generation. |
| Outcome: | The proposed framework matches or outperforms state-of-the-art performance on Reddit-based benchmarks and extends to self-harm screening. |
Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness (2026.acl-long)
Copied to clipboard
Amin Banayeeanzade, Ala N. Tak, Fatemeh Bahrani, Anahita Bolourani, Leonardo Blas, Emilio Ferrara, Jonathan Gratch, Sai Praneeth Karimireddy
| Challenge: | Using a model with a high degree of emotion and personality control, large language models can be used to control socially interactive interactions. |
| Approach: | They propose a Psychologically-informed benchmark to evaluate LLM steering effectiveness and trustworthiness across emotion and personality domains. |
| Outcome: | The framework establishes the first holistic evaluation of emotion and personality steering, offering insights into its interpretability and reliability for socially interactive applications. |
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation (2026.eacl-long)
Copied to clipboard
Abeer Badawi, Elahe Rahimi, Md Tahmid Rahman Laskar, Sheri Grach, Lindsay Bertrand, Lames Danok, Prathiba Dhanesh, Jimmy Huang, Frank Rudzicz, Elham Dolatabadi
| Challenge: | Existing benchmarks for large language models are limited in scale, authenticity, and reliability due to the emotionally complex nature of therapeutic dialogue. |
| Approach: | They propose two benchmarks that provide a framework for evaluating large language models for mental health support. |
| Outcome: | The proposed framework provides a framework for generation and evaluation of large-scale authentic dialogue datasets and judge-reliability assessments. |
InMind: Evaluating LLMs in Capturing and Applying Individual Human Reasoning Styles (2025.emnlp-main)
Copied to clipboard
Zizhen Li, Chuanhao Li, Yibin Wang, Qi Chen, Diping Song, Yukang Feng, Jianwen Sun, Jiaxin Ai, Fanrui Zhang, Mingzhu Sun, Kaipeng Zhang
| Challenge: | Recent large language models (LLMs) have demonstrated strong reasoning abilities across complex mathematical and scientific domains. |
| Approach: | They propose a framework to assess whether LLMs can capture and apply personalized reasoning styles in social deduction games. |
| Outcome: | The proposed framework evaluates LLMs on the game Avalon and shows that they can capture and apply individualized reasoning styles. |
ReflAct: World-Grounded Decision Making in LLM Agents via Goal-State Reflection (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in LLMs have significantly enhanced their reasoning capabilities, enabling LLM-based agents to perform complex multi-step decision making beyond static problem solving. |
| Approach: | They propose a novel reasoning backbone that shifts reasoning from merely planning next actions to continuously reflecting on the agent’s state relative to its goal. |
| Outcome: | The proposed model outperforms ReAct by 27.7% on average, achieving a 93.3% success rate in ALFWorld. |