Challenge: standardized questionnaires are essential tools for mental health screening, but computational approaches bypass these tools in favor of black-box classification.
Approach: They propose a questionnaire-guided screening framework that bridges psychological practice and computational methods through adaptive Retrieval-Augmented Generation.
Outcome: The proposed framework matches or outperforms state-of-the-art performance on Reddit-based benchmarks and extends to self-harm screening.

Similar Papers

LLM Questionnaire Completion for Automatic Psychiatric Assessment (2024.findings-emnlp)

Copied to clipboard

Challenge: Psychiatric evaluations are heavily based on patient verbal reports of disturbed feelings, thoughts, behaviors, and their changes over time.
Approach: They employ a Large Language Model to convert unstructured psychological interviews into structured questionnaires spanning various psychiatric and personality domains.
Outcome: The proposed model improves diagnostic accuracy compared to baselines.
ALBA: Adaptive Language-Based Assessments for Mental Health (2024.naacl-long)

Copied to clipboard

Challenge: Adaptive language-based assessments require a substantial sample of words per person for accuracy.
Approach: They propose an adaptive language-based assessment task that involves ordering questions and scoring latent psychological trait using limited language responses to previous questions.
Outcome: The proposed methods improve over non-adaptive baselines, but are more accurate and scalable with fewer questions.
When Can We Trust LLMs in Mental Health? Large-Scale Benchmarks for Reliable LLM Evaluation (2026.eacl-long)

Copied to clipboard

Challenge: Existing benchmarks for large language models are limited in scale, authenticity, and reliability due to the emotionally complex nature of therapeutic dialogue.
Approach: They propose two benchmarks that provide a framework for evaluating large language models for mental health support.
Outcome: The proposed framework provides a framework for generation and evaluation of large-scale authentic dialogue datasets and judge-reliability assessments.
The Colorful Future of LLMs: Evaluating and Improving LLMs as Emotional Supporters for Queer Youth (2024.naacl-long)

Copied to clipboard

Challenge: Queer youth face increased mental health risks, such as depression, anxiety, and suicidal ideation.
Approach: They propose a scale that is inspired by psychological standards and expert input to evaluate LLM's interactions with queer-related content.
Outcome: The proposed scale outperforms human responses to queer-related content and outperformed LLMs in the qualitative and quantitative analysis.
Explainable Depression Detection in Clinical Interviews with Personalized Retrieval-Augmented Generation (2025.findings-acl)

Copied to clipboard

Challenge: Existing systems rely on black-box neural networks, which lack interpretability, which is crucial in mental health contexts.
Approach: They propose a Retrieval-augmented generation framework for Explainable depression detection that retrieves evidence from clinical interview transcripts, providing explanations for predictions.
Outcome: The proposed framework retrieves evidence from clinical interview transcripts, providing explanations for predictions.
LLM-Independent Adaptive RAG: Let the Question Speak for Itself (2025.emnlp-main)

Copied to clipboard

Challenge: Existing methods to retrieve Large Language Models (LLMs) are inefficient and impractical.
Approach: They propose a lightweight adaptive retrieval method that leverages external information to achieve comparable quality while achieving significant efficiency gains.
Outcome: The proposed methods achieve comparable quality while achieving significant efficiency gains on 6 QA datasets.
Can AI Relate: Testing Large Language Model Response for Mental Health Support (2024.findings-emnlp)

Copied to clipboard

Challenge: Large language models (LLMs) are already being piloted for clinical use in hospitals . recent failures of the Tessa chatbot have led to doubts about their reliability in high-stakes settings.
Approach: They propose safety guidelines for the potential deployment of large language models for mental health response.
Outcome: The proposed framework measures equity in empathy and adherence of LLM responses to motivational interviewing theory.
Towards Interpretable Mental Health Analysis with Large Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies on large language models lack adequate evaluations and prompting strategies for explainability.
Approach: They evaluate the mental health analysis and emotional reasoning ability of large language models (LLMs) using 11 datasets across 5 tasks.
Outcome: The proposed model shows strong in-context learning ability but still has a significant gap with advanced task-specific methods.
Fact, Fetch, and Reason: A Unified Evaluation of Retrieval-Augmented Generation (2025.naacl-long)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have significantly enhanced their capabilities across various cognitive tasks.
Approach: They propose a high-quality evaluation dataset to test LLMs' ability to provide factual responses, assess retrieval capabilities, and evaluate the reasoning required to generate final answers.
Outcome: The proposed framework improves performance in end-to-end RAG scenarios.
Are LLM-based Evaluators Confusing NLG Quality Criteria? (2024.acl-long)

Copied to clipboard

Challenge: Existing studies show that LLMs confuse evaluation criteria, which reduces their reliability.
Approach: They propose a hierarchical classification system for 11 common aspects with corresponding different evaluation criteria.
Outcome: The proposed system is based on 11 common aspects with different evaluation criteria.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations