Leveraging Language Models for Summarizing Mental State Examinations: A Comprehensive Evaluation and Dataset Release (2025.coling-main)
Copied to clipboard
| Challenge: | Mental health disorders affect a significant portion of the global population . access to mental health support is limited in developing countries . |
| Approach: | They evaluated a 12-item descriptive MSE questionnaire and five well-known summarization models . they found that language models can generate coherent MSE summaries for doctors . |
| Outcome: | The proposed model can generate coherent summaries from MSEs in a conversational format. |
Similar Papers
MentSum: A Resource for Exploring Summarization of Mental Health Online Posts (2022.lrec-1)
Copied to clipboard
| Challenge: | Mental health remains a significant challenge of public health worldwide . many use online platforms to share their mental health conditions and seek help . |
| Approach: | They analyze a dataset of over 24k user posts from Reddit and 43 mental health subreddits to generate a short summarization. |
| Outcome: | The proposed dataset compared over 24k user posts and 43 mental health subreddits . it shows that the summarization of these posts is faster and more accurate than previous studies. |
Towards Interpretable Mental Health Analysis with Large Language Models (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies on large language models lack adequate evaluations and prompting strategies for explainability. |
| Approach: | They evaluate the mental health analysis and emotional reasoning ability of large language models (LLMs) using 11 datasets across 5 tasks. |
| Outcome: | The proposed model shows strong in-context learning ability but still has a significant gap with advanced task-specific methods. |
Systematic Evaluation of Auto-Encoding and Large Language Model Representations for Capturing Author States and Traits (2025.findings-acl)
Copied to clipboard
Khushboo Singh, Vasudha Varadarajan, Adithya V Ganesan, August Håkan Nilsson, Nikita Soni, Syeda Mahwish, Pranav Chitale, Ryan L. Boyd, Lyle Ungar, Richard N Rosenthal, H. Schwartz
| Challenge: | Large Language Models (LLMs) are increasingly used in human-centered applications, yet their ability to model diverse psychological constructs is not well understood. |
| Approach: | They evaluated a range of Transformer-LMs to predict psychological variables across five major dimensions: affect, substance use, mental health, sociodemographics, and personality. |
| Outcome: | The models predict affect, substance use, mental health, sociodemographics, and personality across five major dimensions. |
FineSurE: Fine-grained Summarization Evaluation using LLMs (2024.acl-long)
Copied to clipboard
| Challenge: | Existing methods for text summarization evaluation do not correlate well with human judgments . evaluators that use Likert scale scores are limited in their ability to perform deeper analysis. |
| Approach: | They propose a fine-grained evaluator specifically tailored for the summarization task using large language models. |
| Outcome: | The proposed method improves on open-source and proprietary LLMs and shows better completeness and conciseness than existing methods. |
SMHD: a Large-Scale Resource for Exploring Online Language Usage for Multiple Mental Health Conditions (C18-1)
Copied to clipboard
| Challenge: | Existing methods to label mental health conditions are based on high-precision diagnosis patterns and carefully selected control users. |
| Approach: | They propose to use high-precision diagnosis patterns to identify self-reported diagnoses of nine different mental health conditions and obtain high-quality labeled data without manual labelling. |
| Outcome: | The proposed dataset is two orders of magnitude larger than the largest published similar resource. |
Towards Comprehensive Language Analysis for Clinically Enriched Spontaneous Dialogue (2024.lrec-main)
Copied to clipboard
| Challenge: | Contemporary NLP has progressed from feature-based classification to fine-tuning and prompt-based techniques . many of these techniques remain understudied in the context of real-world, clinically enriched spontaneous dialogue. |
| Approach: | They investigate the efficacy and overall performance of a range of NLP techniques on transcribed speech from patients with schizophrenia and other disorders. |
| Outcome: | The proposed methods are effective in analyzing transcribed speech from patients with schizophrenia and healthy controls taking a clinically-validated language test. |
Re-Evaluating Evaluation for Multilingual Summarization (2024.emnlp-main)
Copied to clipboard
Jessica Forde, Ruochen Zhang, Lintang Sutawika, Alham Aji, Samuel Cahyawijaya, Genta Winata, Minghao Wu, Carsten Eickhoff, Stella Biderman, Ellie Pavlick
| Challenge: | Existing studies have shown that automated evaluation approaches correlate with human ratings in English, but this is unclear for other languages. |
| Approach: | They construct a small-scale pilot dataset containing article-summary pairs and human ratings in English, Chinese and Indonesian to measure the strength of summaries. |
| Outcome: | The results show that standard metrics are unreliable measures of quality in Chinese and Indonesian. |
Reasoning Is Not All You Need: Examining LLMs for Multi-Turn Mental Health Conversations (2026.acl-long)
Copied to clipboard
| Challenge: | Existing evaluation frameworks focus on diagnostic accuracy and win-rates and often overlook alignment with patient-specific goals, values, and personalities required for meaningful conversations. |
| Approach: | They propose a framework for synthetically generating realistic, multi-turn mental health sensemaking conversations and a dataset to examine their models in healthcare settings. |
| Outcome: | The proposed framework synthesizes a dataset comprising over 2,200 patient–LLM conversations and evaluates them using human-centric criteria. |
LLM Questionnaire Completion for Automatic Psychiatric Assessment (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Psychiatric evaluations are heavily based on patient verbal reports of disturbed feelings, thoughts, behaviors, and their changes over time. |
| Approach: | They employ a Large Language Model to convert unstructured psychological interviews into structured questionnaires spanning various psychiatric and personality domains. |
| Outcome: | The proposed model improves diagnostic accuracy compared to baselines. |
Towards Multi-dimensional Evaluation of LLM Summarization across Domains and Languages (2025.acl-long)
Copied to clipboard
Hyangsuk Min, Yuho Lee, Minjeong Ban, Jiaqi Deng, Nicole Hee-Yeon Kim, Taewon Yun, Hang Su, Jason Cai, Hwanjun Song
| Challenge: | Existing evaluation frameworks for text summarization lack domain-specific assessment criteria and are predominantly English-centric. |
| Approach: | They propose a multi-dimensional, multi-domain evaluation of summarization in English and Chinese that incorporates specialized assessment criteria for each domain and leverages a debate system to enhance annotation quality. |
| Outcome: | The proposed evaluation framework provides a multi-dimensional, multi-domain evaluation of summarization in English and Chinese. |