The Objective and Subjective Sleepiness Voice Corpora (2020.lrec-1)

Copied to clipboard

Challenge: Following chronic sleep disorders involves multiple appointments between doctors and patients which often results in episodic follow-ups with unevenly spaced interviews.
Approach: They propose to use a large database to assess the sleepiness level of highly phenotyped patients that complain from excessive daytime sleepiness instead of healthy subjects.
Outcome: The proposed model is based on recordings from patients suffering from excessive daytime sleepiness instead of healthy subjects and incites them to sleep contrary to existing stressing sleepiness deprivation paradigms.

Similar Papers

Extracting Symptoms and their Status from Clinical Conversations (P19-1)

Copied to clipboard

Challenge: Existing models for extracting symptoms from clinical conversations are inherently difficult.
Approach: They propose two new deep learning models tailored for a new application . they propose a hierarchical span-attribute tagging model and a sequence-to-sequence model .
Outcome: The proposed models perform well under different conditions and are compared to existing models.
Still Not Quite There! Evaluating Large Language Models for Comorbid Mental Health Diagnosis (2024.emnlp-main)

Copied to clipboard

Challenge: ANGST is a benchmark for depression-anxiety comorbidity classification from social media posts.
Approach: They propose a social media-based benchmark for depression-anxiety comorbidity classification . ANGST enables multi-label classification, allowing each post to be simultaneously identified as indicating depression and/or anxiety.
Outcome: The proposed dataset enables multi-label classification of depression and anxiety . it outperforms existing models but none achieves an F1 score exceeding 72% .
How to Compare Automatically Two Phonological Strings: Application to Intelligibility Measurement in the Case of Atypical Speech (2020.lrec-1)

Copied to clipboard

Challenge: Atypical speech productions must be evaluated with regard to "typical" or "expected" productions . a first test of this method among healthy speakers and patients treated for cancer has proved its validity .
Approach: They propose a method to evaluate "atypical" speech productions based on phonological transcriptions . authors propose to use phonology to compute distances between phonologic forms produced and expected .
Outcome: The proposed method has been validated in a large population of healthy speakers and patients with cancer . it computes distances between phonological forms produced and expected from cost matrices based on features of phonemes .
Semi-Supervised Diseased Detection from Speech Dialogues with Multi-Level Data Modeling (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for learning from weakly-supervised speech data are hampered by severe data scarcity and the subjective nature of clinical annotations.
Approach: They propose a framework that explicitly models pathological traits by jointly learning from frame-level, segment-level and session-level representations within unsegmented clinical dialogues.
Outcome: The proposed framework is model-agnostic, robust across languages and conditions, and highly data-efficient.
Comprehensive Benchmarking of Long-Form Speech Generation in Diverse Scenarios (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluation benchmarks for long-form speech are limited to limited domains, creating a significant gap with the diverse downstream applications.
Approach: They propose a benchmark that decomposes "long-form speech quality" into specific, disentangled dimensions.
Outcome: The proposed benchmark decomposes “long-form speech quality” into specific, disentangled dimensions.
Dash-M5H: An Interactive Dashboard for Multi-Modal, Multi-Model Mental Health Assessment (2026.acl-demo)

Copied to clipboard

Challenge: Dash-M5H integrates transcript text, audio, and facial behavior with a clinically grounded VLM prediction pipeline that produces DSM-5-aligned depression predictions.
Approach: They propose a dashboard that integrates multimodal behavioral data with multi-model signal outputs of recorded clinical interviews.
Outcome: Dash-M5H is an interactive dashboard for *multi-modal, multi-model mental health assessment that integrates transcript text, audio, and facial behavior with a clinically grounded VLM prediction pipeline.
Beyond Discrete Personas: Personality Modeling Through Journal Intensive Conversations (2025.coling-main)

Copied to clipboard

Challenge: Existing LLMs rely on static, predefined personas to capture dynamic and evolving nature of human personalities.
Approach: They propose a dataset with 400,000 conversations and a framework for generating personalized conversations using long-form journal entries from Reddit.
Outcome: The proposed framework generates high-quality, personality-rich dialogues grounded in reddit journal entries.
IM^2: an Interpretable and Multi-category Integrated Metric Framework for Automatic Dialogue Evaluation (2022.emnlp-main)

Copied to clipboard

Challenge: Evaluation metrics for dialogue systems are expensive and time-consuming . current evaluation metrics focus on a single quality or several qualities .
Approach: They propose an interpretable, multi-faceted, and controllable framework to combine dialogue metrics which are good at measuring different qualities.
Outcome: The proposed framework integrates a large number of evaluation metrics to improve the performance of the model.
SMHD: a Large-Scale Resource for Exploring Online Language Usage for Multiple Mental Health Conditions (C18-1)

Copied to clipboard

Challenge: Existing methods to label mental health conditions are based on high-precision diagnosis patterns and carefully selected control users.
Approach: They propose to use high-precision diagnosis patterns to identify self-reported diagnoses of nine different mental health conditions and obtain high-quality labeled data without manual labelling.
Outcome: The proposed dataset is two orders of magnitude larger than the largest published similar resource.
Interpretable Assessment of Speech Intelligibility Using Deep Learning: A Case Study on Speech Disorders Due to Head and Neck Cancers (2024.lrec-main)

Copied to clipboard

Challenge: Using deep learning, speech disorders can be evaluated by perceptual measures, but they are subject to subjectivity and lack of reproducibility.
Approach: They propose to use deep-learning to explain hidden representations in a deep- learning speech model to provide a deeper understanding of the final intelligibility assessment of patients with Head and Neck Cancers.
Outcome: The proposed approach predicts speech intelligibility and severity of patients with Head and Neck Cancers while giving relevant interpretations of the final assessment at the phonemes and phonetic feature levels.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations