Carcinologic Speech Severity Index Project: A Database of Speech Disorder Productions to Assess Quality of Life Related to Speech After Cancer (L18-1)
Copied to clipboard
Corine Astésano, Mathieu Balaguer, Jérôme Farinas, Corinne Fredouille, Pascal Gaillard, Alain Ghio, Imed Laaridh, Muriel Lalain, Benoît Lepage, Julie Mauclair, Olivier Nocaudie, Julien Pinquier, Oriol Pont, Gilles Pouchoulin, Michèle Puech, Danièle Robert, Etienne Sicard, Virginie Woisard
| Challenge: | Increasing mortality in cancerology highlights the importance of reducing the impact on the Quality of Life after cancer treatment. |
| Approach: | They collect a large database of french speech recordings aimed at validating Disorder Severity Indexes. |
| Outcome: | The collected data will be available to the scientific community through the GIS Parolotheque. |
Similar Papers
Interpretable Assessment of Speech Intelligibility Using Deep Learning: A Case Study on Speech Disorders Due to Head and Neck Cancers (2024.lrec-main)
Copied to clipboard
Sondes Abderrazek, Corinne Fredouille, Alain Ghio, Muriel Lalain, Christine Meunier, Mathieu Balaguer, Virginie Woisard
| Challenge: | Using deep learning, speech disorders can be evaluated by perceptual measures, but they are subject to subjectivity and lack of reproducibility. |
| Approach: | They propose to use deep-learning to explain hidden representations in a deep- learning speech model to provide a deeper understanding of the final intelligibility assessment of patients with Head and Neck Cancers. |
| Outcome: | The proposed approach predicts speech intelligibility and severity of patients with Head and Neck Cancers while giving relevant interpretations of the final assessment at the phonemes and phonetic feature levels. |
How to Compare Automatically Two Phonological Strings: Application to Intelligibility Measurement in the Case of Atypical Speech (2020.lrec-1)
Copied to clipboard
| Challenge: | Atypical speech productions must be evaluated with regard to "typical" or "expected" productions . a first test of this method among healthy speakers and patients treated for cancer has proved its validity . |
| Approach: | They propose a method to evaluate "atypical" speech productions based on phonological transcriptions . authors propose to use phonology to compute distances between phonologic forms produced and expected . |
| Outcome: | The proposed method has been validated in a large population of healthy speakers and patients with cancer . it computes distances between phonological forms produced and expected from cost matrices based on features of phonemes . |
The MonPaGe_HA Database for the Documentation of Spoken French Throughout Adulthood (L18-1)
Copied to clipboard
| Challenge: | Existing studies on life-span changes in the speech of adults are mainly based on English speakers and few studies have compared more than two extreme age groups. |
| Approach: | They describe a MonPaGe_HealthyAdults database of spoken french with 405 speakers aged from 20 to 93 years old. |
| Outcome: | The proposed database includes 405 speakers aged 20 to 93 years old and includes 4 regiolects. |
SpeechLLM-as-Judges: Towards General and Interpretable Speech Quality Evaluation (2026.acl-long)
Copied to clipboard
Hui Wang, Jinghua Zhao, Yifan Yang, Shujie Liu, Junyang Chen, Yanzhe Zhang, Shiwan Zhao, Jinyu Li, Jiaming Zhou, Haoqin Sun, Yan Lu, Yong Qin
| Challenge: | Existing methods for evaluating the perceptual quality of synthetic speech are limited due to the complexity of perceptual quality factors and the diversity of speech generation tasks. |
| Approach: | They propose a new paradigm for enabling large language models to conduct structured speech quality evaluation using a large-scale dataset. |
| Outcome: | The proposed model performs well across tasks and languages. |
Corpora of Disordered Speech in the Light of the GDPR: Two Use Cases from the DELAD Initiative (2020.lrec-1)
Copied to clipboard
| Challenge: | Corpora of disordered speech (CDS) are costly to collect and difficult to share due to personal data protection and IP issues. |
| Approach: | a new paper examines the legal grounds for processing corpora of disordered speech . it illustrates how consent and public interest are taken into consideration . the paper also examines how public interest research can be used to obtain consent . |
| Outcome: | a new study examines the legal grounds for processing corpora of disordered speech (CDS) two use cases illustrate the legal basis for processing CDS in light of the GDPR . |
Subjective Evaluation of Comprehensibility in Movie Interactions (2020.lrec-1)
Copied to clipboard
| Challenge: | Various studies have dealt with the comprehensibility of textual, audio, or audiovisual documents. |
| Approach: | They aim to build a corpus of human annotations that could help to study human perceptions of comprehensibility of audiovisual documents. |
| Outcome: | The proposed corpus of human annotations will help to study human perceptions of comprehensibility of audiovisual documents. |
Recent Advances in Speech Language Models: A Survey (2025.acl-long)
Copied to clipboard
Wenqian Cui, Dianzhi Yu, Xiaoqi Jiao, Ziqiao Meng, Guangyan Zhang, Qichao Wang, Steven Y. Guo, Irwin King
| Challenge: | Text-based Large Language Models (LLMs) are a promising solution for end-to-end speech interaction. |
| Approach: | They propose to build a framework that allows users to input text and translate it into speech . they propose to use a text-only LLM and a "textto-speech" framework to generate a response based on this transcription . |
| Outcome: | The survey offers an overview of recent approaches to building SpeechLMs . it outlines core architectural components, training methodologies, evaluation strategies and challenges . |
Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (2022.naacl-main)
Copied to clipboard
| Challenge: | Silver es high quality training datasets for d AI training data labelilng services academic institutions engaged in ce R&D and application research to processing (NLP), voice recognition hesis (TTS), and computer vision (CV). |
| Approach: | Silver es provides high quality training datasets for d AI training data labelilng services academic institutions engaged in ce R&D and application research to processing (NLP), voice recognition hesis (TTS), and computer vision (CV). |
| Outcome: | Silver es high quality training datasets for d AI training data labelilng services academic institutions engaged in ce R&D and application research to processing (NLP), voice recognition hesis (TTS), and computer vision (CV). |
PRODIS - a Speech Database and a Phoneme-based Language Model for the Study of Predictability Effects in Polish (2024.lrec-main)
Copied to clipboard
| Challenge: | acoustic predictability is operationalised by surprisal in Polish, but cross-linguistic differences depend on prosodic system. |
| Approach: | They present a speech database and a phoneme-level language model of Polish . they aim to study contextual predictability effects on acoustic distinctiveness . |
| Outcome: | The proposed model is the first large, publicly available speech database of Polish . it is based on a light GPT architecture and can be expanded to other languages . |
Towards Holistic Evaluation of Large Audio-Language Models: A Comprehensive Survey (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent advances in large audio-language models (LALMs) have expanded their impact beyond natural language processing (NLP) to multimodal domains. |
| Approach: | They propose a systematic taxonomy for LALM evaluations, categorizing them into four dimensions based on their objectives: (1) General Auditory Awareness and Processing, (2) Knowledge and Reasoning, (3) Dialogue-oriented Ability, and (4) Fairness, Safety, and Trustworthiness. |
| Outcome: | The proposed taxonomy categorizes LALM evaluations into four dimensions based on their objectives and highlights challenges in this field. |