A Swedish Cookie-Theft Corpus (L18-1)

Copied to clipboard

Challenge: Language disturbances can be a diagnostic marker for neurodegenerative diseases, such as Alzheimer's disease, at earlier stages.
Approach: They develop a corpus of audio recordings of the Cookie-theft, a standardized test that has been used in studies in the past.
Outcome: The proposed corpus is based on audio recordings of the Cookie-theft . it provides a rich resource for future research and experimentation in many areas .

Similar Papers

The Slovak Autistic and Non-Autistic Child Speech Corpus:Task-Oriented Child-Adult Interactions (2024.lrec-main)

Copied to clipboard

Challenge: Presented is the Slovak Autistic and Non-Autistic Child Speech Corpus . corpus contains over 15 hours of speech .
Approach: They present a Slovak autistic and non-autistic child speech corpus . the corpus was primarily recorded to investigate lexical alignment .
Outcome: The Slovak Autistic and Non-Autistic Child Speech Corpus contains over 15 hours of speech . the corpus can be shared with researchers and replicated in future research .
LoSST-AD: A Longitudinal Corpus for Tracking Alzheimer’s Disease Related Changes in Spontaneous Speech (2024.lrec-main)

Copied to clipboard

Challenge: Language-based biomarkers have shown promising results in differentiating those with Alzheimer’s disease (AD) diagnosis from healthy individuals, but the earliest changes in language are thought to start years or even decades before the diagnosis.
Approach: They propose to use transcripts of public interviews with 20 famous figures to track language change over several decades to validate their corpus.
Outcome: The proposed corpus can provide a valuable starting point for the development of early detection tools and enhance our understanding of how AD affects language over time.
Evaluating Sentence Segmentation in Different Datasets of Neuropsychological Language Tests in Brazilian Portuguese (2020.lrec-1)

Copied to clipboard

Challenge: Using automated analysis of connected speech is a promising direction for diagnosing cognitive impairments.
Approach: They propose to use a novel model to segment impaired speech transcriptions . they propose to include a Linear Chain CRF and a self-attention mechanism .
Outcome: The proposed system performs better than the existing model with three new datasets used to diagnose cognitive impairments.
SLaCAD: A Spoken Language Corpus for Early Alzheimer’s Disease Detection (2024.lrec-main)

Copied to clipboard

Challenge: Recent studies show that cerebrospinal fluid (CSF) levels serve as useful early biomarkers for identifying early AD, but CSF biomarker collection is challenging.
Approach: They propose to use speech data to identify early Alzheimer's disease (AD) trajectory to identify cognitive deficits in early disease stages.
Outcome: The proposed dataset relates speech and speech characteristics with CSF and plasma biomarkers to clinical diagnoses, CSF levels, and biomarker scores.
Elderly Conversational Speech Corpus with Cognitive Impairment Test and Pilot Dementia Detection Experiment Using Acoustic Characteristics of Speech in Japanese Dialects (2022.lrec-1)

Copied to clipboard

Challenge: Several studies have explored using only the acoustic and linguistic information of conversational speech as diagnostic material, with some success.
Approach: They propose to use acoustic features of conversational speech to detect dementia even when dialects are present.
Outcome: The proposed method detects dementia using acoustic features even when dialects are present even when spoken from two regions.
Reformulating NLP tasks to Capture Longitudinal Manifestation of Language Disorders in People with Dementia. (2023.emnlp-main)

Copied to clipboard

Challenge: Dementia is associated with language disorders which impede communication.
Approach: They propose to use a pre-trained language model to automatically learn linguistic disorder patterns by forcing it to focus on reformulated natural language processing (NLP) tasks and associated linguistic patterns.
Outcome: The proposed communication marker outperforms existing linguistic approaches and shows external validity via significant correlation with clinical markers of behaviour.
Multilingual prediction of Alzheimer’s disease through domain adaptation and concept-based language modelling (N19-1)

Copied to clipboard

Challenge: Existing work on speech and language models has been limited by the size of available datasets.
Approach: They propose to augment a small French dataset with a much larger English dataset to augment the language model to model the order in which information units are produced by dementia patients and controls.
Outcome: The proposed model improves classification performance in English and French separately.
Analyzing Gambling Addictions: A Spanish Corpus for Understanding Pathological Behavior (2025.findings-emnlp)

Copied to clipboard

Challenge: a new study examines the interaction between natural language use and gambling disorders.
Approach: They build a new corpus of sentences that are searched and compared using top-k pooling to form the assessment pools of sentences.
Outcome: The proposed model is based on a new corpus of sentences in spanish .
Augmenting word2vec with latent Dirichlet allocation within a clinical application (N19-1)

Copied to clipboard

Challenge: Existing models that combine latent Dirichlet allocation and word embedding for distinguishing between speakers with and without Alzheimer’s disease from transcripts of picture descriptions are not suitable for clinical binary text classification tasks.
Approach: They propose three models that combine latent Dirichlet allocation and word embedding for distinguishing between speakers with and without Alzheimer’s disease from transcripts of picture descriptions.
Outcome: The proposed models outperform word2vec and LDA models on a clinical binary text classification task.
Towards Domain-Agnostic and Domain-Adaptive Dementia Detection from Spoken Language (2023.acl-long)

Copied to clipboard

Challenge: Domain adaptation (DA) techniques have been used to improve performance of NLP systems for healthcare tasks due to numerous complexities of data.
Approach: They propose to use domain adaptation techniques to improve generalizability across diverse datasets for dementia detection.
Outcome: The proposed model achieves a 22% increase in accuracy adapting from a conversational to task-oriented dataset compared to a jointly trained baseline.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations