Challenge: Existing studies on the detection of Alzheimer's disease focus on the diagnosis of dementia instead.
Approach: They propose to use a dataset to develop NLP models for the detection of Mild Cognitive Impairment (MCI) MCI is a progressive neurodegenerative disorder associated with memory loss and declines in major brain functions including semantic and pragmatic levels of language processing.
Outcome: The proposed model performs best on 74.1% of the 3 topics studied.

Similar Papers

SLaCAD: A Spoken Language Corpus for Early Alzheimer’s Disease Detection (2024.lrec-main)

Copied to clipboard

Challenge: Recent studies show that cerebrospinal fluid (CSF) levels serve as useful early biomarkers for identifying early AD, but CSF biomarker collection is challenging.
Approach: They propose to use speech data to identify early Alzheimer's disease (AD) trajectory to identify cognitive deficits in early disease stages.
Outcome: The proposed dataset relates speech and speech characteristics with CSF and plasma biomarkers to clinical diagnoses, CSF levels, and biomarker scores.
Enriching Neural Models with Targeted Features for Dementia Detection (P19-2)

Copied to clipboard

Challenge: In the United States, adults over 65 are expected to comprise one-fifth of the population by 2030, and a larger proportion of the . population than those under 18 by 2035.
Approach: They propose a neural model that takes into account both long language samples and hand-crafted linguistic features to distinguish between dementia affected and healthy patients.
Outcome: The proposed model achieves an F1 score of 0.929 on the DementiaBank dataset and the state-of-the-art on the dataset.
An LLM-based Temporal-spatial Data Generation and Fusion Approach for Early Detection of Late Onset Alzheimer’s Disease (LOAD) Stagings Especially in Chinese and English-speaking Populations (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches struggle with temporal-spatial challenges in capturing subtle linguistic shifts across different disease stages.
Approach: They propose a large language model-driven T-S fusion framework that integrates multilingual LLMs, contrastive learning and interpretable marker discovery to revolutionize late onset AD detection.
Outcome: The proposed framework achieves state-of-the-art performance in late onset AD detection while enabling cross-linguistic diagnostics.
A Hierarchical Neural Attention-based Text Classifier (D18-1)

Copied to clipboard

Challenge: Existing hierarchical classification models are unable to handle large corpora and the number of categories increases with increasing corpus.
Approach: They propose to use external knowledge to introduce a hierarchical neural attention-based classifier to help with the classification of documents.
Outcome: The proposed model performs better than or comparable to state-of-the-art hierarchical models at significantly lower computational cost while maintaining high interpretability.
Transforming Brainwaves into Language: EEG Microstates Meet Text Embedding Models for Dementia Detection (2025.acl-srw)

Copied to clipboard

Challenge: Dementia is recognised as the seventh leading cause of mortality globally and plays a major role in increasing disability and dependence among older adults.
Approach: They propose to represent electroencephalography microstates as symbolic, language-like sequences and use text embedding and time-series deep learning models for classification.
Outcome: The proposed method achieves a high accuracy of 94.31% on 1001 EEG data from multiple countries and eliminates fixed configurations and costly/invasive modalities.
Detecting Primary Progressive Aphasia (PPA) from Text: A Benchmarking Study (2026.findings-eacl)

Copied to clipboard

Challenge: Primary progressive aphasia (PPA) is a neurodegenerative disorder characterized by progressive language deficits as the primary symptom.
Approach: They benchmarked the performance of traditional machine learning models with various feature extraction techniques, transformer-based models, and large language models (LLMs) they found that transformer-Based models exceeded chance-level performance in terms of balanced accuracy, while MLP using MentalBert’s embeddings achieved the highest accuracy.
Outcome: The proposed models outperform chance-level models in terms of balanced accuracy while using MentalBert’s embeddings achieve the highest accuracy.
Multi-View Attention Multiple-Instance Learning Enhanced by LLM Reasoning for Cognitive Distortion Detection (2026.acl-long)

Copied to clipboard

Challenge: Cognitive distortions are systematic errors in thinking that occur when individuals perceive and interpret external information, leading to a negative conclusion that does not correspond to reality.
Approach: They propose a framework that combines Large Language Models with a Multiple-Instance Learning architecture to enhance interpretability and expression-level reasoning.
Outcome: The proposed framework improves interpretability and expression-level reasoning on Korean and English datasets.
Multilingual prediction of Alzheimer’s disease through domain adaptation and concept-based language modelling (N19-1)

Copied to clipboard

Challenge: Existing work on speech and language models has been limited by the size of available datasets.
Approach: They propose to augment a small French dataset with a much larger English dataset to augment the language model to model the order in which information units are produced by dementia patients and controls.
Outcome: The proposed model improves classification performance in English and French separately.
Augmenting word2vec with latent Dirichlet allocation within a clinical application (N19-1)

Copied to clipboard

Challenge: Existing models that combine latent Dirichlet allocation and word embedding for distinguishing between speakers with and without Alzheimer’s disease from transcripts of picture descriptions are not suitable for clinical binary text classification tasks.
Approach: They propose three models that combine latent Dirichlet allocation and word embedding for distinguishing between speakers with and without Alzheimer’s disease from transcripts of picture descriptions.
Outcome: The proposed models outperform word2vec and LDA models on a clinical binary text classification task.
CDA: A Contrastive Data Augmentation Method for Alzheimer’s Disease Detection (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods for detecting AD are challenging and time-consuming due to lack of data and generalizability of the models.
Approach: They propose a contrastive data augmentation method which simulates the cognitive impairment of a patient by randomly deleting a proportion of text from the transcript to create negative samples.
Outcome: The proposed method achieves the best performance among language-based models on the benchmark ADReSS Challenge dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations