Challenge: Pre-trained neural language models fine-tuned on AD transcripts perform well, but little research has explored the effects of the gender of the speakers represented by these transcripts.
Approach: They propose to use the Extended Confounding Filter and the Dual Filter to isolate and ablate weights associated with gender in dementia datasets.
Outcome: The proposed methods overfit to training data distributions and disrupt gender-related weights, with the trade-off of slightly reduced dementia detection performance.

Similar Papers

GPT-D: Inducing Dementia-related Linguistic Anomalies by Deliberate Degradation of Artificial Neural Language Models (2022.acl-long)

Copied to clipboard

Challenge: Existing methods for fine-tuning large numbers of model parameters have shown impressive performance on the task of discriminating between language produced by cognitively healthy individuals and those with Alzheimer’s disease (AD).
Approach: They propose to use a Transformer DL model pre-trained on general English text to combine an artificially degraded version of itself with a model that generalizes well to spontaneous conversations.
Outcome: The proposed method generalizes well to spontaneous conversations and generates text with characteristics associated with AD, demonstrating the induction of dementia-related linguistic anomalies.
Enriching Neural Models with Targeted Features for Dementia Detection (P19-2)

Copied to clipboard

Challenge: In the United States, adults over 65 are expected to comprise one-fifth of the population by 2030, and a larger proportion of the . population than those under 18 by 2035.
Approach: They propose a neural model that takes into account both long language samples and hand-crafted linguistic features to distinguish between dementia affected and healthy patients.
Outcome: The proposed model achieves an F1 score of 0.929 on the DementiaBank dataset and the state-of-the-art on the dataset.
Augmenting word2vec with latent Dirichlet allocation within a clinical application (N19-1)

Copied to clipboard

Challenge: Existing models that combine latent Dirichlet allocation and word embedding for distinguishing between speakers with and without Alzheimer’s disease from transcripts of picture descriptions are not suitable for clinical binary text classification tasks.
Approach: They propose three models that combine latent Dirichlet allocation and word embedding for distinguishing between speakers with and without Alzheimer’s disease from transcripts of picture descriptions.
Outcome: The proposed models outperform word2vec and LDA models on a clinical binary text classification task.
Detecting Linguistic Characteristics of Alzheimer’s Dementia by Interpreting Neural Models (N18-2)

Copied to clipboard

Challenge: Current diagnoses often involve lengthy medical evaluations.
Approach: They apply neural models based on CNNs, LSTM-RNNs, and their combination to classify AD and control language samples.
Outcome: The proposed model achieves independent benchmark accuracy for the AD classification task.
An LLM-based Temporal-spatial Data Generation and Fusion Approach for Early Detection of Late Onset Alzheimer’s Disease (LOAD) Stagings Especially in Chinese and English-speaking Populations (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing approaches struggle with temporal-spatial challenges in capturing subtle linguistic shifts across different disease stages.
Approach: They propose a large language model-driven T-S fusion framework that integrates multilingual LLMs, contrastive learning and interpretable marker discovery to revolutionize late onset AD detection.
Outcome: The proposed framework achieves state-of-the-art performance in late onset AD detection while enabling cross-linguistic diagnostics.
A Tale of Two Perplexities: Sensitivity of Neural Language Models to Lexical Retrieval Deficits in Dementia of the Alzheimer’s Type (2020.acl-main)

Copied to clipboard

Challenge: Recent studies show that cognitive manifestations of future dementia may appear as early as 18 years prior to clinical diagnosis . lack of clear diagnosis and prognosis, possibly for an Alzheimer's type, is a major limitation of current methods for identifying dementia-specific cognitive markers.
Approach: They propose to interrogate neural LMs trained on participants with and without dementia by manipulating lexical frequency.
Outcome: The proposed model improves upon the current state-of-the-art for models trained on transcripts of speech produced by healthy participants and those with dementia.
Using Artificial French Data to Understand the Emergence of Gender Bias in Transformer Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies have demonstrated the ability of neural language models to learn linguistic properties without direct supervision.
Approach: They propose to use an artificial corpus generated by a PCFG to control the gender distribution in training data and determine under which conditions a model correctly captures gender information.
Outcome: The proposed approach allows to control the gender distribution in training data and determine under which conditions a model correctly captures gender information or appears gender-biased.
How to Generalize the Detection of AI-Generated Text: Confounding Neurons (2025.findings-emnlp)

Copied to clipboard

Challenge: Linguistic and domain confounders introduce spurious correlations, leading to poor out-of-distribution (OOD) performance.
Approach: They propose a novel post-hoc, neuron-level intervention framework to disentangle AI-generated text detection factors from data-specific biases.
Outcome: The proposed framework reduces topic-specific biases by encoding individual neurons within transformers-based detectors rather than task-specific signals.
CDA: A Contrastive Data Augmentation Method for Alzheimer’s Disease Detection (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods for detecting AD are challenging and time-consuming due to lack of data and generalizability of the models.
Approach: They propose a contrastive data augmentation method which simulates the cognitive impairment of a patient by randomly deleting a proportion of text from the transcript to create negative samples.
Outcome: The proposed method achieves the best performance among language-based models on the benchmark ADReSS Challenge dataset.
Mitigating Gender Bias Amplification in Distribution by Posterior Regularization (2020.acl-main)

Copied to clipboard

Challenge: Recent studies show that data-driven machine learning models carry societal biases in the dataset they trained on.
Approach: They propose to calibrate top predictions of a model by injecting corpus-level constraints to ensure that the gender disparity is not amplified.
Outcome: The proposed method can almost remove bias amplification in the distribution with little loss of performance.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations