Adversarial Removal of Demographic Attributes Revisited (D19-1)

Copied to clipboard

Challenge: Several approaches have been proposed to learn classifiers that are invariant (unbiased with respect) to protected attributes.
Approach: They propose to use a diagnostic classifier trained on a held-out subsample to find protected attributes for mention detection at above-chance levels.
Outcome: The proposed classifier generalizes poorly to new in-domain and new domains, suggesting it relies on correlations specific to their particular data sample.

Similar Papers

Adversarial Removal of Demographic Attributes from Text Data (D18-1)

Copied to clipboard

Challenge: Recent advances in Representation Learning and Adversarial Training remove unwanted features from the learned representation.
Approach: They show that demographic information of authors is encoded in the intermediate representations learned by text-based neural classifiers.
Outcome: The proposed approach achieves higher accuracies on the same dataset, the authors show . they show that the proposed approach is effective in removing unwanted features from the learned representations.
BLIND: Bias Removal With No Demographics (2023.acl-long)

Copied to clipboard

Challenge: Numerous methods to mitigate social biases require prior knowledge of the demographics in the dataset, such as gender or race.
Approach: They propose a method for bias removal without prior knowledge of demographics in the dataset.
Outcome: Experiments with racial and gender biases in sentiment classification and occupation classification tasks show that BLIND mitigates biase . BLINT is competitive with methods that require demographic information and sometimes surpasses them.
Masking Actor Information Leads to Fairer Political Claims Detection (2020.acl-main)

Copied to clipboard

Challenge: In recent years, NLP methods have found increasing adoption in the social sciences . however, CSS must be crucially interested in the algorithmic fairness of the underlying methods .
Approach: They propose two methods which mask proper names and pronouns during training of the model, thus removing personal information bias.
Outcome: The proposed methods decrease frequency bias while keeping the overall performance stable.
Measuring Mechanistic Independence: Can Bias Be Removed Without Erasing Demographics? (2026.eacl-long)

Copied to clipboard

Challenge: Using multi-task evaluation, we examine how independent demographic bias mechanisms are from general demographic recognition in language models.
Approach: They compare attribution-based and correlation-based methods for locating bias features in language models to find out which features are independent from general demographic recognition.
Outcome: The proposed method reduces bias without degrading recognition performance.
More than Minorities and Majorities: Understanding Multilateral Bias in Language Generation (2024.findings-acl)

Copied to clipboard

Challenge: Existing studies on bias dataset construction and mitigation focus on one demographic group . in real-world applications, there are more than two demographic groups at risk of the same bias.
Approach: They propose to analyze and reduce biases across multiple demographic groups using a multi-demographic bias dataset.
Outcome: The proposed method can mitigate biases among multiple demographic groups effectively, the authors show .
When and Why Does Bias Mitigation Work? (2023.findings-emnlp)

Copied to clipboard

Challenge: Neural models exploit shallow surface features to perform language understanding tasks, rather than learning the deeper language understanding and reasoning skills that practitioners desire.
Approach: They propose to use model debiasing techniques to pressure models away from spurious features and to use them to learn useful representations instead.
Outcome: The proposed methods increase models' reliance on hidden biases instead of learning robust features that help them solve a task.
When Debiasing Backfires: Counterintuitive Side Effects of Preprocessing-Based Stereotype Mitigation (2026.findings-acl)

Copied to clipboard

Challenge: Preprocessing-based methods for stereotype mitigation are widely used in NLP . preprocessing methods cause unintended shifts in attention flow, authors say .
Approach: They propose to use preprocessing-based methods to reduce stereotypes for targeted groups . they find that stereotyping or counter-stereotyping can increase for other demographics .
Outcome: The proposed methods often induce unintended shifts across demographics, the authors show . they show that such side effects are not accompanied by large changes in attention flow .
Improving Bias Mitigation through Bias Experts in Natural Language Understanding (2023.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to mitigate the detrimental effect of bias on the network include debiasing methods that down-weight the biased examples identified by an auxiliary model, which is trained with explicit bias labels.
Approach: They propose a framework that introduces binary classifiers between the auxiliary model and main model, coined bias experts, to reduce the detrimental effect of bias on the network.
Outcome: The proposed approach outperforms the state-of-the-art on various datasets while achieving high performance on in-distribution data.
Applying Intrinsic Debiasing on Downstream Tasks: Challenges and Considerations for Machine Translation (2024.emnlp-main)

Copied to clipboard

Challenge: In this study, we examine three considerations for intrinsic debiasing in neural machine translation models.
Approach: They propose to measure the extrinsic bias of neural machine translation models by embedding them in a neural embeddable space and using different tokens to debias them.
Outcome: The proposed methods over-rely on gender stereotypes and over-represent them in their models.
Modular and On-demand Bias Mitigation with Attribute-Removal Subnetworks (2023.findings-acl)

Copied to clipboard

Challenge: Existing studies show that pre-trained language models can be used to mitigate societal biases and stereotypes.
Approach: They propose a modular bias mitigation approach that integrates debiasing modules into the core model on-demand at inference time.
Outcome: The proposed approach improves on-par with baseline finetuning on gender, race, and age protected attributes on three classification tasks with gender, age, and race as protected attributes.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations