Identifying and Reducing Gender Bias in Word-Level Language Models (N19-3)

Copied to clipboard

Challenge: Existing discriminatory biases in training data can be amplified by models . text corpora exhibit socially problematic biase .
Approach: They propose a metric to measure gender bias and a regularization loss term to minimize embeddings onto an embeddable subspace that encodes gender.
Outcome: The proposed method reduces gender bias up to an optimal weight assigned to the loss term, and the model becomes unstable as the perplexity increases.

Similar Papers

Reducing Gender Bias in Word-Level Language Models with a Gender-Equalizing Loss Function (P19-2)

Copied to clipboard

Challenge: Existing methods to reduce gender bias in natural language datasets are inadequate.
Approach: They propose a loss function modification approach which equalizes the probabilities of male and female words in the output.
Outcome: The proposed approach outperforms existing methods in several aspects, especially in reducing gender bias in occupation words.
Mitigating Gender Bias in Natural Language Processing: Literature Review (P19-1)

Copied to clipboard

Challenge: NLP models propagate and may even amplify gender bias found in text corpora . methods to mitigate gender bias in NLP are relatively nascent .
Approach: They propose to analyze gender bias based on four forms of representation bias and discuss the advantages and drawbacks of existing gender debiasing methods.
Outcome: The proposed methods are based on four forms of representation bias and have advantages and drawbacks.
Examining Gender Bias in Languages with Grammatical Gender (D19-1)

Copied to clipboard

Challenge: Existing studies on gender bias in word embeddings focus on English . however, these studies cannot be extended to languages with morphological agreement on gender .
Approach: They propose new metrics to evaluate gender bias in word embeddings of English and Spanish . they extend existing approaches to mitigate gender bias while preserving original embeddables .
Outcome: The proposed methods reduce gender bias while preserving the original embeddings.
Lipstick on a Pig: Debiasing Methods Cover up Systematic Gender Biases in Word Embeddings But do not Remove Them (N19-1)

Copied to clipboard

Challenge: Existing methods to remove gender bias from word embeddings are insufficient, we argue . existing methods for gender-neutral modeling are ineffective, we conclude .
Approach: They propose methods to reduce gender bias in word embeddings by debiasing them using text corpora.
Outcome: The proposed methods show that they can reduce gender bias in word embeddings . the proposed methods are insufficient and should not be trusted, the authors argue .
Leveraging Pre-trained Language Models for Gender Debiasing (2022.lrec-1)

Copied to clipboard

Challenge: Existing methods to reduce gender bias in natural language are costly and time-consuming.
Approach: They propose a method to generate gender variants for a given text using pre-trained language models as the resource without any task-specific labelled data.
Outcome: The proposed method can reduce gender bias in a language generation context without a task-specific labelled data.
Multi-Dimensional Gender Bias Classification (2020.emnlp-main)

Copied to clipboard

Challenge: a novel framework decomposes gender bias in text along several pragmatic and semantic dimensions . language is a primary means by which people communicate, express identities and categorize themselves . unwanted gender biases can affect downstream applications, leading to poor user experiences .
Approach: They propose a framework that decomposes gender bias in text along several dimensions . they annotate eight large scale datasets with gender information and collect a benchmark .
Outcome: The proposed framework decomposes gender bias in text along several pragmatic and semantic dimensions.
Reducing Gender Bias in Abusive Language Detection (D18-1)

Copied to clipboard

Challenge: Abusive language detection models tend to be biased toward identity words of a certain group of people . recent studies have raised concerns about the robustness of such systems .
Approach: They propose to use debiased word embeddings, gender swap data augmentation to reduce model bias . they also propose to fine-tune models with a larger corpus to correct such bias if needed .
Outcome: The proposed methods reduce model bias by 90-98% and can be extended to correct model bias in other scenarios.
Gender-preserving Debiasing for Pre-trained Word Embeddings (P19-1)

Copied to clipboard

Challenge: Existing methods for debiasing word embeddings have shown discriminative biases . word embeds learnt from social media have shown to encode racist, offensive and discriminative language usage.
Approach: They propose a method that preserves gender-related information while removing stereotypical gender biases from pre-trained word embeddings.
Outcome: The proposed method preserves gender-related information while removing stereotypical discriminative gender biases from pre-trained word embeddings.
Double-Hard Debias: Tailoring Word Embeddings for Gender Bias Mitigation (2020.acl-main)

Copied to clipboard

Challenge: Existing methods to debias word embeddings from human-generated corpora inherit strong gender bias . prior work has suggested removing gender component from pre-trained word embeds or compressing gender information into a few dimensions of the embeddable space .
Approach: They propose a technique that purifies word embeddings against inferred gender subspaces . they propose to preserve distributional semantics of pre-trained word embeds while reducing gender bias .
Outcome: The proposed technique preserves distributional semantics of pre-trained word embeddings while reducing gender bias to a larger degree than prior approaches.
Neutralizing Gender Bias in Word Embeddings with Latent Disentanglement and Counterfactual Generation (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent research shows word embeddings have strong gender biases in embeddable spaces . a proposed method can be used to debiase word embeds without loss of semantic information .
Approach: They propose a latent disentanglement method with a siamese auto-encoder structure with an adapted gradient reversal layer to debiase word embeddings.
Outcome: The proposed method can preserve semantic information during debiasing while minimizing loss of semantic information for extrinsic NLP tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations