Challenge: Recent natural language processing systems use large language models as the backbone . however, societal biases are encoded in these models and transferred to downstream applications .
Approach: They propose to use two categories to measure fairness in natural language processing tasks . they find intrinsic and extrinsic metrics do not correlate in their original setting .
Outcome: The proposed metrics do not correlate in their original setting, the authors show . they find that they are not accurate when correcting for metric misalignments and noise .

Similar Papers

Measuring Fairness with Biased Rulers: A Comparative Study on Bias Metrics for Pre-trained Language Models (2022.naacl-main)

Copied to clipboard

Challenge: An increasing awareness of biased patterns in natural language processing resources such as BERT has motivated many metrics to quantify ‘bias’ and ‘fairness’.
Approach: They combine literature survey, correlation analysis and empirical evaluations to evaluate compatibility of fairness metrics for pre-trained language models and their downstream tasks.
Outcome: The proposed measures are not compatible with each other and highly depend on (i) templates, (ii) attribute and target seeds and (iv) the choice of embeddings.
Intrinsic Bias Metrics Do Not Correlate with Application Bias (2021.acl-long)

Copied to clipboard

Challenge: a recent survey of bias in natural language processing found that a coreference system makes more errors in an anti-stereotypical coreferent than in a pro-sterereotype one.
Approach: They compare intrinsic and extrinsic bias metrics across hundreds of trained models . they urge researchers to focus on extrindic measures of bias, not easy to measure .
Outcome: a new intrinsic metric and an annotated test set on gender bias in hate speech are tested . authors urge researchers to focus on extrinsic measures of bias, and to make them more feasible .
Quantifying Social Biases in NLP: A Generalization and Empirical Comparison of Extrinsic Fairness Metrics (2021.tacl-1)

Copied to clipboard

Challenge: Existing fairness metrics quantify the differences in a model’s behaviour across a range of demographic groups.
Approach: They propose to unify existing fairness metrics and compare them to three generalized fairness measures to reveal the connections between them.
Outcome: The proposed measures can be explained by differences in parameter choices, and the results are consistent with previous studies.
Bias and Fairness in Natural Language Processing (D19-2)

Copied to clipboard

Challenge: a tutorial will review the history of bias and fairness studies in machine learning and language processing .
Approach: This tutorial reviews the history of bias and fairness studies in machine learning and language processing . it presents recent community effort to quantify and mitigat bias in natural language processing models .
Outcome: This tutorial reviews the history of bias and fairness studies in machine learning and language processing . it aims to quantify and mitigate bias in natural language processing models for a wide spectrum of tasks .
How Gender Debiasing Affects Internal Model Representations, and Why It Matters (2022.naacl-main)

Copied to clipboard

Challenge: Existing studies of gender bias in NLP focus on extrinsic or intrinsic bias, but the relationship between extrindic and intrinsic bias is relatively unknown.
Approach: They propose a framework to measure extrinsic and intrinsic bias together and propose metric to measure debiasing and intrinsic debiases.
Outcome: The proposed framework provides a comprehensive perspective on bias in NLP models, which can be applied to deploy NLP systems in a more informed manner.
How Far Can It Go? On Intrinsic Gender Bias Mitigation for Text Classification (2023.eacl-main)

Copied to clipboard

Challenge: a growing interest in exploring how gender bias pertains in contextualized language models has been generated . intrinsic mitigation strategies and bias metrics have been proposed to mitigate gender bias in contextualised language models .
Approach: They propose to use different intrinsic bias mitigation strategies to mitigate gender bias in contextualized language models.
Outcome: The proposed probe shows that some mitigation techniques can hide gender bias . the probe also shows that not all mitigation techniques fool extrinsic bias despite their use .
On Measures of Biases and Harms in NLP (2022.findings-aacl)

Copied to clipboard

Challenge: Recent studies show that natural language processing (NLP) technologies propagate societal biases about demographic groups associated with attributes such as gender, race, and nationality.
Approach: They propose a framework for harms and questions to help practitioners understand biases . they propose measurable measures to detect and mitigate biased groups .
Outcome: The proposed framework provides a framework for harms and questions for practitioners to answer to guide the development of bias measures.
Biases Propagate in Encoder-based Vision-Language Models: A Systematic Analysis From Intrinsic Measures to Zero-shot Retrieval Outcomes (2025.findings-acl)

Copied to clipboard

Challenge: Existing encoder-based vision-language models (VLMs) contain intrinsic biases that manifest in biased outputs.
Approach: They propose a framework to measure intrinsic bias propagation by correlating intrinsic bias with extrinsic bias in zero-shot text-to-image and image-totext retrieval.
Outcome: The proposed framework shows that larger/better-performing models exhibit greater bias propagation, raising concerns given the trend towards increasingly complex AI models.
Why Don’t Prompt-Based Fairness Metrics Correlate? (2024.acl-long)

Copied to clipboard

Challenge: Existing methods to assess fairness using prompts have low correlations between fairness metrics.
Approach: They propose a method to enhance the correlation between fairness metrics by using pre-trained language models.
Outcome: The proposed method improves the correlation between fairness metrics by using pre-trained language models.
Quantifying Metric and Model Agreement in Bias Evaluation of Large Language Models (2026.acl-long)

Copied to clipboard

Challenge: a systematic way to measure agreement across bias metrics and models is lacking . a lack of agreement between metrics and model results may be a problem .
Approach: They introduce Metric Agreement Score and Model Agreement Score to measure agreement across bias metrics and models.
Outcome: The proposed measures show that metrics within the same category behave independently of each other.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations