Challenge: a recent study shows that models often make less reliable or overconfident predictions for marginalized groups.
Approach: They evaluate performance and reliability disparities across demographic, regional, and legal attributes across four jurisdictions using the FairLex benchmark.
Outcome: The FairLex benchmark shows that pre-training improves performance and reliability for underrepresented groups.

Similar Papers

FairLex: A Multilingual Benchmark for Evaluating Fairness in Legal Text Processing (2022.acl-long)

Copied to clipboard

Challenge: Using pre-trained language models, we evaluate performance group disparities while none of these techniques guarantee fairness, nor consistently mitigate group disparity.
Approach: They present a benchmark suite of four datasets for evaluating the fairness of pre-trained language models and the techniques used to fine-tune them for downstream tasks.
Outcome: The proposed methods show that performance group disparities are vibrant in many cases, while none of these techniques guarantee fairness, nor consistently mitigate group disparity.
Benchmarking Intersectional Biases in NLP (2022.naacl-main)

Copied to clipboard

Challenge: Recent work on fairness of machine learning models has focused on how to debias, but research on the fairness and performance of biased/debiased models on downstream prediction tasks has been limited.
Approach: They assess intersectional bias - fairness across multiple demographic dimensions . they highlight possible causes and make recommendations for future NLP debiasing research.
Outcome: The proposed approaches fare well in terms of fairness-accuracy trade-off, but are unable to effectively alleviate bias in downstream tasks.
Quantifying Social Biases in NLP: A Generalization and Empirical Comparison of Extrinsic Fairness Metrics (2021.tacl-1)

Copied to clipboard

Challenge: Existing fairness metrics quantify the differences in a model’s behaviour across a range of demographic groups.
Approach: They propose to unify existing fairness metrics and compare them to three generalized fairness measures to reveal the connections between them.
Outcome: The proposed measures can be explained by differences in parameter choices, and the results are consistent with previous studies.
Bias and Fairness in Natural Language Processing (D19-2)

Copied to clipboard

Challenge: a tutorial will review the history of bias and fairness studies in machine learning and language processing .
Approach: This tutorial reviews the history of bias and fairness studies in machine learning and language processing . it presents recent community effort to quantify and mitigat bias in natural language processing models .
Outcome: This tutorial reviews the history of bias and fairness studies in machine learning and language processing . it aims to quantify and mitigate bias in natural language processing models for a wide spectrum of tasks .
Fair Enough: Standardizing Evaluation and Model Selection for Fairness Research in NLP (2023.eacl-main)

Copied to clipboard

Challenge: Modern NLP systems exhibit a range of biases, which a growing literature on model debiasing attempts to correct.
Approach: They propose to clarify the current situation and plot a course for meaningful progress in fair learning by making clear inter-relations among the current gamut of methods and their relation to fairness theory.
Outcome: The proposed approach addresses the practical problem of model selection, which involves a trade-off between fairness and accuracy and has led to systemic issues in fairness research.
Reliability Testing for Natural Language Processing Systems (2021.acl-long)

Copied to clipboard

Challenge: a lack of rigorous testing and ML implicit assumption of identical training and testing distributions may result in systems that discriminate against minorities.
Approach: They argue that reliability testing is needed to address the issue of demographics . they argue that adversarial attacks can be reframed for this goal .
Outcome: The proposed framework will enable rigorous and targeted testing and aid in the enactment and enforcement of industry standards.
Re-contextualizing Fairness in NLP: The Case of India (2022.aacl-main)

Copied to clipboard

Challenge: Recent research has revealed undesirable biases in NLP data and models . however, these efforts focus of social disparities in the West and are not directly portable to other geo-cultural contexts.
Approach: They propose a framework to re-contextualize NLP fairness research for the Indian context . they build resources for fairness evaluation in the Indian and delve deeper into social stereotypes for Region and Religion .
Outcome: The proposed framework can be generalized to other geo-cultural contexts.
NLP Needs Diversity outside of ‘Diversity’ (2025.findings-emnlp)

Copied to clipboard

Challenge: a new position paper argues that diversity in NLP is concentrated on a small number of areas surrounding fairness .
Approach: a new position paper argues that diversity in NLP is disproportionately concentrated on fairness areas.
Outcome: a new position paper argues that diversity in NLP is disproportionately concentrated on fairness areas.
Fairness in Language Models Beyond English: Gaps and Challenges (2023.findings-eacl)

Copied to clipboard

Challenge: Language models are inequitable at encoding and re-presentation, but there is much to be studied and criticism for the existing research that remains to be addressed.
Approach: They propose to survey fairness in multilingual and non-English contexts . they argue that it is infeasible to achieve comprehensive coverage in terms of fairness datasets based on English .
Outcome: The proposed methods are infeasible to scale across languages and cultures, the authors argue . they argue that the current methods are too narrowly focused on specific dimensions and types of biases and cannot scale across cultures.
Fairness in Automatic Speech Recognition Isn’t a One-Size-Fits-All (2025.findings-emnlp)

Copied to clipboard

Challenge: Pre-trained speech models like Whisper exhibit inconsistent group-level performance that varies across domains.
Approach: They fine-tune a Whisper model on the Fair-Speech corpus using basic fine- tuning, demographic rebalancing, gender-swapped data augmentation and a novel contrastive learning objective.
Outcome: The proposed method achieves stable, cross-domain fairness improvements without changes to the training data distribution and with minimal accuracy trade-offs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations