Challenge: a prototypical NLP experiment trains a standard architecture on labeled English data . a recent study shows that research often goes beyond the square one setup .
Approach: They argue that the prototypical NLP experiment trains a standard architecture on labeled English data and optimizes for accuracy without accounting for other dimensions such as fairness, interpretability, or computational efficiency.
Outcome: The proposed model steers and biases the research dynamics in the NLP community, the authors argue . they show that the prototype biased recent NLP research on English data is true .

Similar Papers

Fair Enough: Standardizing Evaluation and Model Selection for Fairness Research in NLP (2023.eacl-main)

Copied to clipboard

Challenge: Modern NLP systems exhibit a range of biases, which a growing literature on model debiasing attempts to correct.
Approach: They propose to clarify the current situation and plot a course for meaningful progress in fair learning by making clear inter-relations among the current gamut of methods and their relation to fairness theory.
Outcome: The proposed approach addresses the practical problem of model selection, which involves a trade-off between fairness and accuracy and has led to systemic issues in fairness research.
Bias and Fairness in Natural Language Processing (D19-2)

Copied to clipboard

Challenge: a tutorial will review the history of bias and fairness studies in machine learning and language processing .
Approach: This tutorial reviews the history of bias and fairness studies in machine learning and language processing . it presents recent community effort to quantify and mitigat bias in natural language processing models .
Outcome: This tutorial reviews the history of bias and fairness studies in machine learning and language processing . it aims to quantify and mitigate bias in natural language processing models for a wide spectrum of tasks .
Language (Technology) is Power: A Critical Survey of “Bias” in NLP (2020.acl-main)

Copied to clipboard

Challenge: 146 papers analyzing "bias" in NLP systems lack normative reasoning, we find . authors propose three recommendations for work analyzing “bias” in Nlp systems .
Approach: They propose three recommendations for analyzing "bias" in NLP systems . they propose to focus on what kinds of system behaviors are harmful, in what ways, to whom, and why .
Outcome: The proposed methods for measuring or mitigating “bias” are poorly matched to their motivations and do not engage critically with literature outside of NLP.
Cognitive Effects and Biases in Large Language Models (2026.eacl-tutorials)

Copied to clipboard

Challenge: This tutorial bridges psychology and NLP to clarify cognitive effects and biases in large language models.
Approach: This tutorial bridges psychology and NLP to clarify cognitive effects and biases in large language models.
Outcome: This tutorial bridges psychology and NLP to clarify cognitive effects and biases in large language models.
Re-contextualizing Fairness in NLP: The Case of India (2022.aacl-main)

Copied to clipboard

Challenge: Recent research has revealed undesirable biases in NLP data and models . however, these efforts focus of social disparities in the West and are not directly portable to other geo-cultural contexts.
Approach: They propose a framework to re-contextualize NLP fairness research for the Indian context . they build resources for fairness evaluation in the Indian and delve deeper into social stereotypes for Region and Religion .
Outcome: The proposed framework can be generalized to other geo-cultural contexts.
Should We Ban English NLP for a Year? (2022.emnlp-main)

Copied to clipboard

Challenge: aaron carroll: two thirds of NLP research is devoted to developing technology for speakers of English . carroll says this bias feeds into consumer technologies to widen existing inequality gaps . he says we need to consider more concrete measures to mitigate climate change .
Approach: a new paper argues that NLP is contributing to global inequalities through a digital language divide . a carbon tax, cap-and-trade and car-free Sundays are examples of measures to mitigate climate change .
Outcome: a new paper argues that NLP is contributing to global inequalities through a digital language divide . a carbon tax, cap-and-trade and car-free Sundays are examples of measures to mitigate climate change .
Quantifying Social Biases in NLP: A Generalization and Empirical Comparison of Extrinsic Fairness Metrics (2021.tacl-1)

Copied to clipboard

Challenge: Existing fairness metrics quantify the differences in a model’s behaviour across a range of demographic groups.
Approach: They propose to unify existing fairness metrics and compare them to three generalized fairness measures to reveal the connections between them.
Outcome: The proposed measures can be explained by differences in parameter choices, and the results are consistent with previous studies.
NLP Needs Diversity outside of ‘Diversity’ (2025.findings-emnlp)

Copied to clipboard

Challenge: a new position paper argues that diversity in NLP is concentrated on a small number of areas surrounding fairness .
Approach: a new position paper argues that diversity in NLP is disproportionately concentrated on fairness areas.
Outcome: a new position paper argues that diversity in NLP is disproportionately concentrated on fairness areas.
Not All Claims are Created Equal: Choosing the Right Statistical Approach to Assess Hypotheses (2020.acl-main)

Copied to clipboard

Challenge: Empirical research in natural language processing has adopted a narrow set of principles for assessing hypotheses . alternative approaches to assess hypothese rely on p-value computation, which suffers from several known issues.
Approach: They propose to compare different methods for assessing hypotheses . they argue that practitioners should first decide their target hypothesis before choosing a method .
Outcome: The proposed method differs from other methods, but is not widely used in NLP . the proposed method is based on a p-value computation, but has a small gap in accuracy .
Benchmarking Intersectional Biases in NLP (2022.naacl-main)

Copied to clipboard

Challenge: Recent work on fairness of machine learning models has focused on how to debias, but research on the fairness and performance of biased/debiased models on downstream prediction tasks has been limited.
Approach: They assess intersectional bias - fairness across multiple demographic dimensions . they highlight possible causes and make recommendations for future NLP debiasing research.
Outcome: The proposed approaches fare well in terms of fairness-accuracy trade-off, but are unable to effectively alleviate bias in downstream tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations