Challenge: omnipresence of large pre-trained language models has fueled concerns regarding systematic biases carried over from underlying data into the applications they are used in.
Approach: They propose to compare social biases with non-social biase masked by alternate constructions that maintain the essence of their social bias.
Outcome: The proposed benchmarks underestimate or overestimate the social bias in a given model.

Similar Papers

Quantifying Social Biases in NLP: A Generalization and Empirical Comparison of Extrinsic Fairness Metrics (2021.tacl-1)

Copied to clipboard

Challenge: Existing fairness metrics quantify the differences in a model’s behaviour across a range of demographic groups.
Approach: They propose to unify existing fairness metrics and compare them to three generalized fairness measures to reveal the connections between them.
Outcome: The proposed measures can be explained by differences in parameter choices, and the results are consistent with previous studies.
BBQ: A hand-built bias benchmark for question answering (2022.findings-acl)

Copied to clipboard

Challenge: NLP models learn social biases, but little work has been done on how these biase manifest in outputs for applied tasks like question answering (QA).
Approach: They propose a dataset that highlights attested social biases against people belonging to protected classes along nine social dimensions relevant for U.S. English-speaking contexts.
Outcome: The proposed dataset highlights attested social biases against people belonging to protected classes along nine social dimensions relevant for U.S. English-speaking contexts.
LLMs Are Biased Towards Output Formats! Systematically Evaluating and Mitigating Output Format Bias of LLMs (2025.naacl-long)

Copied to clipboard

Challenge: Using format-following capabilities, state-of-the-art large language models (LLMs) can be leveraged to tailor outputs to specific task formats.
Approach: They propose to define a format bias evaluation metric and establish effective strategies to reduce it.
Outcome: The proposed evaluation reduces the variance in ChatGPT’s performance among wrapping formats from 235.33 to 0.71 (%2)
Bias in Language Models: Beyond Trick Tests and Towards RUTEd Evaluation (2025.acl-long)

Copied to clipboard

Challenge: Standard bias benchmarks are used for large language models to measure the association between social attributes and single-word outputs.
Approach: They adapt three standard bias metrics of next-word prediction to measure gender-occupation bias and develop an analogous RUTEd evaluation in three contexts of real-world LLM use.
Outcome: The proposed benchmarks are robust to lengthening model outputs via a more realistic user prompt in the domain of gender-occupation bias.
Sense Embeddings are also Biased – Evaluating Social Biases in Static and Contextualised Sense Embeddings (2022.acl-long)

Copied to clipboard

Challenge: Existing studies have evaluated social biases in word embeddings, but they are understudied.
Approach: They propose to evaluate the social biases in sense embeddings using a benchmark dataset for word embedders.
Outcome: The proposed measures show that even when no biases are found at word-level, there are still worrying levels of social biase at sense-level which are often ignored by the word- level bias evaluation measures.
Bias and Fairness in Natural Language Processing (D19-2)

Copied to clipboard

Challenge: a tutorial will review the history of bias and fairness studies in machine learning and language processing .
Approach: This tutorial reviews the history of bias and fairness studies in machine learning and language processing . it presents recent community effort to quantify and mitigat bias in natural language processing models .
Outcome: This tutorial reviews the history of bias and fairness studies in machine learning and language processing . it aims to quantify and mitigate bias in natural language processing models for a wide spectrum of tasks .
Blind Men and the Elephant: Diverse Perspectives on Gender Stereotypes in Benchmark Datasets (2025.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks for measuring gender stereotypical bias in language models are inconsistencies . lack of explicit standards in data gathering can have detrimental effects on results .
Approach: They propose that currently available benchmarks capture only partial facets of gender stereotypes . they apply a framework from social psychology to balance data across components of gender stereotypes based on stereotypical benchmarks.
Outcome: The proposed framework improves correlation between different benchmarks by using simple balancing techniques.
Robust Bias Detection in MLMs and its Application to Human Trait Ratings (2025.findings-naacl)

Copied to clipboard

Challenge: Existing methods to assess demographic bias in MLMs ignore random variability of templates and target concepts, and neglect bias quantification.
Approach: They propose a systematic statistical approach to assess bias in MLMs using mixed models to account for random effects, pseudo-perplexity weights for sentences derived from templates and quantify bias using statistical effect sizes.
Outcome: The proposed method matches on bias scores in magnitude and direction with small to medium effect sizes.
Cognitive Effects and Biases in Large Language Models (2026.eacl-tutorials)

Copied to clipboard

Challenge: This tutorial bridges psychology and NLP to clarify cognitive effects and biases in large language models.
Approach: This tutorial bridges psychology and NLP to clarify cognitive effects and biases in large language models.
Outcome: This tutorial bridges psychology and NLP to clarify cognitive effects and biases in large language models.
Benchmarking Meta-embeddings: What Works and What Does Not (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to build meta-embeddings have been evaluated using a variety of methods and datasets, which makes it difficult to draw meaningful conclusions regarding the merits of each approach.
Approach: They propose a unified framework for a fair and objective meta-embedding evaluation using intrinsic and extrinsic tasks.
Outcome: The proposed framework outperforms existing methods on intrinsic and extrinsic evaluation benchmarks and outperformed existing methods.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations