Challenge: Social biases manifest in language agency, but there is no comprehensive benchmark for evaluating such biase in language models.
Approach: They propose a benchmark to evaluate language agency biases in large language models . they propose 'Mitigation via Selective Rewrite' to selectively revise parts of generated texts .
Outcome: The proposed language agency bias evaluation benchmark identifies gender, racial, and intersectional biases in 3 recent LLMs.

Similar Papers

Who Gets Which Message? Auditing Demographic Bias in LLM-Generated Targeted Text (2026.findings-acl)

Copied to clipboard

Challenge: Large language models generate demographically conditioned persuasive texts at scale . authors argue that such capabilities raise questions about fairness and representational bias in automated communication.
Approach: They propose a framework for evaluating demographic-conditioned targeted messages . they find gender- and age-based asymmetries in male- and youth-targeted messages a .
Outcome: The proposed framework evaluates generated messages across three dimensions: lexical content, language style, and persuasive framing.
Towards Implicit Bias Detection and Mitigation in Multi-Agent LLM Interactions (2024.findings-emnlp)

Copied to clipboard

Challenge: a recent study shows that large language models are susceptible to societal biases due to their exposure to human-generated data.
Approach: They propose two strategies to mitigate implicit gender biases in large language models . they create scenarios where implicit gender is present and develop a metric to assess the presence of biase .
Outcome: The proposed methods mitigate implicit biases with self-reflection and fine-tuning.
Social Bias Probing: Fairness Benchmarking for Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods for evaluating social biases in language models have been limited to binary association tests on small datasets.
Approach: They propose a framework for probing language models for social biases by assessing disparate treatment . they use a large-scale benchmark to examine the diversity of identities and stereotypes .
Outcome: The proposed framework expands the analysis beyond the binary comparison of stereotypical versus anti-stereotypical identities to include a diverse range of identities and stereotypes.
ROBBIE: Robust Bias Evaluation of Large Generative Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: generative large language models (LLMs) are becoming more performant and prevalent . we need tools to measure and improve their fairness, authors say .
Approach: They propose to compare 6 different prompt-based bias and toxicity metrics across 12 demographic axes and 5 families of generative large language models.
Outcome: The proposed model can be tested on more datasets to better characterize and mitigate biases . the study compared 6 prompt-based bias and toxicity metrics across 12 demographic axes and 5 families of generative large language models.
Intersectional Stereotypes in Large Language Models: Dataset and Analysis (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies on intersectional stereotypes focus on broader, individual categories . current studies focus on single-group stereotypes, such as racial bias against African Americans .
Approach: They propose to use a dataset of intersectional stereotypes curated with the ChatGPT model to analyze propagation in three contemporary LLMs.
Outcome: The proposed dataset enables analysis of stereotype propagation in three contemporary LLMs.
Causally Testing Gender Bias in LLMs: A Case Study on Occupational Bias (2025.findings-naacl)

Copied to clipboard

Challenge: Existing studies have shown that large language models can cause harmful, human-like biases against various demographics.
Approach: They propose a causal formulation for bias measurement in generative language models based on a list of desiderata for designing robust bias benchmarks and a bias-measuring procedure to investigate occupational gender bias.
Outcome: The proposed framework is generalizable and can be extended to include other datasets.
“Kelly is a Warm Person, Joseph is a Role Model”: Gender Biases in LLM-Generated Reference Letters (2023.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) are an effective tool to assist individuals in writing documents.
Approach: They examine gender biases in large language models (LLMs)-generated reference letters . they find that models are biased because they are hallucinated .
Outcome: The proposed model-generated reference letters are evaluated on 2 popular LLMs- ChatGPT and Alpaca.
Social Bias Benchmark for Generation: A Comparison of Generation and QA-Based Evaluations (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for assessing social bias in large language models (LLMs) do not capture nuanced and context-dependent nature of natural language generation.
Approach: They propose a Bias Benchmark for Generation (BBG) that evaluates social bias in long-form generation by having LLMs generate continuations of story prompts.
Outcome: The proposed benchmark is based on the English BBQ and Korean BBQ datasets and compares it with multiplechoice BBQ evaluation.
Ask LLMs Directly, “What shapes your bias?”: Measuring Social Bias in Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: Existing methods to evaluate social bias in large language models have limitations . et al., 1995: stereotypes shape social perceptions without objective basis .
Approach: They propose a method to intuitively quantify social perceptions and suggest metrics to evaluate biases within LLMs.
Outcome: The proposed metrics capture the multi-dimensional aspects of social bias, the paper shows . they show that the proposed metrics can be used to evaluate bias in large language models .
Confronting LLMs with Traditional ML: Rethinking the Fairness of Large Language Models in Tabular Classifications (2024.naacl-long)

Copied to clipboard

Challenge: Recent studies suggest using large language models to make tabular classifications . however, LLMs have been shown to exhibit harmful social biases based on stereotypes and inequalities present in society.
Approach: They propose to use large language models to make tabular classifications . they show that LLMs inherit biases from their training data .
Outcome: The proposed models exhibit harmful biases that reflect stereotypes and inequalities in society.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations