Challenge: Past studies have shown biases in natural language generation systems but there has been little work on evaluating the bias evaluation approaches.
Approach: They propose a method for evaluating biases in natural language generation systems by paraphrasing syntactic prompts with different syntaktic structures and paraphrazing them to evaluate demographic bias.
Outcome: The proposed method is more robust and shows that some syntactic structures prompt more toxic content while others could prompt less biased generation.

Similar Papers

Adapting Bias Evaluation to Domain Contexts using Generative Models (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to assess social bias in NLP systems face limitations in scalability and fidelity across domains.
Approach: They propose a domain-adaptive framework that uses prompting with Large Language Models to automatically transform template-based bias datasets into domain-specific variants.
Outcome: The proposed framework improves the accuracy and contextual relevance of bias evaluations in socially relevant datasets.
Towards Controllable Biases in Language Generation (2020.findings-emnlp)

Copied to clipboard

Challenge: a new method to induce societal biases in natural language generation is being developed . a method to equalize the amount of biased text across demographics is effective .
Approach: They propose a method to induce societal biases in natural language generation by using demographic inequalities.
Outcome: The proposed method is effective at equalizing biases across demographics while generating less negatively biased text overall.
Using Natural Sentence Prompts for Understanding Biases in Language Models (2022.naacl-main)

Copied to clipboard

Challenge: Recent work has shown that language models are susceptible to biases present in the training dataset.
Approach: They propose to use natural sentence prompts to analyze gender-occupation biases in language models.
Outcome: The proposed dataset can be used to analyze gender-occupation biases in language models.
This prompt is measuring <mask>: evaluating bias evaluation in language models (2023.findings-acl)

Copied to clipboard

Challenge: a growing body of work uses prompts and templates to assess bias in language models . authors examine the scope of possible bias types and identify those under-researched .
Approach: They draw on a measurement modelling framework to create a bias taxonomy . they show that bias tests are often unstated or ambiguous, carry implicit assumptions .
Outcome: The proposed taxonomy shows that bias tests are often unstated or ambiguous . the analysis illuminates the scope of possible bias types the field can measure .
Take Care of Your Prompt Bias! Investigating and Mitigating Prompt Bias in Factual Knowledge Extraction (2024.lrec-main)

Copied to clipboard

Challenge: Recent research shows that pre-trained language models suffer from “prompt bias” in factual knowledge extraction.
Approach: They propose a representation-based approach to mitigate prompt bias during inference time by querying the model and removing it from its internal representations to generate debiased representations.
Outcome: The proposed approach corrects the overfitted performance caused by prompt bias and significantly improves prompt retrieval capability.
Beyond One-Size-Fits-All: Inversion Learning for Highly Effective NLG Evaluation Prompts (2026.tacl-1)

Copied to clipboard

Challenge: Evaluating natural language generation systems is challenging due to the diversity of valid outputs.
Approach: They propose an inversion learning method that learns effective reverse mappings from model outputs back to their input instructions.
Outcome: The proposed method requires only a single evaluation sample and eliminates manual prompt engineering.
Social Bias Evaluation for Large Language Models Requires Prompt Variations (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent studies have tried to evaluate and mitigate social biases accurately using limited prompts.
Approach: They investigate the sensitivity of Large Language Models when changing prompt variations . they found that LLM rankings fluctuate across prompts for both task performance and social bias .
Outcome: The results show that LLM rankings fluctuate when changing prompt variations .
Evaluating the Evaluation of Diversity in Natural Language Generation (2021.eacl-main)

Copied to clipboard

Challenge: Existing methods for controlling diversity by tuning a “decoding parameter” affect form but not meaning.
Approach: They propose a framework that measures correlation between a diversity metric and a parameter that controls some aspect of diversity in generated text.
Outcome: The proposed framework outperforms existing methods in estimating diversity . it shows that humans outperformed existing methods but affect form but not meaning .
Quantifying Bias from Decoding Techniques in Natural Language Generation (2022.coling-1)

Copied to clipboard

Challenge: Natural language generation (NLG) models can propagate social bias towards particular demography.
Approach: They propose to examine whether bias metrics like toxicity and sentiment are impacted by decoding techniques that use stochastic decoding.
Outcome: The proposed methods reveal the imperative of testing inference time bias and provide evidence on the usefulness of inspecting the entire decoding spectrum.
ROBBIE: Robust Bias Evaluation of Large Generative Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: generative large language models (LLMs) are becoming more performant and prevalent . we need tools to measure and improve their fairness, authors say .
Approach: They propose to compare 6 different prompt-based bias and toxicity metrics across 12 demographic axes and 5 families of generative large language models.
Outcome: The proposed model can be tested on more datasets to better characterize and mitigate biases . the study compared 6 prompt-based bias and toxicity metrics across 12 demographic axes and 5 families of generative large language models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations