SAGED: A Holistic Bias-Benchmarking Pipeline for Language Models with Customisable Fairness Calibration (2025.coling-main)
Copied to clipboard
Xin Guan, Nate Demchak, Saloni Gupta, Ze Wang, Ediz Ertekin Jr., Adriano Koshiyama, Emre Kazim, Zekun Wu
| Challenge: | Existing benchmarks for large language models fail to detect bias due to limited scope, contamination, and lack of a fairness baseline. |
| Approach: | They propose a benchmarking pipeline to detect biases in large language models . they use metrics for max disparity, impact ratio, and bias concentration to analyze disparity . |
| Outcome: | SAGED(bias) is the first holistic benchmarking pipeline to address biases in large language models. |
Similar Papers
ROBBIE: Robust Bias Evaluation of Large Generative Language Models (2023.emnlp-main)
Copied to clipboard
David Esiobu, Xiaoqing Tan, Saghar Hosseini, Megan Ung, Yuchen Zhang, Jude Fernandes, Jane Dwivedi-Yu, Eleonora Presani, Adina Williams, Eric Smith
| Challenge: | generative large language models (LLMs) are becoming more performant and prevalent . we need tools to measure and improve their fairness, authors say . |
| Approach: | They propose to compare 6 different prompt-based bias and toxicity metrics across 12 demographic axes and 5 families of generative large language models. |
| Outcome: | The proposed model can be tested on more datasets to better characterize and mitigate biases . the study compared 6 prompt-based bias and toxicity metrics across 12 demographic axes and 5 families of generative large language models. |
RubricBench: Aligning Model-Generated Rubrics with Human Standards (2026.acl-long)
Copied to clipboard
Junyi Zhou, Qiyuan Zhang, Yufei Wang, Fuyuan Lyu, Yidong Ming, Can Xu, Qingfeng Sun, Kai Zheng, Peng Kang, Xue Liu, Chen Ma
| Challenge: | Existing benchmarks lack discriminative complexity and ground-truth rubric annotations required for rigorous evaluation. |
| Approach: | They propose a curated benchmark with 1,147 pairwise comparisons to assess the reliability of rubric-based evaluation. |
| Outcome: | The proposed benchmarks show that they support diverse domains, exhibit discriminative ability, provide high-quality annotations, and include human-authored rubrics. |
Beyond Performance: Quantifying and Mitigating Label Bias in LLMs (2024.naacl-long)
Copied to clipboard
| Challenge: | Large language models exhibit undesirable preference toward predicting certain answers over others, despite their adaptability to diverse tasks. |
| Approach: | They propose a label bias calibration method that outperforms recent calibration approaches for improving performance and mitigating label bias. |
| Outcome: | The proposed method outperforms calibration approaches for improving performance and mitigating label bias. |
LLMs Are Biased Towards Output Formats! Systematically Evaluating and Mitigating Output Format Bias of LLMs (2025.naacl-long)
Copied to clipboard
Do Xuan Long, Ngoc-Hai Nguyen, Tiviatis Sim, Hieu Dao, Shafiq Joty, Kenji Kawaguchi, Nancy F. Chen, Min-Yen Kan
| Challenge: | Using format-following capabilities, state-of-the-art large language models (LLMs) can be leveraged to tailor outputs to specific task formats. |
| Approach: | They propose to define a format bias evaluation metric and establish effective strategies to reduce it. |
| Outcome: | The proposed evaluation reduces the variance in ChatGPT’s performance among wrapping formats from 235.33 to 0.71 (%2) |
Uncovering Stereotypes in Large Language Models: A Task Complexity-based Approach (2024.eacl-long)
Copied to clipboard
| Challenge: | Recent Large Language Models (LLMs) have unlocked unprecedented applications of AI. |
| Approach: | They propose to use a social benchmark to evaluate the bias protection provided by Large Language Models (LLMs) with a variety of tasks with varying complexities to assess their effectiveness. |
| Outcome: | The proposed benchmark shows that both ChatGPT and GPT-4 have strong biases with respect to nationality, gender, race, and religion. |
Subtle Biases Need Subtler Measures: Dual Metrics for Evaluating Representative and Affinity Bias in Large Language Models (2024.acl-long)
Copied to clipboard
| Challenge: | Representative bias is a tendency of Large Language Models to generate outputs that mirror the experiences of certain identity groups, and affinity bias is an evaluative preference for specific narratives. |
| Approach: | They propose two new metrics to measure representative bias and affinity bias within large language models and present a new set of tasks designed with customized rubrics to detect these biases. |
| Outcome: | The proposed model identifies representative biases in prominent LLMs, with a preference for identities associated with being white, straight, and men. |
“I’m sorry to hear that”: Finding New Biases in Language Models with a Holistic Descriptor Dataset (2022.emnlp-main)
Copied to clipboard
| Challenge: | Language models are increasingly important to measure all possible demographic markers of identity . many datasets for measuring bias are limited in their coverage of demographic axes . |
| Approach: | They propose a bias measurement dataset that includes nearly 600 descriptor terms across 13 demographic axes. |
| Outcome: | The proposed dataset explores, detects, and reduces biases in language models. |
Ready to Translate, Not to Represent? Bias and Performance Gaps in Multilingual LLMs Across Language Families and Domains (2026.findings-acl)
Copied to clipboard
Md. Faiyaz Abdullah Sayeedi, Subhey Sadi Rahman, Md. Mahbub Alam, Md. Adnanul Islam, Jannatul Ferdous Deepti, Tasnim Mohiuddin, Md Mofijul Islam, Swakkhar Shatabda
| Challenge: | Large Language Models (LLMs) have redefined Machine Translation, enabling context-aware and fluent translations across hundreds of languages and textual domains. |
| Approach: | They propose a framework and dataset to evaluate the translation quality and fairness of open-source LLMs. |
| Outcome: | The proposed framework and dataset evaluates translation quality and fairness of open-source LLMs. |
DiFair: A Benchmark for Disentangled Assessment of Gender Knowledge and Bias (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to mitigate gender bias in pre-trained language models are often evaluated on datasets that check the extent to which the model is gender-neutral in its predictions. |
| Approach: | They propose to use a manually curated dataset to measure gender bias and to measure useful gender knowledge. |
| Outcome: | The proposed dataset aims to quantify gender biases and to assess their impact on useful gender knowledge. |
A Scalable Entity-Based Framework for Auditing Bias in Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to bias evaluation in large language models trade ecological validity for statistical control, or use artificial prompts that lack scale and rigor. |
| Approach: | They propose a framework that uses named entities as probes to measure bias in large language models. |
| Outcome: | The proposed framework reproduces bias patterns observed in natural text, enabling large-scale analysis. |