Challenge: Due to the nature of speech modality, social bias in Spoken Language Models (SLMs) can emerge from two distinct sources: 1) content aspect and 2) acoustic aspect.
Approach: They propose a dataset that measures social bias by presenting ambiguous or disambiguated contexts followed by questions that may elicit stereotypical responses.
Outcome: The proposed dataset converts every BBQ context into controlled voice conditions, enabling per-axis accuracy, bias, and consistency scores comparable to the original text benchmark.

Similar Papers

PakBBQ: A Culturally Adapted Bias Benchmark for QA (2025.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are widely adopted in language processing applications, but they often perpetuate harmful societal biases.
Approach: They propose a culturally and regionally adapted extension of the original Bias Benchmark for Question Answering dataset to address this gap.
Outcome: The proposed model gains 12% accuracy with disambiguation and stronger counter bias behaviors in Urdu than in English.
BBQ: A hand-built bias benchmark for question answering (2022.findings-acl)

Copied to clipboard

Challenge: NLP models learn social biases, but little work has been done on how these biase manifest in outputs for applied tasks like question answering (QA).
Approach: They propose a dataset that highlights attested social biases against people belonging to protected classes along nine social dimensions relevant for U.S. English-speaking contexts.
Outcome: The proposed dataset highlights attested social biases against people belonging to protected classes along nine social dimensions relevant for U.S. English-speaking contexts.
BasqBBQ: A QA Benchmark for Assessing Social Biases in LLMs for Basque, a Low-Resource Language (2025.coling-main)

Copied to clipboard

Challenge: Existing pre-trained language models can propagate social biases in under-resourced languages like Basque.
Approach: They propose a benchmark to assess biases in Basque using a multiple-choice question-answering task.
Outcome: The proposed dataset is the first to assess biases in Basque across eight domains . larger models achieve better accuracy, but ambiguous cases remain challenging .
Bias in the Ear of the Listener: Assessing Sensitivity in Audio Language Models Across Linguistic, Demographic, and Positional Variations (2026.findings-eacl)

Copied to clipboard

Challenge: Recent advances extend language understanding beyond text to speech, enabling unified reasoning across modalities.
Approach: They construct and release a speech-augmented benchmark based on Global MMLU Lite and a data set spanning English, Chinese, and Korean.
Outcome: The proposed model is robust to demographic factors but sensitive to language and option order, suggesting that speech can amplify structural biases.
LLMs Are Biased Towards Output Formats! Systematically Evaluating and Mitigating Output Format Bias of LLMs (2025.naacl-long)

Copied to clipboard

Challenge: Using format-following capabilities, state-of-the-art large language models (LLMs) can be leveraged to tailor outputs to specific task formats.
Approach: They propose to define a format bias evaluation metric and establish effective strategies to reduce it.
Outcome: The proposed evaluation reduces the variance in ChatGPT’s performance among wrapping formats from 235.33 to 0.71 (%2)
Artie Bias Corpus: An Open Dataset for Detecting Demographic Bias in Speech Applications (2020.lrec-1)

Copied to clipboard

Challenge: A speech technology exhibits demographic bias when performance is worse for one demographic group relative to another.
Approach: They create an English dataset of expert-validated audio, transcript> pairs with demographic tags for age, gender, accent and open software which may be used to detect demographic bias in Automatic Speech Recognition systems.
Outcome: The Artie Bias Corpus is a curated subset of the Mozilla Common Voice corpus, which is released under a Creative Commons CC0 license .
RedditBias: A Real-World Resource for Bias Evaluation and Debiasing of Conversational Language Models (2021.acl-long)

Copied to clipboard

Challenge: Recent work has focused on measuring and mitigating bias in pretrained language models.
Approach: They propose a dataset that measures and mitigates bias across gender,race, religion, and queerness . they compare REDDITBIAS to a widely used conversational DialoGPT model .
Outcome: The proposed framework measures and mitigates bias across gender,race, religion, and queerness dimensions.
Evaluating Open-Source ASR Systems: Performance Across Diverse Audio Conditions and Error Correction Methods (2025.coling-main)

Copied to clipboard

Challenge: Automated speech recognition (ASR) systems are able to transcribe spontaneous human conversations with high accuracy.
Approach: They evaluate the accuracy of open source automatic speech recognition systems across conversational speech datasets and explore the potential of ASR ensembling and post-ASR correction methods to improve transcription accuracy.
Outcome: The proposed methods highlight the need for robust error correction techniques and address demographic biases to enhance ASR performance and inclusivity.
VIGNETTE: Socially Grounded Bias Evaluation for Vision-Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing studies on VLM bias focus on portrait-style images and gender-occupation associations . existing studies ignore broader and more complex social stereotypes and their implied harm .
Approach: They propose a large-scale VQA benchmark for evaluating bias in vision-language models . they use a question-answering framework that spans factuality, perception, stereotyping, and decision making .
Outcome: The proposed framework examines bias in vision-language models using 30M+ images . findings reveal subtle, multifaceted, and surprising stereotypical patterns .
CBBQ: A Chinese Bias Benchmark Dataset Curated with Human-AI Collaboration for Large Language Models (2024.lrec-main)

Copied to clipboard

Challenge: a dataset of Chinese large language models is used to measure societal biases . many studies have shown that LLMs exhibit harmful societal biased outputs despite human data .
Approach: They present a Chinese Bias Benchmark dataset that includes over 100K questions constructed by human experts and generative language models.
Outcome: The proposed dataset covers stereotypes and societal biases in 14 social dimensions related to Chinese culture and values.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations