VoiceBBQ: Investigating Effect of Content and Acoustics in Social Bias of Spoken Language Model (2025.emnlp-main)
Copied to clipboard
| Challenge: | Due to the nature of speech modality, social bias in Spoken Language Models (SLMs) can emerge from two distinct sources: 1) content aspect and 2) acoustic aspect. |
| Approach: | They propose a dataset that measures social bias by presenting ambiguous or disambiguated contexts followed by questions that may elicit stereotypical responses. |
| Outcome: | The proposed dataset converts every BBQ context into controlled voice conditions, enabling per-axis accuracy, bias, and consistency scores comparable to the original text benchmark. |
Similar Papers
PakBBQ: A Culturally Adapted Bias Benchmark for QA (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are widely adopted in language processing applications, but they often perpetuate harmful societal biases. |
| Approach: | They propose a culturally and regionally adapted extension of the original Bias Benchmark for Question Answering dataset to address this gap. |
| Outcome: | The proposed model gains 12% accuracy with disambiguation and stronger counter bias behaviors in Urdu than in English. |
BBQ: A hand-built bias benchmark for question answering (2022.findings-acl)
Copied to clipboard
Alicia Parrish, Angelica Chen, Nikita Nangia, Vishakh Padmakumar, Jason Phang, Jana Thompson, Phu Mon Htut, Samuel Bowman
| Challenge: | NLP models learn social biases, but little work has been done on how these biase manifest in outputs for applied tasks like question answering (QA). |
| Approach: | They propose a dataset that highlights attested social biases against people belonging to protected classes along nine social dimensions relevant for U.S. English-speaking contexts. |
| Outcome: | The proposed dataset highlights attested social biases against people belonging to protected classes along nine social dimensions relevant for U.S. English-speaking contexts. |
BasqBBQ: A QA Benchmark for Assessing Social Biases in LLMs for Basque, a Low-Resource Language (2025.coling-main)
Copied to clipboard
| Challenge: | Existing pre-trained language models can propagate social biases in under-resourced languages like Basque. |
| Approach: | They propose a benchmark to assess biases in Basque using a multiple-choice question-answering task. |
| Outcome: | The proposed dataset is the first to assess biases in Basque across eight domains . larger models achieve better accuracy, but ambiguous cases remain challenging . |
Bias in the Ear of the Listener: Assessing Sensitivity in Audio Language Models Across Linguistic, Demographic, and Positional Variations (2026.findings-eacl)
Copied to clipboard
| Challenge: | Recent advances extend language understanding beyond text to speech, enabling unified reasoning across modalities. |
| Approach: | They construct and release a speech-augmented benchmark based on Global MMLU Lite and a data set spanning English, Chinese, and Korean. |
| Outcome: | The proposed model is robust to demographic factors but sensitive to language and option order, suggesting that speech can amplify structural biases. |
LLMs Are Biased Towards Output Formats! Systematically Evaluating and Mitigating Output Format Bias of LLMs (2025.naacl-long)
Copied to clipboard
Do Xuan Long, Ngoc-Hai Nguyen, Tiviatis Sim, Hieu Dao, Shafiq Joty, Kenji Kawaguchi, Nancy F. Chen, Min-Yen Kan
| Challenge: | Using format-following capabilities, state-of-the-art large language models (LLMs) can be leveraged to tailor outputs to specific task formats. |
| Approach: | They propose to define a format bias evaluation metric and establish effective strategies to reduce it. |
| Outcome: | The proposed evaluation reduces the variance in ChatGPT’s performance among wrapping formats from 235.33 to 0.71 (%2) |
Artie Bias Corpus: An Open Dataset for Detecting Demographic Bias in Speech Applications (2020.lrec-1)
Copied to clipboard
| Challenge: | A speech technology exhibits demographic bias when performance is worse for one demographic group relative to another. |
| Approach: | They create an English dataset of expert-validated audio, transcript> pairs with demographic tags for age, gender, accent and open software which may be used to detect demographic bias in Automatic Speech Recognition systems. |
| Outcome: | The Artie Bias Corpus is a curated subset of the Mozilla Common Voice corpus, which is released under a Creative Commons CC0 license . |
RedditBias: A Real-World Resource for Bias Evaluation and Debiasing of Conversational Language Models (2021.acl-long)
Copied to clipboard
| Challenge: | Recent work has focused on measuring and mitigating bias in pretrained language models. |
| Approach: | They propose a dataset that measures and mitigates bias across gender,race, religion, and queerness . they compare REDDITBIAS to a widely used conversational DialoGPT model . |
| Outcome: | The proposed framework measures and mitigates bias across gender,race, religion, and queerness dimensions. |
Evaluating Open-Source ASR Systems: Performance Across Diverse Audio Conditions and Error Correction Methods (2025.coling-main)
Copied to clipboard
| Challenge: | Automated speech recognition (ASR) systems are able to transcribe spontaneous human conversations with high accuracy. |
| Approach: | They evaluate the accuracy of open source automatic speech recognition systems across conversational speech datasets and explore the potential of ASR ensembling and post-ASR correction methods to improve transcription accuracy. |
| Outcome: | The proposed methods highlight the need for robust error correction techniques and address demographic biases to enhance ASR performance and inclusivity. |
VIGNETTE: Socially Grounded Bias Evaluation for Vision-Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Existing studies on VLM bias focus on portrait-style images and gender-occupation associations . existing studies ignore broader and more complex social stereotypes and their implied harm . |
| Approach: | They propose a large-scale VQA benchmark for evaluating bias in vision-language models . they use a question-answering framework that spans factuality, perception, stereotyping, and decision making . |
| Outcome: | The proposed framework examines bias in vision-language models using 30M+ images . findings reveal subtle, multifaceted, and surprising stereotypical patterns . |
CBBQ: A Chinese Bias Benchmark Dataset Curated with Human-AI Collaboration for Large Language Models (2024.lrec-main)
Copied to clipboard
| Challenge: | a dataset of Chinese large language models is used to measure societal biases . many studies have shown that LLMs exhibit harmful societal biased outputs despite human data . |
| Approach: | They present a Chinese Bias Benchmark dataset that includes over 100K questions constructed by human experts and generative language models. |
| Outcome: | The proposed dataset covers stereotypes and societal biases in 14 social dimensions related to Chinese culture and values. |