Characterizing Selective Refusal Bias in Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: a recent study shows that safety guardrails in large language models can inadvertently introduce or reflect new biases as they may refuse to generate harmful content targeting some demographic groups and not others.
Approach: They examine the selective refusal bias in large language models by examining demographics and responses.
Outcome: The proposed model fails to defend against an indirect attack on previously refused groups in 89% of the trials.

Similar Papers

Understanding Large Language Model Vulnerabilities to Social Bias Attacks (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable linguistic capabilities across tasks . however, there is a growing concern about their potential to perpetuate social biases .
Approach: They evaluate LLMs across gender, racial, and religious bias types . they also explore cross-bias and multiple-biases attacks .
Outcome: The proposed models are more susceptible to gender bias attacks than racial or religious biases.
Safety Is Not Universal: The Selective Safety Trap in LLM Alignment (2026.findings-acl)

Copied to clipboard

Challenge: Existing safety evaluations of large language models aggregate harms under generic categories such as "Identity Hate" a bilingual benchmark identifies a selective safety trap, where defense rates vary by up to 42% within the same model solely based on the target group.
Approach: They propose a bilingual adversarial benchmark to audit selective safety in large language models . defense rates vary by up to 42% within the same model solely based on target group .
Outcome: The proposed benchmark identifies a selective safety trap in large language models . defense rates vary by up to 42% within the same model solely based on the target group.
Safety of Large Language Models Beyond English: A Systematic Literature Review of Risks, Biases, and Safeguards (2026.eacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have a growing number of applications that generate harmful, biased, or unsafe content.
Approach: They synthesize findings from recent studies that evaluate their robustness across languages . they highlight gaps in multilingual safety research and recommend future work .
Outcome: The systematic review examines the multilingual safety of large language models in English . it identifies challenges such as dataset availability and evaluation biases .
A Chinese Dataset for Evaluating the Safeguards in Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: a recent study has shown that large language models can produce harmful responses, exposing users to unexpected risks.
Approach: They propose a dataset for the safety evaluation of Chinese LLMs in Mandarin Chinese . they extend the dataset to better identify false negative and false positive examples .
Outcome: The proposed dataset is for the safety evaluation of Chinese LLMs, and is based on a Chinese dataset.
From Representational Harms to Quality-of-Service Harms: A Case Study on Llama 2 Safety Safeguards (2024.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models have also introduced additional safety risks and raised concerns regarding their detrimental impact on already marginalized populations.
Approach: They propose to use LLMs to evaluate their safety responses on already mitigated biases by evaluating models on already encoded assumptions.
Outcome: The proposed model can encode harmful assumptions, but it can also be harmful for certain demographic groups.
Let Them Down Easy! Contextual Effects of LLM Guardrails on User Perceptions and Preferences (2025.findings-emnlp)

Copied to clipboard

Challenge: Current LLMs are trained to refuse potentially harmful input queries regardless of intent . a study of 480 participants evaluating 3,840 query-response pairs reveals that response strategy largely shapes user experience .
Approach: They examine how different refusal strategies affect user perceptions across varying motivations . they find partial compliance reduces negative user perception by over 50% to flat-out refusals a 480 participants study .
Outcome: The study examines the perceptions of LLMs on user intents and their response strategies . it shows that partial compliance reduces negative user perceptions by over 50% to flat refusals .
Exploring Safety-Utility Trade-Offs in Personalized Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Prior studies have shown that large language models can exhibit bias against specific demographic groups and engage in the generation of stereotypical responses.
Approach: They propose a framework to evaluate LLM performance along two axes: safety and utility.
Outcome: The proposed framework evaluates the performance of LLMs along two axes: safety and utility.
Do Large Language Models Reflect Demographic Pluralism in Safety? (2026.findings-eacl)

Copied to clipboard

Challenge: Existing datasets that focus on demographics and safety are narrow in their annotator pools.
Approach: They propose to decouple value framing from responses by modeling pluralism directly at the prompt level.
Outcome: Demo-SafetyBench decouples value framing from responses to model pluralism at the prompt level.
Missing the Margins: A Systematic Literature Review on the Demographic Representativeness of LLMs (2025.findings-acl)

Copied to clipboard

Challenge: 211 studies on the demographic representativeness of large language models have conflicting results . 29% of the studies report positive conclusions on the representativeness, 30% do not evaluate LLMs across multiple demographic categories or within demographic subcategories.
Approach: 211 papers review the representativeness of large language models . authors recommend more precise evaluation methods and comprehensive documentation of demographic attributes .
Outcome: 211 studies on the representativeness of large language models are reviewed . 29% of the studies report positive conclusions, but 30% fail to specify subcategories . authors recommend more precise evaluation methods and documentation of demographic attributes .
The Language Barrier: Dissecting Safety Challenges of LLMs in Multilingual Contexts (2024.findings-acl)

Copied to clipboard

Challenge: Recent studies show that malicious prompt instructions could solicit objectionable content from LLMs.
Approach: They compare how state-of-the-art LLMs respond to malicious prompts in different languages . they find that LLM's generate unsafe responses more often when a prompt is written in a lower-resource language .
Outcome: The proposed model can generate unsafe responses more often when a malicious prompt is written in a lower-resource language, and less irrelevant responses when written in lower-source languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations