Challenge: Recent advances in large language models have also introduced additional safety risks and raised concerns regarding their detrimental impact on already marginalized populations.
Approach: They propose to use LLMs to evaluate their safety responses on already mitigated biases by evaluating models on already encoded assumptions.
Outcome: The proposed model can encode harmful assumptions, but it can also be harmful for certain demographic groups.

Similar Papers

Exploring Safety-Utility Trade-Offs in Personalized Language Models (2025.naacl-long)

Copied to clipboard

Challenge: Prior studies have shown that large language models can exhibit bias against specific demographic groups and engage in the generation of stereotypical responses.
Approach: They propose a framework to evaluate LLM performance along two axes: safety and utility.
Outcome: The proposed framework evaluates the performance of LLMs along two axes: safety and utility.
A Chinese Dataset for Evaluating the Safeguards in Large Language Models (2024.findings-acl)

Copied to clipboard

Challenge: a recent study has shown that large language models can produce harmful responses, exposing users to unexpected risks.
Approach: They propose a dataset for the safety evaluation of Chinese LLMs in Mandarin Chinese . they extend the dataset to better identify false negative and false positive examples .
Outcome: The proposed dataset is for the safety evaluation of Chinese LLMs, and is based on a Chinese dataset.
Multitask-Bench: Unveiling and Mitigating Safety Gaps in LLMs Fine-tuning (2025.coling-main)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have led to their adoption across a wide range of tasks, ranging from code generation to machine translation and sentiment analysis.
Approach: They propose to fine-tune LLMs on benign (non-harmful) data to ensure safe outputs.
Outcome: The proposed model reduces attack success rates across a range of tasks without compromising its usefulness.
Safety of Large Language Models Beyond English: A Systematic Literature Review of Risks, Biases, and Safeguards (2026.eacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have a growing number of applications that generate harmful, biased, or unsafe content.
Approach: They synthesize findings from recent studies that evaluate their robustness across languages . they highlight gaps in multilingual safety research and recommend future work .
Outcome: The systematic review examines the multilingual safety of large language models in English . it identifies challenges such as dataset availability and evaluation biases .
Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety (2025.emnlp-main)

Copied to clipboard

Challenge: Existing surveys focus on interpretation or safety, but safety and understanding are core motivations for interpretation research.
Approach: They propose a framework that connects interpretation methods, enhancements they inform, and tools that operationalize them.
Outcome: The proposed framework summarizes nearly 70 studies at their intersections and concludes with open challenges and future directions.
Do-Not-Answer: Evaluating Safeguards in LLMs (2024.findings-eacl)

Copied to clipboard

Challenge: a dataset evaluating harmful capabilities in large language models is available at https://github.com/Libr-AI/do-not-answer.
Approach: They collect an open-source dataset to evaluate the safeguards in large language models . they find that simple BERT-style classifiers can achieve results comparable to GPT-4 .
Outcome: The proposed dataset compares the safety of six popular LLMs to GPT-4 on automatic safety evaluation.
SLM as Guardian: Pioneering AI Safety with Small Language Model (2024.emnlp-industry)

Copied to clipboard

Challenge: Prior safety research on large language models focused on aligning them to safety requirements, but internalizing such safeguard features into larger models brought challenges of higher training cost and unintended degradation of helpfulness.
Approach: They propose a multi-task learning mechanism that integrates harmful query detection and safeguard response into a single model.
Outcome: The proposed approach outperforms the publicly available LLMs in harmful query detection and safeguard response generation.
The Language Barrier: Dissecting Safety Challenges of LLMs in Multilingual Contexts (2024.findings-acl)

Copied to clipboard

Challenge: Recent studies show that malicious prompt instructions could solicit objectionable content from LLMs.
Approach: They compare how state-of-the-art LLMs respond to malicious prompts in different languages . they find that LLM's generate unsafe responses more often when a prompt is written in a lower-resource language .
Outcome: The proposed model can generate unsafe responses more often when a malicious prompt is written in a lower-resource language, and less irrelevant responses when written in lower-source languages.
Characterizing Selective Refusal Bias in Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: a recent study shows that safety guardrails in large language models can inadvertently introduce or reflect new biases as they may refuse to generate harmful content targeting some demographic groups and not others.
Approach: They examine the selective refusal bias in large language models by examining demographics and responses.
Outcome: The proposed model fails to defend against an indirect attack on previously refused groups in 89% of the trials.
Auditing LLM Responses to Harmful Stereotypes Targeting Mental Health Groups (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) can exhibit imbalanced biases against vulnerable groups, but how they rationalize stereotypes and rights restrictions targeting mental health entities remains underexplored.
Approach: They audit a suite of open-weight LLMs on stereotype-justification prompts tied to mental health identities.
Outcome: The proposed models endorse harmful stereotypes when explicitly asked to justify them, with endorsement varying across model families, versions, and mental health conditions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations