RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models (2026.eacl-long)
Copied to clipboard
| Challenge: | Language models fail to selectively refuse to answer based on flawed context, study finds . current benchmarks fail to evaluate complex capabilities like selective refusal . |
| Approach: | They propose a framework that generates diagnostic test cases through controlled linguistic perturbation. |
| Outcome: | The proposed framework employs 176 perturbation strategies across six categories of uncertainty and three intensity levels. |
Similar Papers
Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing safety alignment benchmarks fail to evaluate Safe Completion: the model’s ability to maximise helpfulness on dual-use or borderline queries without crossing into actionable harm. |
| Approach: | They propose a large-scale benchmark to measure Over-Refusal and Safe Completion quality in healthcare. |
| Outcome: | The framework evaluates 30 state-of-the-art LLMs including GPT-5 and Claude-4. |
COVER: Context-Driven Over-Refusal Verification in LLMs (2025.findings-acl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) have become increasingly prevalent in the field of Natural Language Processing (NLP), achieving unprecedented performance across linguistic tasks. |
| Approach: | They propose a framework to quantify and analyze context-driven over-refusal . they find that over-fusals depend on the task, system prompts, model family, and the number of retrieved documents. |
| Outcome: | The proposed framework quantifyes and analyzes the concept of context-driven over-refusal on two public corpora. |
E-Bench: Towards Evaluating the Ease-of-Use of Large Language Models (2025.coling-main)
Copied to clipboard
| Challenge: | E-Bench is a framework for easy-to-use research on large language models. |
| Approach: | They propose to evaluate the ease-of-use of large language models and construct an E-Bench . they simulate human use from synonymous and typographical perturbations . |
| Outcome: | The proposed model is able to resist synonymous expressions and typos and improves performance. |
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models (2024.naacl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) are now being used by millions of people across the world. |
| Approach: | They propose a test suite called XSTest to identify such eXaggerated Safety behaviours in a systematic way. |
| Outcome: | The proposed test suite identifies eXaggerated Safety behaviours in a systematic way. |
Do not Abstain! Identify and Solve the Uncertainty (2025.acl-long)
Copied to clipboard
| Challenge: | Existing solutions rely on evasive responses when confronting uncertain scenarios. |
| Approach: | They propose a benchmark to assess LLMs' ability to recognize and address uncertainty . they generate context-aware inquiries that highlight the confusing aspect of the original query . |
| Outcome: | Experiments with ConfuseBench show that LLMs struggle to identify root cause of uncertainty and solve it. |
Refusal-Aware Red Teaming: Exposing Inconsistency in Safety Evaluations (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) require rigorous safety evaluations to be effective. |
| Approach: | They propose a red teaming framework that detects internal model refusals and contrasts them with judgments from an external safety evaluator to generate test cases that expose such discrepancies. |
| Outcome: | The proposed framework outperforms existing reinforcement learning-based approaches in generating diverse test cases and achieves a substantially higher discovery rate of refusal gaps. |
RiddleBench: A New Generative Reasoning Benchmark for LLMs (2026.findings-eacl)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) show remarkable capabilities, but complex reasoning skills require deeper investigation. |
| Approach: | They propose a benchmark of 1,737 puzzles to test reasoning beyond simple pattern matching. |
| Outcome: | The proposed model performs poorly when faced with reordered constraints or irrelevant information. |
Dynamic Evaluation for Oversensitivity in LLMs (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Existing benchmarks rely on static datasets that degrade over time as models evolve, leading to data contamination and diminished evaluative power. |
| Approach: | They construct a framework that generates model-specific challenging datasets and aggregates them across diverse LLM families. |
| Outcome: | The framework captures emerging defensive patterns and aligns with each model’s unique behavior. |
Please refuse to answer me! Mitigating Over-Refusal in Large Language Models via Adaptive Contrastive Decoding (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods for mitigating over-refusal can't maintain low refusal ratio for harmless queries while keeping high for malicious queries. |
| Approach: | They propose a model-agnostic approach to mitigate over-refusal in large language models . they propose an adaptive contrastive decoding strategy that incorporates or removes the refusal token distribution . |
| Outcome: | The proposed approach reduces the refusal ratio for over-refusal queries by 10.35% while increasing the refusal rate for malicious queries by 0.13%. |
The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models (2025.naacl-long)
Copied to clipboard
Seungone Kim, Juyoung Suk, Ji Yong Cho, Shayne Longpre, Chaeeun Kim, Dongkeun Yoon, Guijin Son, Yejin Cho, Sheikh Shafayat, Jinheon Baek, Sue Hyun Park, Hyeonbin Hwang, Jinkyung Jo, Hyowon Cho, Haebin Shin, Seongyun Lee, Hanseok Oh, Noah Lee, Namgyu Ho, Se June Joo, Miyoung Ko, Yoonjoo Lee, Hyungjoo Chae, Jamin Shin, Joel Jang, Seonghyeon Ye, Bill Yuchen Lin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, Minjoon Seo
| Challenge: | a recent study evaluated language models using abstract evaluation criteria that lack the flexibility and granularity of human assessment. |
| Approach: | They propose a benchmark to evaluate nine distinct language models' capabilities . they use instance-specific evaluation criteria to mirror human evaluation . |
| Outcome: | The proposed benchmark evaluates nine distinct capabilities of language models across 77 tasks. |