Deep Research with Open-Domain Evaluation and Multi-Stage Guardrails for Safety (2026.acl-long)
Copied to clipboard
Wei-Chieh Huang, Henry Peng Zou, Yaozu Wu, Dongyuan Li, Yankai Chen, Weizhi Zhang, Yangning Li, Angelo Zangari, Jizhou Guo, Chunyu Miao, Liancheng Fang, Langzhou He, Yinghui Li, Renhe Jiang, Philip S. Yu
| Challenge: | Existing deep research frameworks lack adequate evaluation procedures and stage-specific protections. |
| Approach: | They propose a framework with open-domain evaluation and a stage-wise safety benchmark to address this oversight. |
| Outcome: | The proposed framework improves defense success rates by 16.53% while reducing over-refusal rates to approximately 6%. |
Similar Papers
DeepGuard: Secure Code Generation via Multi-Layer Semantic Aggregation (2026.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) can replicate insecure patterns from training data. |
| Approach: | They propose a framework that leverages distributed security-relevant cues by aggregating representations from multiple upper layers via an attention-based module. |
| Outcome: | Experiments show that the framework improves the secure-and-correct generation rate by 11.9% over baselines. |
DeepResearch Retail: Benchmarking Tool-Augmented Deep Research in the E-Commerce Domain (2026.acl-industry)
Copied to clipboard
| Challenge: | Existing DR systems are largely web-centric and do not incorporate structured, domain-specific, and personalized information accessible through internal API tools. |
| Approach: | They propose a framework grounded in real-world e-commerce data for assessing Deep Research with tools in realistic commercial settings. |
| Outcome: | The proposed framework evaluates factual faithfulness and multidimensional response quality when reasoning over heterogeneous web and internal data sources. |
XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing safety evaluations rely on binary labels, overlooking the nuanced risk these outputs pose. |
| Approach: | They propose a framework to assess the severity of extremist content generated by Large Language Models (LLMs) it categorizes model responses into five danger levels (0–4) defined by degree of extremism endorsement . |
| Outcome: | The proposed framework categorizes model responses into five danger levels (0–4) defined by degree of extremist endorsement, enabling nuanced analysis of failure frequency and severity. |
SafeLawBench: Towards Safe Alignment of Large Language Models (2025.findings-acl)
Copied to clipboard
Chuxue Cao, Han Zhu, Jiaming Ji, Qichao Sun, Zhenghao Zhu, Wu Yinyu, Josef Dai, Yaodong Yang, Sirui Han, Yike Guo
| Challenge: | Recent studies indicate that large language models (LLMs) may exhibit risks, including threats to the protection of private data and the generation of hallucinations. |
| Approach: | They propose to evaluate LLMs from a legal perspective using the SafeLawBench benchmark. |
| Outcome: | The proposed framework categorizes safety risks into three levels based on legal standards and includes 24,860 multi-choice questions and 1,106 open-domain question-answering tasks. |
Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation (2026.acl-long)
Copied to clipboard
Fangda Ye, Kuicai Dong, Xie Zhifei, Yuxin Hu, Yihang Yin, Shurui Huang, Shikai Dong, Chen Zhang, Jianzhu Bao, Shuicheng Yan
| Challenge: | Recent agentic search frameworks are text-centric, overlooking multimodal evidence . a pressing task is multimodal long-form generation, a new paper argues . |
| Approach: | They propose a unified agentic framework for grounded multimodal long-form generation. |
| Outcome: | The proposed framework is based on a unified agentic framework for grounded multimodal long-form generation. |
GuardBench: A Large-Scale Benchmark for Guardrail Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Lack of a standard benchmark for guardrail models poses significant evaluation issues . lack of standardized benchmark makes it hard to compare results across scientific publications. |
| Approach: | They propose a large-scale benchmark for guardrail models comprising 40 safety evaluation datasets. |
| Outcome: | The proposed model achieves competitive results without specific fine-tuning without the need for specific fine tuning. |
Guardrails and Security for LLMs: Safe, Secure and Controllable Steering of LLM Applications (2025.acl-tutorials)
Copied to clipboard
Traian Rebedea, Leon Derczynski, Shaona Ghosh, Makesh Narsimhan Sreedhar, Faeze Brahman, Liwei Jiang, Bo Li, Yulia Tsvetkov, Christopher Parisien, Yejin Choi
| Challenge: | Pretrained generative models provide novel ways for users to interact with computers. |
| Approach: | This tutorial provides an overview of key guardrail mechanisms developed for LLMs along with evaluation methodologies and a detailed security assessment protocol. |
| Outcome: | This tutorial provides an overview of key guardrail mechanisms developed for LLMs, along with evaluation methodologies and a detailed security assessment protocol. |
WaterBench: Towards Holistic Evaluation of Watermarks for Large Language Models (2024.acl-long)
Copied to clipboard
| Challenge: | Recent studies have developed watermarking algorithms which restrict the generation process to leave an invisible trace for watermark detection. |
| Approach: | They propose a benchmarking procedure that compares different methods to ensure consistent watermarking strength and jointly evaluates their generation and detection performance. |
| Outcome: | The proposed benchmark compares 4 open-source watermarks on 2 LLMs under 2 watermarking strengths and observes the common struggles for current methods on maintaining the generation quality. |
Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing safety alignment benchmarks fail to evaluate Safe Completion: the model’s ability to maximise helpfulness on dual-use or borderline queries without crossing into actionable harm. |
| Approach: | They propose a large-scale benchmark to measure Over-Refusal and Safe Completion quality in healthcare. |
| Outcome: | The framework evaluates 30 state-of-the-art LLMs including GPT-5 and Claude-4. |
AnalystBench: Benchmarking professional long-form report generation with web-mined multimodal tasks (2026.findings-acl)
Copied to clipboard
Chau Minh Pham, Zichao Wang, Puneet Mathur, Alexa Siu, Akriti Jain, Aparna Garimella, Ananya B. Sai, Nedim Lipka, Mohit Iyyer, Varun Manjunatha
| Challenge: | Existing benchmarks decompose the end-to-end professional report generation into individual components. |
| Approach: | They propose a benchmarking tool that evaluates 20 real-world professional report generation tasks grounded in multimodal document collections. |
| Outcome: | The proposed model outperforms closed-source models on executive summarization tasks but drops significantly on long-horizon synthesis tasks. |