Challenge: Existing deep research frameworks lack adequate evaluation procedures and stage-specific protections.
Approach: They propose a framework with open-domain evaluation and a stage-wise safety benchmark to address this oversight.
Outcome: The proposed framework improves defense success rates by 16.53% while reducing over-refusal rates to approximately 6%.

Similar Papers

DeepGuard: Secure Code Generation via Multi-Layer Semantic Aggregation (2026.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) can replicate insecure patterns from training data.
Approach: They propose a framework that leverages distributed security-relevant cues by aggregating representations from multiple upper layers via an attention-based module.
Outcome: Experiments show that the framework improves the secure-and-correct generation rate by 11.9% over baselines.
DeepResearch Retail: Benchmarking Tool-Augmented Deep Research in the E-Commerce Domain (2026.acl-industry)

Copied to clipboard

Challenge: Existing DR systems are largely web-centric and do not incorporate structured, domain-specific, and personalized information accessible through internal API tools.
Approach: They propose a framework grounded in real-world e-commerce data for assessing Deep Research with tools in realistic commercial settings.
Outcome: The proposed framework evaluates factual faithfulness and multidimensional response quality when reasoning over heterogeneous web and internal data sources.
XGUARD: A Graded Benchmark for Evaluating Safety Failures of Large Language Models on Extremist Content (2026.findings-acl)

Copied to clipboard

Challenge: Existing safety evaluations rely on binary labels, overlooking the nuanced risk these outputs pose.
Approach: They propose a framework to assess the severity of extremist content generated by Large Language Models (LLMs) it categorizes model responses into five danger levels (0–4) defined by degree of extremism endorsement .
Outcome: The proposed framework categorizes model responses into five danger levels (0–4) defined by degree of extremist endorsement, enabling nuanced analysis of failure frequency and severity.
SafeLawBench: Towards Safe Alignment of Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Recent studies indicate that large language models (LLMs) may exhibit risks, including threats to the protection of private data and the generation of hallucinations.
Approach: They propose to evaluate LLMs from a legal perspective using the SafeLawBench benchmark.
Outcome: The proposed framework categorizes safety risks into three levels based on legal standards and includes 24,860 multi-choice questions and 1,106 open-domain question-answering tasks.
Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation (2026.acl-long)

Copied to clipboard

Challenge: Recent agentic search frameworks are text-centric, overlooking multimodal evidence . a pressing task is multimodal long-form generation, a new paper argues .
Approach: They propose a unified agentic framework for grounded multimodal long-form generation.
Outcome: The proposed framework is based on a unified agentic framework for grounded multimodal long-form generation.
GuardBench: A Large-Scale Benchmark for Guardrail Models (2024.emnlp-main)

Copied to clipboard

Challenge: Lack of a standard benchmark for guardrail models poses significant evaluation issues . lack of standardized benchmark makes it hard to compare results across scientific publications.
Approach: They propose a large-scale benchmark for guardrail models comprising 40 safety evaluation datasets.
Outcome: The proposed model achieves competitive results without specific fine-tuning without the need for specific fine tuning.
Guardrails and Security for LLMs: Safe, Secure and Controllable Steering of LLM Applications (2025.acl-tutorials)

Copied to clipboard

Challenge: Pretrained generative models provide novel ways for users to interact with computers.
Approach: This tutorial provides an overview of key guardrail mechanisms developed for LLMs along with evaluation methodologies and a detailed security assessment protocol.
Outcome: This tutorial provides an overview of key guardrail mechanisms developed for LLMs, along with evaluation methodologies and a detailed security assessment protocol.
WaterBench: Towards Holistic Evaluation of Watermarks for Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Recent studies have developed watermarking algorithms which restrict the generation process to leave an invisible trace for watermark detection.
Approach: They propose a benchmarking procedure that compares different methods to ensure consistent watermarking strength and jointly evaluates their generation and detection performance.
Outcome: The proposed benchmark compares 4 open-source watermarks on 2 LLMs under 2 watermarking strengths and observes the common struggles for current methods on maintaining the generation quality.
Health-ORSC-Bench: A Benchmark for Measuring Over-Refusal and Safety Completion in Health Context (2026.findings-acl)

Copied to clipboard

Challenge: Existing safety alignment benchmarks fail to evaluate Safe Completion: the model’s ability to maximise helpfulness on dual-use or borderline queries without crossing into actionable harm.
Approach: They propose a large-scale benchmark to measure Over-Refusal and Safe Completion quality in healthcare.
Outcome: The framework evaluates 30 state-of-the-art LLMs including GPT-5 and Claude-4.
AnalystBench: Benchmarking professional long-form report generation with web-mined multimodal tasks (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks decompose the end-to-end professional report generation into individual components.
Approach: They propose a benchmarking tool that evaluates 20 real-world professional report generation tasks grounded in multimodal document collections.
Outcome: The proposed model outperforms closed-source models on executive summarization tasks but drops significantly on long-horizon synthesis tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations