Challenge: Existing benchmarks for deep search agents rely on blackbox web search APIs . dynamic and opaque web APIs hinder reproducibility and fair comparisons - authors .
Approach: They propose a benchmark that employs a fixed corpus for controlled retrieval for deep search agents.
Outcome: The new benchmark shows that agents that combine large language models with retrieval tools excel at complex, reasoning-intensive queries.

Similar Papers

Benchmarking Deep Search over Heterogeneous Enterprise Data (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing methods struggle to conduct deep searches and retrieve all necessary evidence.
Approach: They propose a benchmark for evaluating deep search, a retrieval-augmented generation that requires source-aware, multi-hop reasoning over diverse, sparsed, but related sources.
Outcome: The proposed benchmarks show that even the best-performing agentic RAG methods achieve an average performance score of 32.96 on the benchmark.
WebWalker: Benchmarking LLMs in Web Traversal (2025.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated impressive capabilities across a wide range of natural language processing tasks.
Approach: They propose a benchmark to assess the ability of LLMs to perform web traversal by using an explore-critic paradigm.
Outcome: The proposed framework mimics human-like web navigation through an explore-critic paradigm and demonstrates the effectiveness of RAG combined with WebWalker in real-world scenarios.
Towards Reliable Agents: Benchmarking Customized LLM-Based Retrieval-Augmented Generation Frameworks with Deployment Validation (2025.naacl-industry)

Copied to clipboard

Challenge: Existing benchmarks for general-purpose RAG systems, such as CRAG, RGB, MultiHop-RAG, and CRUD-RAGG, are limited and lack a benchmark specifically tailored to evaluate frameworks.
Approach: They evaluated OpenAI’s Assistants API versus a RAG assistant built with Langchain and deployed a system based on benchmark insights as a course assistant over a two-year span.
Outcome: The proposed benchmarks show that domain-specific retrieval impacts response accuracy and highlight key challenges in real-world deployment.
AutoSearch: Adaptive Search Depth for Efficient Agentic RAG via Reinforcement Learning (2026.findings-acl)

Copied to clipboard

Challenge: Prior work limits search depth to reduce cost, but this often leads to underexploration of complex questions.
Approach: They propose a reinforcement learning framework that evaluates each search step via self-generated intermediate answers.
Outcome: Extensive experiments on multiple benchmarks show that AutoSearch achieves a superior accuracy-efficiency trade-off, alleviating over-searching while preserving search quality.
MAS-Bench: A Unified Benchmark for Shortcut-Augmented Hybrid Mobile GUI Agents (2026.acl-long)

Copied to clipboard

Challenge: Shortcuts such as APIs and deep-links have emerged as efficient complements to flexible GUI operations, but systematic evaluation of GUI–shortcut hybrid agents remains underexplored.
Approach: They propose a benchmark that evaluates GUI-shortcut hybrid agents with a specific focus on the mobile domain.
Outcome: MAS-Bench evaluates agent's ability to generate shortcuts by discovering and creating reusable, low-cost workflows.
Rerank Before You Reason: Analyzing Reranking Tradeoffs through Effective Token Cost in Deep Search Agents (2026.findings-acl)

Copied to clipboard

Challenge: Recent work emphasizes improving efficiency in LLM-based systems, especially for longcontext and multi-step reasoning.
Approach: They analyze the role of listwise reranking in deep search pipelines and compare their results to a novel ETC metric to determine model scale and reasoning effort.
Outcome: The proposed model scale, reasoning effort, reranking depth, and total token cost (ETC) metric improve retrieval and end-to-end accuracy and moderate reranked agents achieve comparable accuracy at substantially lower cost.
A Functionality-Grounded Benchmark for Evaluating Web Agents in E-commerce Domains (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks focus on product search tasks, but ignore potential risks.
Approach: They propose a data generation pipeline that leverages webpage content and interactive elements to create diverse, functionality-grounded user queries.
Outcome: The proposed framework assesses the performance and safety of web agents under dynamic, real-world e-commerce environments.
AgenticRAGTracer: A Hop-Aware Benchmark for Diagnosing Multi-Step Retrieval Reasoning in Agentic RAG (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks provide only final questions and answers, while lacking intermediate hop-level questions that gradually connect atomic questions to the final multi-hop query.
Approach: They propose to build a multi-hop reasoning model that is primarily constructed automatically by large language models and designed to support step-by-step validation.
Outcome: The proposed benchmark spans multiple domains, contains 1,305 data points, and has no overlap with existing mainstream benchmarks.
AgentGym2: Benchmarking Large Language Model Agents in De-Idealized Real-World Environments (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks evaluate agents in simplified, idealized settings, relying on pre-packaged tool interfaces, overlooking critical steps, and assume inputs are clean and fully specified.
Approach: They propose a framework that evaluates language agents in simplified, idealized settings . they show that even SOTA systems like Gemini and GPT-5 struggle on AgentGym2 .
Outcome: Experiments on 15 proprietary and open-source models show that even SOTA systems like Gemini and GPT-5 struggle on AgentGym2 .
GTA: Generating Long-horizon Tasks for Web Agents at Scale (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks provide only coarse start–goal annotations without intermediate trajectories . Existing frameworks provide no supervision over the agent's latent decision process .
Approach: They propose a framework that integrates crawling, retrieval-based seeding, in-context generation and automated quality control to produce realistic tasks paired with executable trajectories.
Outcome: The proposed framework decouples crawling from generation for greater efficiency and ensures dense supervision through deterministic replays and systematic validation.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations