Challenge: Existing benchmarks focus on single-document understanding, whereas real scientific workflows require integrating evidence from multiple papers.
Approach: They propose a multi-modal multi-document benchmark for agentic deep research that integrates evidence from multiple documents.
Outcome: Experimental results show that even advanced systems achieve limited scores on PaperScope . paper provides a rigorous benchmark alongside a pipeline for constructing large multi-modal, multi-source deep research datasets.

Similar Papers

PAPERMIND: Benchmarking Agentic Reasoning and Critique over Scientific Papers in Multimodal LLMs (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks assess integrated and agent-oriented scientific reasoning in isolation . Existing systems assess integrated reasoning in isolated tasks .
Approach: They propose a benchmark to evaluate integrated and agent-oriented scientific reasoning over research papers.
Outcome: The proposed benchmark evaluates integrated and agent-oriented scientific reasoning over scientific papers.
WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts (2025.findings-acl)

Copied to clipboard

Challenge: Documents are fundamental to preserving and disseminating information, often incorporating complex layouts, tables, and charts that pose significant challenges for automatic document understanding (DU).
Approach: They propose a benchmark for evaluating cross-modal reasoning over tables and charts extracted from 4,000 Wikipedia pages . they evaluate 12 vision-language models that achieve 70% accuracy when provided with direct context .
Outcome: The proposed benchmark evaluates models with high accuracy over tables and charts extracted from 4,000 Wikipedia pages . proprietary models achieve 70% accuracy when provided with direct context, but open-source models perform worse when retrieval from long documents is required.
SciMMIR: Benchmarking Scientific Multi-modal Information Retrieval (2024.findings-acl)

Copied to clipboard

Challenge: Multi-modal information retrieval (MMIR) is a rapidly evolving field . current benchmarks for image-text pairings overlook the scientific domain .
Approach: They develop a scientific domain-specific MMIR benchmark to evaluate image-text pairings using open-access research paper corpora.
Outcome: The proposed benchmarks are based on 530K image-text pairs extracted from scientific documents with detailed captions.
Paper Circle: An Open-source Multi-agent Research Discovery and Analysis Framework (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in large language models have demonstrated strong potential for understanding user intent . paper describes system architecture, agent roles, retrieval and scoring methods, knowledge graph schema, and evaluation interfaces .
Approach: They propose a multi-agent research discovery and analysis system that integrates multiple agents to reduce the effort required to find, assess, organize, and understand academic literature.
Outcome: The proposed system reduces the effort required to find, assess, organize, and understand academic literature.
SciMDR: Advancing Scientific Multimodal Document Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Current models struggle to provide reliable assistance in real-world scientific workflows because evidence is distributed across long, multimodal documents.
Approach: They propose a framework for QA Synthesis and document-scale regrounding that generates faithful, isolated QA pairs and reasoning on focused segments.
Outcome: The proposed framework achieves significant improvements across multiple QA benchmarks, particularly in tasks requiring complex document-level reasoning.
M3SciQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing evaluation benchmarks for foundation models in understanding scientific literature focus on single-document tasks.
Approach: They propose a multi-modal, multi-document scientific question answering benchmark . it uses expert-annotated questions that span 70 natural language processing paper clusters .
Outcome: The proposed benchmarks underperform human experts in multi-modal reasoning and retrieval of scientific data.
SciVQR: A Multidisciplinary Multimodal Benchmark for Advanced Scientific Reasoning Evaluation (2026.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for multimodal large language models fail to capture complexity and traceability of reasoning processes . SciVQR includes domain-specific visuals and challenges models to combine visual comprehension with reasoning.
Approach: They propose a multimodal benchmark for scientific reasoning covering 54 subfields . SciVQR includes domain-specific visuals and challenges models to combine visual comprehension with reasoning .
Outcome: SciVQR evaluates 54 subfields in mathematics, physics, chemistry, geography, astronomy, and biology . the results highlight the need for improved multi-step reasoning and integration of interdisciplinary knowledge .
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems (2025.findings-acl)

Copied to clipboard

Challenge: SciVerse is a multi-modal scientific evaluation benchmark to assess large multi-models . it examines the scientific knowledge comprehension, multi-mod content interpretation and Chain-of-Thought reasoning . authors examine the scientific proficiency of LMMs in scientific domains based on their work .
Approach: They propose a multi-modal scientific evaluation benchmark to thoroughly assess Large Multi-modal Models across 5,735 test instances in five different versions.
Outcome: The proposed evaluation reveals critical limitations in LMMs' scientific proficiency and provides new insights into future developments.
ToolScope: An Agentic Framework for Vision-Guided and Long-Horizon Tool Use (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models have demonstrated remarkable problem-solving capabilities . however, enabling multimodal large language model to flexibly and efficiently utilize external tools remains a challenge .
Approach: They introduce an agentic framework to unify global planning with local multimodal perception . they evaluate ToolScope on four VQA benchmarks across diverse domains .
Outcome: The proposed framework unifies global planning with local multimodal perception . it adopts a specialized Perceive tool to mitigate visual context degradation in long-horizon VQA task.
DocMMIR: A Framework for Document Multi-modal Information Retrieval (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing multi-modal information retrieval models lack a comprehensive exploration of document-level retrieval . existing models suffer from the absence of cross-domain datasets at this granularity.
Approach: They propose a multi-modal document retrieval framework to unify diverse document formats and domains with a comprehensive retrieval scenario.
Outcome: The proposed framework improves document retrieval performance on a large multimodal dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations