Challenge: Existing document understanding benchmarks only handle a small number of pages . existing models are limited to handling only a limited number of documents .
Approach: They propose a long document understanding benchmark that integrates three primary tasks and 20 sub-tasks based on different primary tasks.
Outcome: The proposed model outperforms existing benchmarks on open-source and closed-source models . the model outpersforms other models on more than 33,000 pages of documents .

Similar Papers

M-LongDoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework (2025.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks for large multimodal models focus on short documents with less than 50 pages and are limited to extraction-based questions.
Approach: They propose a retrieval-aware tuning approach to improve the accuracy of multimodal document reading by 4.6%.
Outcome: The proposed framework improves the accuracy of model responses by 4.6% compared to existing benchmarks on documents with hundreds of pages and longer documents with more complex content.
MMDocIR: Benchmarking Multimodal Retrieval for Long Documents (2025.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks for multimodal document retrieval are lacking for evaluating performance of systems.
Approach: They propose a benchmark that evaluates page-level and layout-level retrieval tasks . they use a rich dataset featuring 1,685 questions annotated by experts .
Outcome: The proposed benchmark outperforms existing benchmarks in page-level and layout-level retrieval tasks.
LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding (2024.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) can only handle texts a few thousand tokens long, limiting their applications on longer sequence inputs, such as books, reports, and codebases.
Approach: They propose a bilingual, multi-task benchmark for long context understanding that extends context windows and more sophisticated memory mechanisms to improve models' long context capabilities.
Outcome: The proposed model outperforms open-source models but struggles on longer contexts.
MuLD: The Multitask Long Document Benchmark (2022.lrec-1)

Copied to clipboard

Challenge: Existing benchmarks for NLP focus on tasks for one or two sentences, but efficient techniques are needed for processing much longer sequences.
Approach: They propose to modify existing NLP tasks to create a long document benchmark which requires models to successfully model long-term dependencies in the text.
Outcome: The proposed benchmark is much more challenging than its ‘short document’ equivalents.
WildDoc: How Far Are We from Achieving Comprehensive and Robust Document Understanding in the Wild? (2025.emnlp-main)

Copied to clipboard

Challenge: Existing benchmarks for document understanding in the wild are based on scanned or digital documents . however, these benchmarks fail to capture the challenges posed by documents in the real world .
Approach: They propose a new benchmark that incorporates a diverse set of manually captured document images reflecting real-world conditions.
Outcome: The proposed model is based on a set of manually captured document images reflecting real-world conditions and is compared with digital or scanned documents.
DocSplit: A Comprehensive Benchmark Dataset and Evaluation Approach for Document Packet Recognition and Splitting (2026.acl-industry)

Copied to clipboard

Challenge: Document understanding in real-world applications often requires processing heterogeneous, multi-page document packets containing multiple documents stitched together.
Approach: They propose to use document packet splitting to improve document understanding in real-world applications.
Outcome: The proposed datasets and evaluation metrics provide a systematic framework for advancing document understanding capabilities essential for legal, financial, healthcare, and other document-intensive domains.
WikiMixQA: A Multimodal Benchmark for Question Answering over Tables and Charts (2025.findings-acl)

Copied to clipboard

Challenge: Documents are fundamental to preserving and disseminating information, often incorporating complex layouts, tables, and charts that pose significant challenges for automatic document understanding (DU).
Approach: They propose a benchmark for evaluating cross-modal reasoning over tables and charts extracted from 4,000 Wikipedia pages . they evaluate 12 vision-language models that achieve 70% accuracy when provided with direct context .
Outcome: The proposed benchmark evaluates models with high accuracy over tables and charts extracted from 4,000 Wikipedia pages . proprietary models achieve 70% accuracy when provided with direct context, but open-source models perform worse when retrieval from long documents is required.
LMOD: A Large Multimodal Ophthalmology Dataset and Benchmark for Large Vision-Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: Existing benchmarks for large vision-language models (LVLMs) are limited to ophthalmology-specific applications.
Approach: They introduce a large-scale multimodal ophthalmology benchmark consisting of 21,993 instances across five ocular imaging modalities and 13 state-of-the-art LVLM representatives from closed-source, open-source and medical domains.
Outcome: The proposed model shows significant performance drop in ophthalmology compared to other domains.
A Survey on MLLM-based Visually Rich Document Understanding: Methods, Challenges, and Emerging Trends (2026.findings-acl)

Copied to clipboard

Challenge: Visually Rich Document Understanding (VRDU) frameworks are a key area of research . early approaches to VRDU relied on manually crafted rules and domain-specific heuristics . conventional deep learning approaches do not integrate the diverse modalities in documents .
Approach: They review recent advances in MLLM-based Visually Rich Document Understanding (VRDU) their findings highlight emerging trends and promising research directions .
Outcome: The proposed frameworks are scalable, reliable, and adaptable, the authors argue . their findings highlight emerging trends and promising research directions .
IW-Bench: Evaluating Large Multimodal Models for Converting Image-to-Web (2025.findings-acl)

Copied to clipboard

Challenge: Existing models have been introduced to improve image comprehension, but there is no robust benchmark for imagetoweb conversion.
Approach: They propose a benchmark to assess imagetoweb conversion proficiency of large multimodal models . they propose to measure layout information of web pages by parsing the Document Object Model tree .
Outcome: The proposed benchmark measures the layout information of web pages—i.e., the positional relationships between elements—which has been overlooked by prior work.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations