Challenge: Recent advances in multimodal large language models have led to progress in tackling complex reasoning tasks that combine textual and visual information.
Approach: They introduce a robust version of the Massive Multi-discipline Multimodal Understanding and Reasoning (MMMU) benchmark.
Outcome: The proposed model performs lower on MMMU-Pro than on the previous benchmark, ranging from 16.8% to 26.9%.

Similar Papers

Uni-MMMU: A Massive Multi-discipline Multimodal Unified Benchmark (2026.acl-long)

Copied to clipboard

Challenge: Existing evaluations treat visual understanding and generation in isolation or overlook tasks that inherently couple them.
Approach: They propose a benchmark that examines the bidirectional synergy between generation and understanding across eight reasoning-centric domains.
Outcome: The proposed model systematically unfolds the bidirectional synergy between generation and understanding across eight reasoning-centric domains.
TIU-Bench: A Benchmark for Evaluating Large Multimodal Models on Text-rich Image Understanding (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing text-rich image understanding benchmarks lack scale and fragmented scenarios . a new full-image structured output format is proposed to enable fine-grained evaluation of perception and reasoning capabilities.
Approach: They propose a large-scale, multilingual benchmark that includes over 100,000 annotations and 22,000 question-answer pairs.
Outcome: The proposed framework provides a comprehensive platform for developing and evaluating next-generation multimodal AI systems.
MMEvalPro: Calibrating Multimodal Benchmarks Towards Trustworthy and Efficient Evaluation (2025.naacl-long)

Copied to clipboard

Challenge: Large Multimodal Models (LMMs) exhibit impressive cross-modal understanding and reasoning abilities, but many benchmarks suffer from systematic biases.
Approach: They propose a benchmark to avoid Type-I errors by creating one perception question and one knowledge anchor question through a meticulous annotation process.
Outcome: The proposed benchmark avoids Type-I errors while maintaining reliability of MCQ evaluations.
Multimodal Inconsistency Reasoning (MMIR): A New Benchmark for Multimodal Reasoning Models (2025.findings-acl)

Copied to clipboard

Challenge: Existing Multimodal Large Language Models (MLLMs) are predominantly trained on consistent visual-textual inputs, leaving open the question of whether they can handle semantic mismatches in layout-rich content.
Approach: They propose to use multimodal inconsistency reasoning to assess MLLMs' ability to reason about semantic mismatches in webpages, presentation slides, and posters.
Outcome: The proposed model outperforms open-source models in detecting inconsistencies in webpages, presentation slides, and posters while remaining vulnerable to inconsistent errors.
JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation (2025.naacl-long)

Copied to clipboard

Challenge: Using culture-agnostic subsets, performance drops in many LMMs when evaluated in Japanese.
Approach: They introduce a Japanese benchmark to evaluate large multimodal models on expert-level tasks based on the Japanese cultural context.
Outcome: The proposed benchmark enables comparisons with other benchmarks in other languages based on cultural contexts.
Reasoning is All You Need for Video Generalization: A Counterfactual Benchmark with Sub-question Evaluation (2025.findings-acl)

Copied to clipboard

Challenge: Existing multimodal benchmarks often overlook counterfactual reasoning, which is crucial for robust video understanding.
Approach: They propose a multidimensional multimodal benchmark that systematically evaluates MLLMs across the abstract-concrete and perception-cognition dimensions.
Outcome: The proposed model decomposes complex queries into structured sub-questions, enabling fine-grained reasoning analysis.
Video-MMMU: Evaluating Knowledge Acquisition from Multidisciplinary Professional Videos (2026.acl-long)

Copied to clipboard

Challenge: Existing video benchmarks do not evaluate the knowledge acquisition capabilities of Large Multimodal Models (LMMs) existing video benchmark focuses on static, general visual understanding tasks, without evaluating whether models can acquire knowledge dynamically.
Approach: They propose a multi-modal, multi-discipline, multitrack benchmark that evaluates Large Multimodal Models’ ability to acquire knowledge from college-level, educational videos.
Outcome: The proposed benchmark reveals a substantial gap between human learners and current Large Multimodal Models (LMMs) and focuses on improving their learning efficiency.
MT-Video-Bench: A Holistic Video Understanding Benchmark for Evaluating Multimodal LLMs in Multi-Turn Dialogues (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluation benchmarks for Multimodal Large Language Models (MLLMs) focus on single-turn question answering, overlooking the complexity of multi-turn dialogues in real-world scenarios.
Approach: They propose a video understanding benchmark for MLLMs in multi-turn dialogues that assesses six core competencies that focus on perceptivity and interactivity.
Outcome: The MT-Video-Bench evaluates 1,000 multi-turn dialogues from diverse domains and reveals significant performance discrepancies and limitations in handling multi-turned video dialogues.
Beyond Perception: Evaluating Abstract Visual Reasoning through Multi-Stage Task (2025.findings-acl)

Copied to clipboard

Challenge: Existing AVR benchmarks focus on single-step reasoning, emphasizing the end result but neglecting the multi-stage nature of reasoning process.
Approach: They propose a multi-stage AVR benchmark based on RAVEN to assess reasoning across varying levels of complexity.
Outcome: The proposed metric considers the correctness of intermediate steps in addition to the final outcomes.
MMTabReal: Real-World Benchmark for Multimodal Table Understanding (2026.findings-acl)

Copied to clipboard

Challenge: Multimodal tables are ubiquitous in real applications but are difficult to evaluate in multimodal large language models.
Approach: They propose a multimodal table benchmark that compares 500 real-world tables with 4021 question–answer pairs.
Outcome: MMtabReal spans four question types, five reasoning categories, and eight structural archetypes.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations