Challenge: Large vision-language models often prioritize language knowledge over image information on visual reasoning tasks, incurring performance degradation.
Approach: They propose a visual reasoning framework that decouples vision-reasoning capabilities and multi-run proactive perception.
Outcome: The proposed framework outperforms existing models on benchmarks for open-source and closed-source models with 13.2% performance gain.

Similar Papers

Can MLLMs Reason Beyond Language? VisReason: A Comprehensive Benchmark for Vision-Centric Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances in multimodal large language models demonstrate strong performance on visual reasoning benchmarks.
Approach: They propose a benchmark for vision-centric reasoning that integrates visual and textual information for non-trivial reasoning.
Outcome: The proposed benchmark exposes gaps between humans and current MLLMs and reveals limited benefits from test-time reasoning strategies.
VL-Calibration: Decoupled Confidence Calibration for Large Vision-Language Models Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Existing verbalized confidence calibration methods for large vision language models optimize a single holistic confidence score using binary answer-level correctness.
Approach: They propose a reinforcement learning framework that explicitly decouples confidence into visual and reasoning confidence.
Outcome: Experiments show that the proposed framework decouples confidence into visual and reasoning confidence while suppressing ungrounded hallucinations while preserving valid perception.
Enhancing Advanced Visual Reasoning Ability of Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in Vision-Language (VL) research have sparked new benchmarks for complex visual reasoning, challenging models’ advanced reasoning ability.
Approach: They propose a novel multi-modal in-context learning methodology to enhance LLMs’ contextual understanding and reasoning.
Outcome: The proposed model achieves SOTA performance among all visual reasoning tasks and achieves a 'higher level of accuracy' than previous models.
EfficientVLM: Fast and Accurate Vision-Language Models via Knowledge Distillation and Modal-adaptive Pruning (2023.findings-acl)

Copied to clipboard

Challenge: Pre-trained vision-language models have achieved impressive results in a range of vision-linguistic tasks.
Approach: They propose a distilling then pruning framework to compress large vision-language models into smaller, faster ones.
Outcome: The proposed framework reduces the size of a pre-trained large vision-language model and improves its performance on vision-linguistic tasks.
MM-Reasoner: A Multi-Modal Knowledge-Aware Framework for Knowledge-Based Visual Question Answering (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent knowledge-based visual question answering approaches miss visual information captured by captions and cannot fully utilize the visual information required to answer the question.
Approach: They propose a framework that extracts visual information from an image and prompts an LLM to extract query-specific knowledge from the extracted textual information.
Outcome: Empirical results show that MM-Reasoner achieves state-of-the-art performance on several KVQA datasets.
Mitigating Visual Forgetting via Take-along Visual Conditioning for Multi-modal Long CoT Reasoning (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) have demonstrated enhanced reasoning capabilities, evolving from simple Chain-of-Thought (CoT) prompting to advanced, product-oriented solutions like OpenAI o1 .
Approach: They propose a strategy that shifts image input to critical reasoning stages and compresses redundant visual tokens via dynamic pruning.
Outcome: The proposed model achieves state-of-the-art on five mathematical reasoning benchmarks (+3.4% vs previous sota) and demonstrates iterative reasoning capabilities for complex multi-step tasks.
LVPruning: An Effective yet Simple Language-Guided Vision Token Pruning Approach for Multi-modal Large Language Models (2025.findings-naacl)

Copied to clipboard

Challenge: Multi-modal Large Language Models (MLLMs) incur significant computational overhead due to the large number of vision tokens processed, limiting their practicality in resource-constrained environments.
Approach: They propose a language-guided vision token pruning method that can be integrated into existing MLLMs with minimal architectural changes.
Outcome: The proposed method reduces vision tokens by 90% and preserves model performance.
RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in multi-modal learning have enhanced MLLMs' ability to reason about visual content.
Approach: They propose a framework that unifies multi-step multimodal reasoning with grounded visual understanding.
Outcome: The proposed framework surpasses state-of-the-art methods by +6.5 gIoU and +9.2 cIou on ReasonSeg and achieves 49.7 mAP on SegInW under zero-shot settings.
VisiPruner: Decoding Discontinuous Cross-Modal Dynamics for Efficient Multimodal LLMs (2025.emnlp-main)

Copied to clipboard

Challenge: Multimodal Large Language Models (MLLMs) suffer from significant computational overhead due to the quadratic growth of attention computations with the number of multimodal tokens.
Approach: They propose a training-free pruning framework that prunes multimodal tokens without a trained pruning method.
Outcome: The proposed pruning framework outperforms existing token pruning methods and generalizes across diverse MLLMs.
LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs (2025.findings-acl)

Copied to clipboard

Challenge: Existing approaches do not emphasize step-wise problem-solving.
Approach: They propose a visual reasoning chain benchmark and a fine-grained reasoning metric that evaluates correctness and logical coherence at each step.
Outcome: The proposed framework outperforms existing models in six benchmarks and is 5x faster during inference scaling.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations