Challenge: Recent advances in text-only "slow thinking" reasoning have prompted efforts to transfer this capability to vision-language models (VLMs).
Approach: They propose a VRM Reflection-V which enhances visual reflection based on reasoning data for cold-start and reward design for reinforcement learning.
Outcome: The proposed model improves visual reflection for cold-start and reward design for reinforcement learning (RL) it maintains a stronger and more consistent reliance on visual information during visual reasoning, indicating effective enhancement in visual reflection capabilities.

Similar Papers

Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Vision-language models have demonstrated strong efficacy as visual assistants . however, evaluation of their reasoning capabilities requires a costly benchmark .
Approach: They propose a pipeline to measure the reasoning consistency of vision-language models . they propose supervised fine-tuning of VLMs and feedback from LLMs .
Outcome: The proposed framework reduces cost while ensuring the generation of a high-quality dataset.
Journey Before Destination: On the importance of Visual Faithfulness in Slow Thinking (2026.eacl-long)

Copied to clipboard

Challenge: Existing evaluations for visual hallucinations are narrow.
Approach: They propose a framework that decomposes reasoning chains into perception versus reasoning steps and uses off-the-shelf VLM judges for step-level faithfulness.
Outcome: The proposed framework reduces Unfaithful Perception Rate while preserving final-answer accuracy.
Enhancing Advanced Visual Reasoning Ability of Large Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Recent advances in Vision-Language (VL) research have sparked new benchmarks for complex visual reasoning, challenging models’ advanced reasoning ability.
Approach: They propose a novel multi-modal in-context learning methodology to enhance LLMs’ contextual understanding and reasoning.
Outcome: The proposed model achieves SOTA performance among all visual reasoning tasks and achieves a 'higher level of accuracy' than previous models.
Vision-Language Models Can Self-Improve Reasoning via Reflection (2025.naacl-long)

Copied to clipboard

Challenge: Chain-of-thought (CoT) has been shown to improve the reasoning capability of large language models (LLMs).
Approach: They propose a framework which iteratively enhances the model’s Vision-language Reasoning by Reflecting on CoT Rationales.
Outcome: The proposed framework improves multimodal reasoning on vision-language tasks by 23% to 60% over baselines.
Visually-Guided Policy Optimization for Multimodal Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Existing RLVRs lack visual faithfulness due to text-dominated reasoning . a novel framework to reinforce visual focus during policy optimization is proposed .
Approach: They propose a framework to reinforce visual focus during policy optimization using visual attention compensation mechanism.
Outcome: The proposed framework exhibits better visual activation and superior performance in multimodal reasoning and visual-dependent tasks.
VReST: Enhancing Reasoning in Large Vision-Language Models through Tree Search and Self-Reward Mechanism (2025.acl-long)

Copied to clipboard

Challenge: Large Vision-Language Models (LVLMs) have shown exceptional performance in multimodal tasks, but their effectiveness in complex visual reasoning is constrained.
Approach: They propose a training-free approach that enhances Reasoning in Large Vision-Language Models . they propose integrating Monte Carlo Tree Search and Self-Reward mechanisms into the reasoning tree .
Outcome: The proposed approach surpasses current prompting methods and secures state-of-the-art performance across three multimodal reasoning benchmarks.
ProgressLM: Towards Progress Reasoning in Vision-Language Models (2026.acl-long)

Copied to clipboard

Challenge: Existing models for task progress estimation lack long-horizon and dynamic reasoning . estimating how much of a task has been completed requires long-term reasoning based on partial information.
Approach: They propose a benchmark for evaluating progress reasoning from a single observation . they instantiate a two-stage paradigm that combines episodic retrieval with mental simulation .
Outcome: The proposed benchmark improves on 14 VLMs on a small scale and shows common failure patterns.
SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data (2025.acl-long)

Copied to clipboard

Challenge: Vision-language models struggle with spatial reasoning, a skill that humans excel at.
Approach: They propose to use a spatial-reasoning Enhanced (SpaRE) VLM to improve spatial reasoning in visual question answering and robotics.
Outcome: The proposed model achieves a 49% performance gain on the What's Up benchmark while maintaining strong results on general tasks.
Look Less, Reason More: Rollout-Guided Adaptive Pixel-Space Reasoning (2026.acl-long)

Copied to clipboard

Challenge: Recent work has shown promise by incorporating pixel-level visual information into the reasoning process, enabling VLMs to access high-resolution visual details during their thought process.
Approach: They propose a framework that dynamically determines necessary pixel-level operations based on the input query.
Outcome: The proposed model achieves 73.4% accuracy on HR-Bench 4K while maintaining a tool usage ratio of only 20.1%, improving accuracy and reducing tool usage by 66.5% compared to the previous methods.
Addressing Overthinking in Large Vision-Language Models via Gated Perception-Reasoning Optimization (2026.findings-acl)

Copied to clipboard

Challenge: Prior work has attempted to mitigate this issue by using adaptive reasoning strategies, but these methods overlook a fundamental bottleneck: visual perception failures.
Approach: They propose a meta-reasoning controller that dynamically routes computation among three decision paths at each generation step.
Outcome: The proposed method outperforms slow-thinking methods while producing shorter responses.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations