Challenge: Recent advances in reinforcement learning (RL) have enhanced the reasoning abilities of large language models, but the impact on multimodal LLMs is limited.
Approach: They propose a two-stage RL framework that enhances visual perception and fosters reasoning capabilities.
Outcome: The proposed framework improves geometric reasoning by 9.7% and problem-solving by 9.1% compared to direct reasoning training approach.

Similar Papers

GR1: Reinforcement-Enhanced LLM for Geoscience Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Recent advances in large language models have demonstrated RL's substantial capacity to enhance multi-step reasoning beyond what supervised instruction tuning achieves.
Approach: They propose a framework that converts multimodal questions into descriptive text . they propose RL-enhanced geoscience reasoning that can be fine-tuned to a text-only level .
Outcome: The proposed framework improves accuracy and accuracy on multimodal questions while preserving answerability and difficulty.
GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning (2021.findings-acl)

Copied to clipboard

Challenge: Existing methods to solve geometric problems are dependent on handcraft rules and limited on small-scale datasets.
Approach: They propose a Geometric Question Answering dataset with 5,010 geometric problems with corresponding annotated programs to illustrate the solving process.
Outcome: The proposed method is significantly lower than human performance on the proposed dataset than on a publicly available dataset.
Beyond the Panorama: Training-Free Hierarchical Perception-Reasoning for Fine-Grained Vision in MLLMs (2026.acl-long)

Copied to clipboard

Challenge: Existing multimodal large language models (MLLMs) face challenges in fine-grained visual tasks.
Approach: They propose a training-free hierarchical perception-reasoning framework that enhances fine-grained visual understanding by simulating human perception mechanisms.
Outcome: The proposed framework enhances fine-grained visual understanding by simulating human perception mechanisms.
Forgotten Polygons: Multimodal Large Language Models are Shape-Blind (2025.findings-acl)

Copied to clipboard

Challenge: Multimodal Large Language Models struggle with visual reasoning, despite strong performance on vision-language tasks.
Approach: They propose a visually cued chain-of-thought prompting that enhances multi-step mathematical reasoning by explicitly referencing visual annotations in diagrams.
Outcome: The proposed model improves GPT-4o's accuracy on an irregular polygon side-counting task from 7% to 93%.
PseudoGD: Enhancing Spatial Reasoning in Vision-Language Models through Pseudo Geometric Knowledge Distillation (2026.findings-acl)

Copied to clipboard

Challenge: Recent Large Vision-Language Models (LVLMs) have shown remarkable success in general semantic understanding, but struggle with 3D spatial reasoning tasks.
Approach: They propose a framework to help vision encoders internalize 3D geometric information using only standard 2D images.
Outcome: The proposed framework achieves State-of-the-Art (SOTA) performance across various model architectures.
SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data (2025.acl-long)

Copied to clipboard

Challenge: Vision-language models struggle with spatial reasoning, a skill that humans excel at.
Approach: They propose to use a spatial-reasoning Enhanced (SpaRE) VLM to improve spatial reasoning in visual question answering and robotics.
Outcome: The proposed model achieves a 49% performance gain on the What's Up benchmark while maintaining strong results on general tasks.
GeoLAN: Geometric Learning of Latent Explanatory Directions in Large Language Models (2026.findings-acl)

Copied to clipboard

Challenge: Large language models lack transparency and are often unable to explain causal relationships .
Approach: They propose a training framework that treats token representations as geometric trajectories and applies stickiness conditions to the Kakeya Conjecture.
Outcome: The proposed training framework maintains task accuracy while improving geometric metrics and reducing fairness biases.
Beyond Lines and Circles: Unveiling the Geometric Reasoning Gap in Large Language Models (2024.findings-emnlp)

Copied to clipboard

Challenge: Recent advances in Large Language Models (LLMs) demonstrate increasing proficiency in complex mathematical and algorithmic tasks, yet their geometric reasoning skills are underexplored.
Approach: They propose a framework that enhances LLMs’ reasoning potential through a multi-agent system conducting internal dialogue.
Outcome: The proposed framework enhances LLMs’ reasoning potential through a multi-agent system conducting internal dialogue.
GeoLaux: A Benchmark for Evaluating MLLMs’ Geometry Performance on Long-Step Problems Requiring Auxiliary Lines (2026.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for Geometry problem solving lack fine-grained evaluation for long-step problems necessitating auxiliary line construction.
Approach: They present a fine-grained annotated dataset with long-step reasoning and auxiliary line construction that provides a detailed evaluation of 23 leading MLLMs.
Outcome: The proposed model performs significantly worse on long-step problems than short-step ones, with 18 models showing a performance drop of over 50%.
Sparkle: Mastering Basic Spatial Capabilities in Vision Language Models Elicits Generalization to Spatial Reasoning (2025.findings-emnlp)

Copied to clipboard

Challenge: Currently, vision-language models excel in many downstream tasks but struggle with spatial reasoning, which is crucial for navigation and interaction with physical environments.
Approach: They propose a framework that generates synthetic data to provide targeted supervision for VLMs across these basic spatial capabilities.
Outcome: The proposed framework disentangles 2D spatial reasoning into three core components: direction comprehension, distance estimation, and localization.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations