GeoPQA: Bridging the Visual Perception Gap in MLLMs for Geometric Reasoning (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in reinforcement learning (RL) have enhanced the reasoning abilities of large language models, but the impact on multimodal LLMs is limited. |
| Approach: | They propose a two-stage RL framework that enhances visual perception and fosters reasoning capabilities. |
| Outcome: | The proposed framework improves geometric reasoning by 9.7% and problem-solving by 9.1% compared to direct reasoning training approach. |
Similar Papers
GR1: Reinforcement-Enhanced LLM for Geoscience Reasoning (2026.findings-acl)
Copied to clipboard
Yule Xie, Jiaxin Ding, Cheng Deng, Shiqing Gao, Junran Zhang, Sibo Zhang, Zeyuan Wang, Ke Wu, Xin Ding, Luoyi Fu, Meng Jin, Xinbing Wang
| Challenge: | Recent advances in large language models have demonstrated RL's substantial capacity to enhance multi-step reasoning beyond what supervised instruction tuning achieves. |
| Approach: | They propose a framework that converts multimodal questions into descriptive text . they propose RL-enhanced geoscience reasoning that can be fine-tuned to a text-only level . |
| Outcome: | The proposed framework improves accuracy and accuracy on multimodal questions while preserving answerability and difficulty. |
GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to solve geometric problems are dependent on handcraft rules and limited on small-scale datasets. |
| Approach: | They propose a Geometric Question Answering dataset with 5,010 geometric problems with corresponding annotated programs to illustrate the solving process. |
| Outcome: | The proposed method is significantly lower than human performance on the proposed dataset than on a publicly available dataset. |
Beyond the Panorama: Training-Free Hierarchical Perception-Reasoning for Fine-Grained Vision in MLLMs (2026.acl-long)
Copied to clipboard
| Challenge: | Existing multimodal large language models (MLLMs) face challenges in fine-grained visual tasks. |
| Approach: | They propose a training-free hierarchical perception-reasoning framework that enhances fine-grained visual understanding by simulating human perception mechanisms. |
| Outcome: | The proposed framework enhances fine-grained visual understanding by simulating human perception mechanisms. |
Forgotten Polygons: Multimodal Large Language Models are Shape-Blind (2025.findings-acl)
Copied to clipboard
William Rudman, Michal Golovanevsky, Amir Bar, Vedant Palit, Yann LeCun, Carsten Eickhoff, Ritambhara Singh
| Challenge: | Multimodal Large Language Models struggle with visual reasoning, despite strong performance on vision-language tasks. |
| Approach: | They propose a visually cued chain-of-thought prompting that enhances multi-step mathematical reasoning by explicitly referencing visual annotations in diagrams. |
| Outcome: | The proposed model improves GPT-4o's accuracy on an irregular polygon side-counting task from 7% to 93%. |
PseudoGD: Enhancing Spatial Reasoning in Vision-Language Models through Pseudo Geometric Knowledge Distillation (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent Large Vision-Language Models (LVLMs) have shown remarkable success in general semantic understanding, but struggle with 3D spatial reasoning tasks. |
| Approach: | They propose a framework to help vision encoders internalize 3D geometric information using only standard 2D images. |
| Outcome: | The proposed framework achieves State-of-the-Art (SOTA) performance across various model architectures. |
SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data (2025.acl-long)
Copied to clipboard
| Challenge: | Vision-language models struggle with spatial reasoning, a skill that humans excel at. |
| Approach: | They propose to use a spatial-reasoning Enhanced (SpaRE) VLM to improve spatial reasoning in visual question answering and robotics. |
| Outcome: | The proposed model achieves a 49% performance gain on the What's Up benchmark while maintaining strong results on general tasks. |
GeoLAN: Geometric Learning of Latent Explanatory Directions in Large Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Large language models lack transparency and are often unable to explain causal relationships . |
| Approach: | They propose a training framework that treats token representations as geometric trajectories and applies stickiness conditions to the Kakeya Conjecture. |
| Outcome: | The proposed training framework maintains task accuracy while improving geometric metrics and reducing fairness biases. |
Beyond Lines and Circles: Unveiling the Geometric Reasoning Gap in Large Language Models (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Recent advances in Large Language Models (LLMs) demonstrate increasing proficiency in complex mathematical and algorithmic tasks, yet their geometric reasoning skills are underexplored. |
| Approach: | They propose a framework that enhances LLMs’ reasoning potential through a multi-agent system conducting internal dialogue. |
| Outcome: | The proposed framework enhances LLMs’ reasoning potential through a multi-agent system conducting internal dialogue. |
GeoLaux: A Benchmark for Evaluating MLLMs’ Geometry Performance on Long-Step Problems Requiring Auxiliary Lines (2026.acl-long)
Copied to clipboard
| Challenge: | Existing benchmarks for Geometry problem solving lack fine-grained evaluation for long-step problems necessitating auxiliary line construction. |
| Approach: | They present a fine-grained annotated dataset with long-step reasoning and auxiliary line construction that provides a detailed evaluation of 23 leading MLLMs. |
| Outcome: | The proposed model performs significantly worse on long-step problems than short-step ones, with 18 models showing a performance drop of over 50%. |
Sparkle: Mastering Basic Spatial Capabilities in Vision Language Models Elicits Generalization to Spatial Reasoning (2025.findings-emnlp)
Copied to clipboard
Yihong Tang, Ao Qu, Zhaokai Wang, Dingyi Zhuang, Zhaofeng Wu, Wei Ma, Shenhao Wang, Yunhan Zheng, Zhan Zhao, Jinhua Zhao
| Challenge: | Currently, vision-language models excel in many downstream tasks but struggle with spatial reasoning, which is crucial for navigation and interaction with physical environments. |
| Approach: | They propose a framework that generates synthetic data to provide targeted supervision for VLMs across these basic spatial capabilities. |
| Outcome: | The proposed framework disentangles 2D spatial reasoning into three core components: direction comprehension, distance estimation, and localization. |