FoREST: Frame of Reference Evaluation in Spatial Reasoning Tasks (2025.emnlp-main)
Copied to clipboard
| Challenge: | Spatial reasoning is a fundamental aspect of human intelligence. |
| Approach: | They propose a framework to assess FoR comprehension in large language models (LLMs) by using the Frame of Reference Evaluation in Spatial Reasoning Tasks benchmark. |
| Outcome: | The proposed method improves overall performance across spatial reasoning tasks. |
Similar Papers
Can LLMs See Without Pixels? Benchmarking Spatial Intelligence from Textual Descriptions (2026.findings-acl)
Copied to clipboard
Zhongbin Guo, Zhen Yang, Yushan Li, Xinyue Zhang, Wenyu Gao, Jiacheng Wang, Chengzhi Li, Xiangrui Liu, Ping Jian
| Challenge: | Existing advances in Spatial Intelligence rely on vision-Language Models . however, a critical question remains: does spatial understanding originate from visual encoders? |
| Approach: | They propose to evaluate the SI performance of Large Language Models without pixel-level input. |
| Outcome: | The proposed benchmark challenges large language models to perform symbolic reasoning rather than visual pattern matching. |
A Benchmark for Reasoning with Spatial Prepositions (2023.emnlp-main)
Copied to clipboard
| Challenge: | Spatial reasoning is a fundamental building block of human cognition . large language models (LLMs) are not on par with advanced aspects of human cognitive domains . |
| Approach: | They propose a benchmark to assess inferential properties of statements with spatial prepositions . they use prompt engineering to test the performance of two large language models . |
| Outcome: | The proposed benchmark shows that none of the models reaches human performance. |
Disentangling Extraction and Reasoning in Multi-hop Spatial Reasoning (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Recent studies highlight the struggles even large language models encounter when it comes to performing spatial reasoning over text. |
| Approach: | They propose to disentangle spatial reasoning over text and compare them to state-of-the-art models with no explicit design for these parts. |
| Outcome: | The proposed models show that they can perform spatial reasoning over text and can generalize within real data domains. |
SPHERE: Unveiling Spatial Blind Spots in Vision-Language Models Through Hierarchical Evaluation (2025.acl-long)
Copied to clipboard
Wenyu Zhang, Wei En Ng, Lixin Ma, Yuwen Wang, Junqi Zhao, Allison Koenecke, Boyang Li, Wanglu Wanglu
| Challenge: | Current vision-language models lack multi-dimensional spatial reasoning capabilities for human-like understanding and applications. |
| Approach: | They propose a hierarchical evaluation framework that probes models across increasing levels of complexity and integrates spatial, visual, and logical understanding. |
| Outcome: | The proposed framework probes models across increasing levels of complexity, from basic skills to multi-skill integration and high-level reasoning that combines spatial, visual, and logical understanding. |
RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks (2025.emnlp-industry)
Copied to clipboard
Amit Agarwal, Hitesh Laxmichand Patel, Srikant Panda, Hansa Meghwani, Jyotika Singh, Karan Dua, Paul Li, Tao Sheng, Sujith Ravi, Dan Roth
| Challenge: | Existing evaluation methods do not explicitly measure this distinction, hindering effective dataset curation and real-world focused model development. |
| Approach: | They introduce a region-based score to quantify a dataset's reliance on global versus local visual information. |
| Outcome: | The proposed model-based score systematically compares model performance on image patches versus full images to determine if tasks require holistic image understanding or can be solved with partial or localized visual cues. |
SpaRC and SpaRP: Spatial Reasoning Characterization and Path Generation for Understanding Spatial Reasoning Capability of Large Language Models (2024.acl-long)
Copied to clipboard
| Challenge: | Existing large language models (LLMs) do not perform well on the datasets. |
| Approach: | They propose to use a Spatial Reasoning Characterization framework and a spatial reasoning path framework to study spatial reasoning. |
| Outcome: | The proposed framework and datasets outperform state-of-the-art models in spatial reasoning. |
KnowDR-REC: Auditing Knowledge-Conditioned Visual Grounding in Referring Expression Comprehension (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluation metrics suggest that Multimodal large language models have acquired fine-grained visual grounding capabilities. |
| Approach: | They propose a benchmark to assess Referring Expression Comprehension (REC) that uses intra-image visual cues to localize target objects and a controllable evaluation mechanism to test sensitivity to fine-grained factual changes. |
| Outcome: | The proposed benchmarks show that multimodal large language models have a high level of performance on the RefCOCO family of benchmarks. |
Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language Models (2024.emnlp-main)
Copied to clipboard
| Challenge: | Despite recent advances in visual language models, their ability to quantitatively reason about object sizes and distances remains underexplored. |
| Approach: | They propose a manually annotated benchmark of 241 questions designed for quantitative spatial reasoning and a zero-shot prompting technique that encourages VLMs to use reference objects as visual cues. |
| Outcome: | The proposed technique improves the performance of the top-performing VLMs by 19 points when a reasoning path using a reference object emerges naturally in the response. |
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs (2026.acl-short)
Copied to clipboard
| Challenge: | Existing multimodal reasoning models lack generalized spatial intelligence, a new study shows . a critical gap exists in the field of vision-centric reasoning, the authors argue . |
| Approach: | They evaluate 16 multimodal reasoning models using Chain-of-Though (CoT) based thinking . they find that CoT prompting consistently degrades performance in visual spatial reasoning . |
| Outcome: | The proposed model hallucinates visual details from textual priors even when the image is absent. |
GraCoRe: Benchmarking Graph Comprehension and Complex Reasoning in Large Language Models (2025.coling-main)
Copied to clipboard
| Challenge: | Existing benchmarks focus primarily on pure graph understanding, lacking a comprehensive evaluation across all graph types and detailed capability definitions. |
| Approach: | They propose a benchmark to evaluate LLMs' graph comprehension and reasoning abilities using a three-tier hierarchical taxonomy and a granular taxonomies. |
| Outcome: | The proposed model includes 11 datasets with 5,140 graphs of varying complexity. |