FoREST: Frame of Reference Evaluation in Spatial Reasoning Tasks (2025.emnlp-main)

Copied to clipboard

Challenge: Spatial reasoning is a fundamental aspect of human intelligence.
Approach: They propose a framework to assess FoR comprehension in large language models (LLMs) by using the Frame of Reference Evaluation in Spatial Reasoning Tasks benchmark.
Outcome: The proposed method improves overall performance across spatial reasoning tasks.

Similar Papers

Can LLMs See Without Pixels? Benchmarking Spatial Intelligence from Textual Descriptions (2026.findings-acl)

Copied to clipboard

Challenge: Existing advances in Spatial Intelligence rely on vision-Language Models . however, a critical question remains: does spatial understanding originate from visual encoders?
Approach: They propose to evaluate the SI performance of Large Language Models without pixel-level input.
Outcome: The proposed benchmark challenges large language models to perform symbolic reasoning rather than visual pattern matching.
A Benchmark for Reasoning with Spatial Prepositions (2023.emnlp-main)

Copied to clipboard

Challenge: Spatial reasoning is a fundamental building block of human cognition . large language models (LLMs) are not on par with advanced aspects of human cognitive domains .
Approach: They propose a benchmark to assess inferential properties of statements with spatial prepositions . they use prompt engineering to test the performance of two large language models .
Outcome: The proposed benchmark shows that none of the models reaches human performance.
Disentangling Extraction and Reasoning in Multi-hop Spatial Reasoning (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent studies highlight the struggles even large language models encounter when it comes to performing spatial reasoning over text.
Approach: They propose to disentangle spatial reasoning over text and compare them to state-of-the-art models with no explicit design for these parts.
Outcome: The proposed models show that they can perform spatial reasoning over text and can generalize within real data domains.
SPHERE: Unveiling Spatial Blind Spots in Vision-Language Models Through Hierarchical Evaluation (2025.acl-long)

Copied to clipboard

Challenge: Current vision-language models lack multi-dimensional spatial reasoning capabilities for human-like understanding and applications.
Approach: They propose a hierarchical evaluation framework that probes models across increasing levels of complexity and integrates spatial, visual, and logical understanding.
Outcome: The proposed framework probes models across increasing levels of complexity, from basic skills to multi-skill integration and high-level reasoning that combines spatial, visual, and logical understanding.
RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks (2025.emnlp-industry)

Copied to clipboard

Challenge: Existing evaluation methods do not explicitly measure this distinction, hindering effective dataset curation and real-world focused model development.
Approach: They introduce a region-based score to quantify a dataset's reliance on global versus local visual information.
Outcome: The proposed model-based score systematically compares model performance on image patches versus full images to determine if tasks require holistic image understanding or can be solved with partial or localized visual cues.
SpaRC and SpaRP: Spatial Reasoning Characterization and Path Generation for Understanding Spatial Reasoning Capability of Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Existing large language models (LLMs) do not perform well on the datasets.
Approach: They propose to use a Spatial Reasoning Characterization framework and a spatial reasoning path framework to study spatial reasoning.
Outcome: The proposed framework and datasets outperform state-of-the-art models in spatial reasoning.
KnowDR-REC: Auditing Knowledge-Conditioned Visual Grounding in Referring Expression Comprehension (2026.findings-acl)

Copied to clipboard

Challenge: Existing evaluation metrics suggest that Multimodal large language models have acquired fine-grained visual grounding capabilities.
Approach: They propose a benchmark to assess Referring Expression Comprehension (REC) that uses intra-image visual cues to localize target objects and a controllable evaluation mechanism to test sensitivity to fine-grained factual changes.
Outcome: The proposed benchmarks show that multimodal large language models have a high level of performance on the RefCOCO family of benchmarks.
Reasoning Paths with Reference Objects Elicit Quantitative Spatial Reasoning in Large Vision-Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Despite recent advances in visual language models, their ability to quantitatively reason about object sizes and distances remains underexplored.
Approach: They propose a manually annotated benchmark of 241 questions designed for quantitative spatial reasoning and a zero-shot prompting technique that encourages VLMs to use reference objects as visual cues.
Outcome: The proposed technique improves the performance of the top-performing VLMs by 19 points when a reasoning path using a reference object emerges naturally in the response.
Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs (2026.acl-short)

Copied to clipboard

Challenge: Existing multimodal reasoning models lack generalized spatial intelligence, a new study shows . a critical gap exists in the field of vision-centric reasoning, the authors argue .
Approach: They evaluate 16 multimodal reasoning models using Chain-of-Though (CoT) based thinking . they find that CoT prompting consistently degrades performance in visual spatial reasoning .
Outcome: The proposed model hallucinates visual details from textual priors even when the image is absent.
GraCoRe: Benchmarking Graph Comprehension and Complex Reasoning in Large Language Models (2025.coling-main)

Copied to clipboard

Challenge: Existing benchmarks focus primarily on pure graph understanding, lacking a comprehensive evaluation across all graph types and detailed capability definitions.
Approach: They propose a benchmark to evaluate LLMs' graph comprehension and reasoning abilities using a three-tier hierarchical taxonomy and a granular taxonomies.
Outcome: The proposed model includes 11 datasets with 5,140 graphs of varying complexity.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations