Challenge: Existing advances in Spatial Intelligence rely on vision-Language Models . however, a critical question remains: does spatial understanding originate from visual encoders?
Approach: They propose to evaluate the SI performance of Large Language Models without pixel-level input.
Outcome: The proposed benchmark challenges large language models to perform symbolic reasoning rather than visual pattern matching.

Similar Papers

How Do LLMs and VLMs Understand Viewpoint Rotation Without Vision? An Interpretability Study (2026.acl-long)

Copied to clipboard

Challenge: Existing studies on spatial intelligence from the perspective of visual-spatial intelligence have not explored whether visual intelligence alone is sufficient to endow models with spatial intelligence.
Approach: They propose to use a linguistic perspective to investigate spatial intelligence from a theoretical perspective.
Outcome: The proposed model performs poorly on the proposed dataset while human can easily achieve 100% accuracy.
UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces (2025.acl-long)

Copied to clipboard

Challenge: Large multimodal models exhibit remarkable intelligence, yet their embodied cognitive abilities during motion in open-ended urban aerial spaces remain to be explored.
Approach: They propose a benchmark to evaluate whether large multimodal models can process continuous first-person visual observations like humans.
Outcome: The proposed model can process first-person visual observations like humans, enabling recall, perception, reasoning, and navigation.
SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data (2025.acl-long)

Copied to clipboard

Challenge: Vision-language models struggle with spatial reasoning, a skill that humans excel at.
Approach: They propose to use a spatial-reasoning Enhanced (SpaRE) VLM to improve spatial reasoning in visual question answering and robotics.
Outcome: The proposed model achieves a 49% performance gain on the What's Up benchmark while maintaining strong results on general tasks.
EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models (2024.acl-short)

Copied to clipboard

Challenge: Recent studies have revealed significant deficiencies of LVLMs in understanding visual contents, leaving the gap between current embodied intelligence and large vision-language models (LVLM) .
Approach: They propose to use a benchmark to evaluate LVLMs' spatial understanding of embodied environments to evaluate their ability to understand visual contents.
Outcome: The proposed benchmark is derived from embodied scenes and covers 6 spatial relationships from an egocentric perspective.
Can LLMs Learn to Map the World from Local Descriptions? (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in large language models have demonstrated strong capabilities in tasks such as code generation and mathematical reasoning.
Approach: They investigate whether large language models can construct coherent global spatial cognition by integrating fragmented relational descriptions.
Outcome: The proposed models can generalize to unseen spatial relationships and exhibit latent representations aligned with real-world spatial distributions.
Exploring Spatial Schema Intuitions in Large Language and Vision Models (2024.findings-acl)

Copied to clipboard

Challenge: Large language models excel in varied NLP tasks, but lack a direct connection between sensory perception and physical action.
Approach: They examine whether large language models capture implicit human intuitions about building blocks of language . they employ spatial cognitive foundations developed through early sensorimotor experiences .
Outcome: The proposed model captures implicit human intuitions about building blocks of language without a tangible connection to embodied experiences.
LLM Meets Scene Graph: Can Large Language Models Understand and Generate Scene Graphs? A Benchmark and Empirical Study (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have demonstrated impressive progress in various text-based tasks, such as question-answering and content generation.
Approach: They propose a benchmark to evaluate Large Language Models’ ability to understand scene graphs and generate them from textual narratives.
Outcome: The proposed model performs well on scene graph understanding but struggles with scene graph generation, particularly for complex narratives.
Where the Cat Sat: A Multilingual Framework for Spatial Language Understanding (2026.acl-long)

Copied to clipboard

Challenge: Existing work exhibits biases toward English and prepositional marking . Existing models are limited in understanding spatial relations across typologically diverse languages .
Approach: They propose a multilingual framework and benchmark for spatial language understanding . they decompose spatial relations into surface elements and semantic components . their results suggest surface parsing does not entail spatial understanding - they argue .
Outcome: The proposed framework and benchmark decomposes spatial relations into surface elements and semantic components.
Things not Written in Text: Exploring Spatial Commonsense from Visual Signals (2022.acl-long)

Copied to clipboard

Challenge: Pretrained language models fail in many NLP tasks, but are ineffective in spatial commonsense reasoning.
Approach: They propose a spatial commonsense benchmark that focuses on relative scales of objects and the positional relationship between people and objects under different actions.
Outcome: The proposed framework outperforms pretrained models in answering spatial questions.
SemVink: Advancing VLMs’ Semantic Understanding of Optical Illusions via Visual Global Thinking (2025.emnlp-main)

Copied to clipboard

Challenge: Vision-language models excel in semantic tasks but fail at detecting hidden content . current architectures prioritize abstract reasoning over low-level visual operations .
Approach: They propose a benchmark to test vision-language models that can detect hidden content . they propose HC-Bench to scale images to low resolutions to unlock 99% accuracy .
Outcome: HC-Bench shows that leading VLMs achieve near-zero accuracy even with explicit prompting . et al.: current models prioritize abstract reasoning over low-level visual operations . they urge a shift toward hybrid models bridging gap between computational vision and human cognition .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations