Defining and Evaluating Visual Language Models’ Basic Spatial Abilities: A Perspective from Psychometrics (2025.acl-long)
Copied to clipboard
| Challenge: | Existing studies assessing the spatial abilities of VLMs lack a solid theoretical foundation and lack measurable data. |
| Approach: | They propose a psychometric framework defining five basic spatial abilities in Visual Language Models. |
| Outcome: | The proposed framework defines five basic spatial abilities in Visual Language Models (VLMs) it provides a comprehensive evaluation benchmark and methodological perspective for embodied AI development . |
Similar Papers
Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models (2026.acl-srw)
Copied to clipboard
| Challenge: | Existing studies have focused on the ability of vision-language models to utilize spatial deictic expressions, which depend on the situation of utterance. |
| Approach: | They develop a benchmark to evaluate the multilingual ability of VLMs to use spatial deictic expressions in four languages. |
| Outcome: | The proposed models use demonstratives in a different manner from humans, particularly in selecting demonstrative based on distance from the object. |
Can LLMs See Without Pixels? Benchmarking Spatial Intelligence from Textual Descriptions (2026.findings-acl)
Copied to clipboard
Zhongbin Guo, Zhen Yang, Yushan Li, Xinyue Zhang, Wenyu Gao, Jiacheng Wang, Chengzhi Li, Xiangrui Liu, Ping Jian
| Challenge: | Existing advances in Spatial Intelligence rely on vision-Language Models . however, a critical question remains: does spatial understanding originate from visual encoders? |
| Approach: | They propose to evaluate the SI performance of Large Language Models without pixel-level input. |
| Outcome: | The proposed benchmark challenges large language models to perform symbolic reasoning rather than visual pattern matching. |
Diagnosing Spatial Consistency across Perspectives and Viewpoints in Large Vision-Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Existing models assess spatial capabilities from a static, single-view and egocentric perspective, failing to capture the dynamic nature of real-world spatial cognition. |
| Approach: | They propose a benchmark to diagnose spatial reasoning capabilities using a 360 field of view. |
| Outcome: | The proposed benchmark evaluates allocentric and egocentric reasoning capabilities from multiple perspectives in high-quality 3D environments. |
UrbanVideo-Bench: Benchmarking Vision-Language Models on Embodied Intelligence with Video Data in Urban Spaces (2025.acl-long)
Copied to clipboard
Baining Zhao, Jianjie Fang, Zichao Dai, Ziyou Wang, Jirong Zha, Weichen Zhang, Chen Gao, Yue Wang, Jinqiang Cui, Xinlei Chen, Yong Li
| Challenge: | Large multimodal models exhibit remarkable intelligence, yet their embodied cognitive abilities during motion in open-ended urban aerial spaces remain to be explored. |
| Approach: | They propose a benchmark to evaluate whether large multimodal models can process continuous first-person visual observations like humans. |
| Outcome: | The proposed model can process first-person visual observations like humans, enabling recall, perception, reasoning, and navigation. |
Exploring Spatial Schema Intuitions in Large Language and Vision Models (2024.findings-acl)
Copied to clipboard
| Challenge: | Large language models excel in varied NLP tasks, but lack a direct connection between sensory perception and physical action. |
| Approach: | They examine whether large language models capture implicit human intuitions about building blocks of language . they employ spatial cognitive foundations developed through early sensorimotor experiences . |
| Outcome: | The proposed model captures implicit human intuitions about building blocks of language without a tangible connection to embodied experiences. |
Visual Spatial Reasoning (2023.tacl-1)
Copied to clipboard
| Challenge: | Existing benchmarks for testing vision-language models (VLMs) are not ideal as they conflate multiple sources of error and do not allow controlled analysis on specific linguistic or cognitive properties. |
| Approach: | They present a dataset containing more than 10k natural text-image pairs with 66 types of spatial relations in English (e.g., under, in front of, facing). |
| Outcome: | The proposed model fails to capture relational information in a visual question answering task and referring expression comprehension tasks. |
Sparkle: Mastering Basic Spatial Capabilities in Vision Language Models Elicits Generalization to Spatial Reasoning (2025.findings-emnlp)
Copied to clipboard
Yihong Tang, Ao Qu, Zhaokai Wang, Dingyi Zhuang, Zhaofeng Wu, Wei Ma, Shenhao Wang, Yunhan Zheng, Zhan Zhao, Jinhua Zhao
| Challenge: | Currently, vision-language models excel in many downstream tasks but struggle with spatial reasoning, which is crucial for navigation and interaction with physical environments. |
| Approach: | They propose a framework that generates synthetic data to provide targeted supervision for VLMs across these basic spatial capabilities. |
| Outcome: | The proposed framework disentangles 2D spatial reasoning into three core components: direction comprehension, distance estimation, and localization. |
SPHERE: Unveiling Spatial Blind Spots in Vision-Language Models Through Hierarchical Evaluation (2025.acl-long)
Copied to clipboard
Wenyu Zhang, Wei En Ng, Lixin Ma, Yuwen Wang, Junqi Zhao, Allison Koenecke, Boyang Li, Wanglu Wanglu
| Challenge: | Current vision-language models lack multi-dimensional spatial reasoning capabilities for human-like understanding and applications. |
| Approach: | They propose a hierarchical evaluation framework that probes models across increasing levels of complexity and integrates spatial, visual, and logical understanding. |
| Outcome: | The proposed framework probes models across increasing levels of complexity, from basic skills to multi-skill integration and high-level reasoning that combines spatial, visual, and logical understanding. |
EmbSpatial-Bench: Benchmarking Spatial Understanding for Embodied Tasks with Large Vision-Language Models (2024.acl-short)
Copied to clipboard
| Challenge: | Recent studies have revealed significant deficiencies of LVLMs in understanding visual contents, leaving the gap between current embodied intelligence and large vision-language models (LVLM) . |
| Approach: | They propose to use a benchmark to evaluate LVLMs' spatial understanding of embodied environments to evaluate their ability to understand visual contents. |
| Outcome: | The proposed benchmark is derived from embodied scenes and covers 6 spatial relationships from an egocentric perspective. |
How Do LLMs and VLMs Understand Viewpoint Rotation Without Vision? An Interpretability Study (2026.acl-long)
Copied to clipboard
Zhen Yang, Ping Jian, Zhongbin Guo, Zuming Zhang, Chengzhi Li, Yonghong Deng, Xinyue Zhang, Wenpeng Lu
| Challenge: | Existing studies on spatial intelligence from the perspective of visual-spatial intelligence have not explored whether visual intelligence alone is sufficient to endow models with spatial intelligence. |
| Approach: | They propose to use a linguistic perspective to investigate spatial intelligence from a theoretical perspective. |
| Outcome: | The proposed model performs poorly on the proposed dataset while human can easily achieve 100% accuracy. |