Challenge: Existing work exhibits biases toward English and prepositional marking . Existing models are limited in understanding spatial relations across typologically diverse languages .
Approach: They propose a multilingual framework and benchmark for spatial language understanding . they decompose spatial relations into surface elements and semantic components . their results suggest surface parsing does not entail spatial understanding - they argue .
Outcome: The proposed framework and benchmark decomposes spatial relations into surface elements and semantic components.

Similar Papers

Can LLMs See Without Pixels? Benchmarking Spatial Intelligence from Textual Descriptions (2026.findings-acl)

Copied to clipboard

Challenge: Existing advances in Spatial Intelligence rely on vision-Language Models . however, a critical question remains: does spatial understanding originate from visual encoders?
Approach: They propose to evaluate the SI performance of Large Language Models without pixel-level input.
Outcome: The proposed benchmark challenges large language models to perform symbolic reasoning rather than visual pattern matching.
Spatial and Temporal Language Understanding: Representation, Reasoning, and Grounding (2024.naacl-tutorials)

Copied to clipboard

Challenge: This tutorial provides an overview of cutting edge research on spatial and temporal language understanding.
Approach: This tutorial provides an overview of cutting edge research on spatial and temporal language understanding.
Outcome: This tutorial provides an overview of cutting edge research on spatial and temporal language understanding.
Can Multimodal Large Language Models Understand Spatial Relations? (2025.acl-long)

Copied to clipboard

Challenge: Spatial relation reasoning is a crucial task for multimodal large language models to understand the objective world.
Approach: They propose a human-annotated spatial relation reasoning benchmark based on COCO2017 to improve MLLMs' spatial relation thinking.
Outcome: The proposed benchmark achieves 48.14% accuracy, far below the human-level accuracy of 98.40%.
Representation, Learning and Reasoning on Spatial Language for Downstream NLP Tasks (2020.emnlp-tutorials)

Copied to clipboard

Challenge: In this tutorial, we discuss the cutting-edge research results and existing challenges related to spatial language understanding including semantic annotations, existing corpora, symbolic and sub-symbolic representations, qualitative spatial reasoning, spatial common sense, deep and structured learning models.
Approach: This tutorial presents cutting-edge research results and current challenges related to spatial language understanding including semantic annotations, existing corpora, symbolic and sub-symbolic representations, qualitative spatial reasoning, spatial common sense, deep and structured learning models.
Outcome: This paper reviews the cutting-edge research results and current challenges related to spatial language understanding including semantic annotations, existing corpora, symbolic and sub-symbolic representations, qualitative spatial reasoning, spatial common sense, deep and structured learning models.
Large-Scale Bitext Corpora Provide New Evidence for Cognitive Representations of Spatial Terms (2024.eacl-long)

Copied to clipboard

Challenge: Recent evidence suggests that there exist two classes of cognitive representations within the spatial terms of a language.
Approach: They propose a pipeline for extracting, isolating, and aligning spatial terms from parallel text . they find evidence that variability in functional terms differs significantly from that of geometric terms .
Outcome: The proposed pipeline extracts, isolates, and aligns spatial terms in basic locative constructions from parallel text.
A Benchmark for Reasoning with Spatial Prepositions (2023.emnlp-main)

Copied to clipboard

Challenge: Spatial reasoning is a fundamental building block of human cognition . large language models (LLMs) are not on par with advanced aspects of human cognitive domains .
Approach: They propose a benchmark to assess inferential properties of statements with spatial prepositions . they use prompt engineering to test the performance of two large language models .
Outcome: The proposed benchmark shows that none of the models reaches human performance.
SpatialWebAgent: Leveraging Large Language Models for Automated Spatial Information Extraction and Map Grounding (2025.acl-demo)

Copied to clipboard

Challenge: Understanding and extracting spatial information from text is vital for a wide range of applications, says nielsen . inherent complexity of geographic expressions in natural language presents significant hurdles for traditional extraction methods.
Approach: They propose a system that leverages large language models to extract spatial information from natural language.
Outcome: SpatialWebAgent is designed to extract, standardize, and ground spatial information from natural language text directly onto maps.
Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models (2026.acl-srw)

Copied to clipboard

Challenge: Existing studies have focused on the ability of vision-language models to utilize spatial deictic expressions, which depend on the situation of utterance.
Approach: They develop a benchmark to evaluate the multilingual ability of VLMs to use spatial deictic expressions in four languages.
Outcome: The proposed models use demonstratives in a different manner from humans, particularly in selecting demonstrative based on distance from the object.
Visual Spatial Reasoning (2023.tacl-1)

Copied to clipboard

Challenge: Existing benchmarks for testing vision-language models (VLMs) are not ideal as they conflate multiple sources of error and do not allow controlled analysis on specific linguistic or cognitive properties.
Approach: They present a dataset containing more than 10k natural text-image pairs with 66 types of spatial relations in English (e.g., under, in front of, facing).
Outcome: The proposed model fails to capture relational information in a visual question answering task and referring expression comprehension tasks.
What’s “up” with vision-language models? Investigating their struggle with spatial reasoning (2023.emnlp-main)

Copied to clipboard

Challenge: Recent work has re-surfaced a concern that has long plagued vision-language models: poor performance on simple tasks like attribute attachment, counting, etc.
Approach: They evaluate 18 vision-language models and find they perform poorly on VQAv2 . they find that popular vision-linguistic pretraining corpora lack reliable data for learning spatial relationships .
Outcome: The new models are compared with existing datasets on what'sup and visual-language models . they achieve 56% accuracy on the new benchmarks compared to 99% for humans .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations