Challenge: Existing studies have focused on the spatial reasoning capabilities of modern language models (LMs) however, there has been limited research into the spatial thinking capabilities of LMs.
Approach: They propose a question-answering (QA) benchmark for spatial reasoning on natural language text which contains more realistic spatial phenomena not covered by prior work.
Outcome: The proposed method significantly improves LMs' ability on spatial understanding, which in turn helps solve two external datasets, bAbI, and boolQ.

Similar Papers

A Benchmark for Reasoning with Spatial Prepositions (2023.emnlp-main)

Copied to clipboard

Challenge: Spatial reasoning is a fundamental building block of human cognition . large language models (LLMs) are not on par with advanced aspects of human cognitive domains .
Approach: They propose a benchmark to assess inferential properties of statements with spatial prepositions . they use prompt engineering to test the performance of two large language models .
Outcome: The proposed benchmark shows that none of the models reaches human performance.
Automatic Generation of a Compositional QA Benchmark for Geospatial Reasoning under Spatial and Entity Constraints (2026.eacl-srw)

Copied to clipboard

Challenge: Recent advances in large language models have enhanced their ability to perform reasoning tasks that integrate linguistic, visual, and factual information.
Approach: They propose a method for constructing compositional geographic question answering datasets that jointly consider spatial and entity constraints.
Outcome: The proposed method performs well on questions involving rich entity grounding, but its accuracy drops on quantitative spatial reasoning questions.
SpaRC and SpaRP: Spatial Reasoning Characterization and Path Generation for Understanding Spatial Reasoning Capability of Large Language Models (2024.acl-long)

Copied to clipboard

Challenge: Existing large language models (LLMs) do not perform well on the datasets.
Approach: They propose to use a Spatial Reasoning Characterization framework and a spatial reasoning path framework to study spatial reasoning.
Outcome: The proposed framework and datasets outperform state-of-the-art models in spatial reasoning.
Representation, Learning and Reasoning on Spatial Language for Downstream NLP Tasks (2020.emnlp-tutorials)

Copied to clipboard

Challenge: In this tutorial, we discuss the cutting-edge research results and existing challenges related to spatial language understanding including semantic annotations, existing corpora, symbolic and sub-symbolic representations, qualitative spatial reasoning, spatial common sense, deep and structured learning models.
Approach: This tutorial presents cutting-edge research results and current challenges related to spatial language understanding including semantic annotations, existing corpora, symbolic and sub-symbolic representations, qualitative spatial reasoning, spatial common sense, deep and structured learning models.
Outcome: This paper reviews the cutting-edge research results and current challenges related to spatial language understanding including semantic annotations, existing corpora, symbolic and sub-symbolic representations, qualitative spatial reasoning, spatial common sense, deep and structured learning models.
SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data (2025.acl-long)

Copied to clipboard

Challenge: Vision-language models struggle with spatial reasoning, a skill that humans excel at.
Approach: They propose to use a spatial-reasoning Enhanced (SpaRE) VLM to improve spatial reasoning in visual question answering and robotics.
Outcome: The proposed model achieves a 49% performance gain on the What's Up benchmark while maintaining strong results on general tasks.
GeoSQA: A Benchmark for Scenario-based Question Answering in the Geography Domain at High School Level (D19-1)

Copied to clipboard

Challenge: SQA is an emerging application of NLP in the medical, geography, and legal domains.
Approach: They propose a dataset of 1,981 scenarios and 4,110 multiple-choice questions in geography domain at high school level.
Outcome: The proposed dataset consists of 1,981 scenarios and 4,110 multiple-choice questions in the geography domain at high school level.
Neuro-symbolic Training for Reasoning over Spatial Language (2025.findings-naacl)

Copied to clipboard

Challenge: Spatial reasoning is essential for everyday human tasks and is crucial for robots to interact with their environment in a human-like manner.
Approach: They propose to train language models to adhere to spatial reasoning rules as constraints . this allows them to capture the necessary level of abstraction for spatial reasoning .
Outcome: The proposed technique improves language models in multi-hop spatial reasoning over text . it achieves higher accuracy than other competitive Spatial Question-answering benchmarks .
FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts (2024.findings-acl)

Copied to clipboard

Challenge: Existing benchmarks for visual question answering lack in visual grounding and complexity, particularly in evaluating spatial reasoning skills.
Approach: They propose to use flowcharts as visual contexts to assess the capabilities of visual question-answering multimodal language models in reasoning.
Outcome: The proposed benchmarks evaluate models' ability to follow visual information without pre-existing knowledge on a suite of open-source and proprietary multimodal language models using various strategies, followed by an analysis of directional bias.
GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning (2021.findings-acl)

Copied to clipboard

Challenge: Existing methods to solve geometric problems are dependent on handcraft rules and limited on small-scale datasets.
Approach: They propose a Geometric Question Answering dataset with 5,010 geometric problems with corresponding annotated programs to illustrate the solving process.
Outcome: The proposed method is significantly lower than human performance on the proposed dataset than on a publicly available dataset.
SPaRC: A Spatial Pathfinding Reasoning Challenge (2025.emnlp-main)

Copied to clipboard

Challenge: Existing reasoning datasets saturate and fail to test abstract, multi-step problems, especially pathfinding and complex rule constraint satisfaction.
Approach: They propose to use a spatial few-shot grid to evaluate spatial and rule-based reasoning with 1,000 2D grid puzzles.
Outcome: The proposed model can be used to evaluate spatial reasoning and improve its accuracy.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations