SPARTQA: A Textual Question Answering Benchmark for Spatial Reasoning (2021.naacl-main)
Copied to clipboard
| Challenge: | Existing studies have focused on the spatial reasoning capabilities of modern language models (LMs) however, there has been limited research into the spatial thinking capabilities of LMs. |
| Approach: | They propose a question-answering (QA) benchmark for spatial reasoning on natural language text which contains more realistic spatial phenomena not covered by prior work. |
| Outcome: | The proposed method significantly improves LMs' ability on spatial understanding, which in turn helps solve two external datasets, bAbI, and boolQ. |
Similar Papers
A Benchmark for Reasoning with Spatial Prepositions (2023.emnlp-main)
Copied to clipboard
| Challenge: | Spatial reasoning is a fundamental building block of human cognition . large language models (LLMs) are not on par with advanced aspects of human cognitive domains . |
| Approach: | They propose a benchmark to assess inferential properties of statements with spatial prepositions . they use prompt engineering to test the performance of two large language models . |
| Outcome: | The proposed benchmark shows that none of the models reaches human performance. |
Automatic Generation of a Compositional QA Benchmark for Geospatial Reasoning under Spatial and Entity Constraints (2026.eacl-srw)
Copied to clipboard
| Challenge: | Recent advances in large language models have enhanced their ability to perform reasoning tasks that integrate linguistic, visual, and factual information. |
| Approach: | They propose a method for constructing compositional geographic question answering datasets that jointly consider spatial and entity constraints. |
| Outcome: | The proposed method performs well on questions involving rich entity grounding, but its accuracy drops on quantitative spatial reasoning questions. |
SpaRC and SpaRP: Spatial Reasoning Characterization and Path Generation for Understanding Spatial Reasoning Capability of Large Language Models (2024.acl-long)
Copied to clipboard
| Challenge: | Existing large language models (LLMs) do not perform well on the datasets. |
| Approach: | They propose to use a Spatial Reasoning Characterization framework and a spatial reasoning path framework to study spatial reasoning. |
| Outcome: | The proposed framework and datasets outperform state-of-the-art models in spatial reasoning. |
Representation, Learning and Reasoning on Spatial Language for Downstream NLP Tasks (2020.emnlp-tutorials)
Copied to clipboard
| Challenge: | In this tutorial, we discuss the cutting-edge research results and existing challenges related to spatial language understanding including semantic annotations, existing corpora, symbolic and sub-symbolic representations, qualitative spatial reasoning, spatial common sense, deep and structured learning models. |
| Approach: | This tutorial presents cutting-edge research results and current challenges related to spatial language understanding including semantic annotations, existing corpora, symbolic and sub-symbolic representations, qualitative spatial reasoning, spatial common sense, deep and structured learning models. |
| Outcome: | This paper reviews the cutting-edge research results and current challenges related to spatial language understanding including semantic annotations, existing corpora, symbolic and sub-symbolic representations, qualitative spatial reasoning, spatial common sense, deep and structured learning models. |
SpaRE: Enhancing Spatial Reasoning in Vision-Language Models with Synthetic Data (2025.acl-long)
Copied to clipboard
| Challenge: | Vision-language models struggle with spatial reasoning, a skill that humans excel at. |
| Approach: | They propose to use a spatial-reasoning Enhanced (SpaRE) VLM to improve spatial reasoning in visual question answering and robotics. |
| Outcome: | The proposed model achieves a 49% performance gain on the What's Up benchmark while maintaining strong results on general tasks. |
GeoSQA: A Benchmark for Scenario-based Question Answering in the Geography Domain at High School Level (D19-1)
Copied to clipboard
| Challenge: | SQA is an emerging application of NLP in the medical, geography, and legal domains. |
| Approach: | They propose a dataset of 1,981 scenarios and 4,110 multiple-choice questions in geography domain at high school level. |
| Outcome: | The proposed dataset consists of 1,981 scenarios and 4,110 multiple-choice questions in the geography domain at high school level. |
Neuro-symbolic Training for Reasoning over Spatial Language (2025.findings-naacl)
Copied to clipboard
| Challenge: | Spatial reasoning is essential for everyday human tasks and is crucial for robots to interact with their environment in a human-like manner. |
| Approach: | They propose to train language models to adhere to spatial reasoning rules as constraints . this allows them to capture the necessary level of abstraction for spatial reasoning . |
| Outcome: | The proposed technique improves language models in multi-hop spatial reasoning over text . it achieves higher accuracy than other competitive Spatial Question-answering benchmarks . |
FlowVQA: Mapping Multimodal Logic in Visual Question Answering with Flowcharts (2024.findings-acl)
Copied to clipboard
Shubhankar Singh, Purvi Chaurasia, Yerram Varun, Pranshu Pandya, Vatsal Gupta, Vivek Gupta, Dan Roth
| Challenge: | Existing benchmarks for visual question answering lack in visual grounding and complexity, particularly in evaluating spatial reasoning skills. |
| Approach: | They propose to use flowcharts as visual contexts to assess the capabilities of visual question-answering multimodal language models in reasoning. |
| Outcome: | The proposed benchmarks evaluate models' ability to follow visual information without pre-existing knowledge on a suite of open-source and proprietary multimodal language models using various strategies, followed by an analysis of directional bias. |
GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning (2021.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to solve geometric problems are dependent on handcraft rules and limited on small-scale datasets. |
| Approach: | They propose a Geometric Question Answering dataset with 5,010 geometric problems with corresponding annotated programs to illustrate the solving process. |
| Outcome: | The proposed method is significantly lower than human performance on the proposed dataset than on a publicly available dataset. |
SPaRC: A Spatial Pathfinding Reasoning Challenge (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing reasoning datasets saturate and fail to test abstract, multi-step problems, especially pathfinding and complex rule constraint satisfaction. |
| Approach: | They propose to use a spatial few-shot grid to evaluate spatial and rule-based reasoning with 1,000 2D grid puzzles. |
| Outcome: | The proposed model can be used to evaluate spatial reasoning and improve its accuracy. |