Challenge: Using a model architecture, we decode text to 2D spatial arrangements in a multi-object and multi-relationship setting.
Approach: They propose a model architecture Spatial-Reasoning Bert that decodes language to 2D spatial arrangements in a multi-object and multi-relationship setting.
Outcome: The proposed model can generate complete abstract scenes if paired with a clip-arts predictor and can generalize to out-of-sample data to a reasonable extent.

Similar Papers

From Spatial Relations to Spatial Configurations (2020.lrec-1)

Copied to clipboard

Challenge: Existing spatial representations are not sufficient for describing complex spatial configurations.
Approach: They propose to integrate existing spatial representation languages with an annotation schema to extend the capabilities of existing ones.
Outcome: The proposed language can represent a large set of spatial concepts crucial for reasoning . it integrates with the Abstract Meaning Representation (AMR) annotation schema and annotates text from diverse datasets .
Representation, Learning and Reasoning on Spatial Language for Downstream NLP Tasks (2020.emnlp-tutorials)

Copied to clipboard

Challenge: In this tutorial, we discuss the cutting-edge research results and existing challenges related to spatial language understanding including semantic annotations, existing corpora, symbolic and sub-symbolic representations, qualitative spatial reasoning, spatial common sense, deep and structured learning models.
Approach: This tutorial presents cutting-edge research results and current challenges related to spatial language understanding including semantic annotations, existing corpora, symbolic and sub-symbolic representations, qualitative spatial reasoning, spatial common sense, deep and structured learning models.
Outcome: This paper reviews the cutting-edge research results and current challenges related to spatial language understanding including semantic annotations, existing corpora, symbolic and sub-symbolic representations, qualitative spatial reasoning, spatial common sense, deep and structured learning models.
Can Multimodal Large Language Models Understand Spatial Relations? (2025.acl-long)

Copied to clipboard

Challenge: Spatial relation reasoning is a crucial task for multimodal large language models to understand the objective world.
Approach: They propose a human-annotated spatial relation reasoning benchmark based on COCO2017 to improve MLLMs' spatial relation thinking.
Outcome: The proposed benchmark achieves 48.14% accuracy, far below the human-level accuracy of 98.40%.
Spatial and Temporal Language Understanding: Representation, Reasoning, and Grounding (2024.naacl-tutorials)

Copied to clipboard

Challenge: This tutorial provides an overview of cutting edge research on spatial and temporal language understanding.
Approach: This tutorial provides an overview of cutting edge research on spatial and temporal language understanding.
Outcome: This tutorial provides an overview of cutting edge research on spatial and temporal language understanding.
A Corpus of Natural Multimodal Spatial Scene Descriptions (L18-1)

Copied to clipboard

Challenge: Existing work on multimodal spatial descriptions combines speech and hand gestures to form a corpus of multimodal descriptions.
Approach: They present a corpus of multimodal spatial descriptions as commonly occurring in route giving tasks.
Outcome: The proposed corpus of multimodal spatial descriptions is more amenable to computational analysis and useable for learning natural computer interfaces.
Finding Structural Knowledge in Multimodal-BERT (2022.acl-long)

Copied to clipboard

Challenge: Several multimodal-BERT models learn contextualized embeddings through training on linguistic data and visual data.
Approach: They propose to make the structure of language and visuals explicit by a dependency parse . they also propose to encode the scene tree in the multimodal-BERT models .
Outcome: The proposed models do not encode the scene trees in the language description.
Understanding Spatial Relations through Multiple Modalities (2020.lrec-1)

Copied to clipboard

Challenge: Existing work on common sense reasoning and understanding of spatial relations is limited.
Approach: They propose a spatial model that uses both textual and visual information to predict spatial relations between two entities in an image.
Outcome: The proposed model improves prediction accuracy and coverage and deals with unseen subjects, objects and relations.
Linear Relational Decoding of Morphology in Language Models (2025.naacl-srw)

Copied to clipboard

Challenge: Recent work has shown that affine transformations on subject representations can faithfully approximate model outputs for certain subject-object relations.
Approach: They propose to use affine transformations to adapt the Bigger Analogy Test Set to test faithfulness of morphological relations.
Outcome: The proposed method achieves 90% faithfulness on morphological relations, with similar findings across languages and models.
Transfer Learning with Synthetic Corpora for Spatial Role Labeling and Reasoning (2022.emnlp-main)

Copied to clipboard

Challenge: Existing datasets on spatial language processing are either synthetic or at small scale.
Approach: They propose a dataset for transfer learning on spatial question answering and spatial role labeling that includes a larger variety of spatial relation types and spatial expressions.
Outcome: The proposed dataset can be used to evaluate spatial language processing models in real-world situations.
Can LLMs Learn to Map the World from Local Descriptions? (2026.acl-long)

Copied to clipboard

Challenge: Recent advances in large language models have demonstrated strong capabilities in tasks such as code generation and mathematical reasoning.
Approach: They investigate whether large language models can construct coherent global spatial cognition by integrating fragmented relational descriptions.
Outcome: The proposed models can generalize to unseen spatial relationships and exhibit latent representations aligned with real-world spatial distributions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations