Decoding Language Spatial Relations to 2D Spatial Arrangements (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Using a model architecture, we decode text to 2D spatial arrangements in a multi-object and multi-relationship setting. |
| Approach: | They propose a model architecture Spatial-Reasoning Bert that decodes language to 2D spatial arrangements in a multi-object and multi-relationship setting. |
| Outcome: | The proposed model can generate complete abstract scenes if paired with a clip-arts predictor and can generalize to out-of-sample data to a reasonable extent. |
Similar Papers
From Spatial Relations to Spatial Configurations (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing spatial representations are not sufficient for describing complex spatial configurations. |
| Approach: | They propose to integrate existing spatial representation languages with an annotation schema to extend the capabilities of existing ones. |
| Outcome: | The proposed language can represent a large set of spatial concepts crucial for reasoning . it integrates with the Abstract Meaning Representation (AMR) annotation schema and annotates text from diverse datasets . |
Representation, Learning and Reasoning on Spatial Language for Downstream NLP Tasks (2020.emnlp-tutorials)
Copied to clipboard
| Challenge: | In this tutorial, we discuss the cutting-edge research results and existing challenges related to spatial language understanding including semantic annotations, existing corpora, symbolic and sub-symbolic representations, qualitative spatial reasoning, spatial common sense, deep and structured learning models. |
| Approach: | This tutorial presents cutting-edge research results and current challenges related to spatial language understanding including semantic annotations, existing corpora, symbolic and sub-symbolic representations, qualitative spatial reasoning, spatial common sense, deep and structured learning models. |
| Outcome: | This paper reviews the cutting-edge research results and current challenges related to spatial language understanding including semantic annotations, existing corpora, symbolic and sub-symbolic representations, qualitative spatial reasoning, spatial common sense, deep and structured learning models. |
Can Multimodal Large Language Models Understand Spatial Relations? (2025.acl-long)
Copied to clipboard
| Challenge: | Spatial relation reasoning is a crucial task for multimodal large language models to understand the objective world. |
| Approach: | They propose a human-annotated spatial relation reasoning benchmark based on COCO2017 to improve MLLMs' spatial relation thinking. |
| Outcome: | The proposed benchmark achieves 48.14% accuracy, far below the human-level accuracy of 98.40%. |
Spatial and Temporal Language Understanding: Representation, Reasoning, and Grounding (2024.naacl-tutorials)
Copied to clipboard
| Challenge: | This tutorial provides an overview of cutting edge research on spatial and temporal language understanding. |
| Approach: | This tutorial provides an overview of cutting edge research on spatial and temporal language understanding. |
| Outcome: | This tutorial provides an overview of cutting edge research on spatial and temporal language understanding. |
A Corpus of Natural Multimodal Spatial Scene Descriptions (L18-1)
Copied to clipboard
| Challenge: | Existing work on multimodal spatial descriptions combines speech and hand gestures to form a corpus of multimodal descriptions. |
| Approach: | They present a corpus of multimodal spatial descriptions as commonly occurring in route giving tasks. |
| Outcome: | The proposed corpus of multimodal spatial descriptions is more amenable to computational analysis and useable for learning natural computer interfaces. |
Finding Structural Knowledge in Multimodal-BERT (2022.acl-long)
Copied to clipboard
| Challenge: | Several multimodal-BERT models learn contextualized embeddings through training on linguistic data and visual data. |
| Approach: | They propose to make the structure of language and visuals explicit by a dependency parse . they also propose to encode the scene tree in the multimodal-BERT models . |
| Outcome: | The proposed models do not encode the scene trees in the language description. |
Understanding Spatial Relations through Multiple Modalities (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing work on common sense reasoning and understanding of spatial relations is limited. |
| Approach: | They propose a spatial model that uses both textual and visual information to predict spatial relations between two entities in an image. |
| Outcome: | The proposed model improves prediction accuracy and coverage and deals with unseen subjects, objects and relations. |
Linear Relational Decoding of Morphology in Language Models (2025.naacl-srw)
Copied to clipboard
| Challenge: | Recent work has shown that affine transformations on subject representations can faithfully approximate model outputs for certain subject-object relations. |
| Approach: | They propose to use affine transformations to adapt the Bigger Analogy Test Set to test faithfulness of morphological relations. |
| Outcome: | The proposed method achieves 90% faithfulness on morphological relations, with similar findings across languages and models. |
Transfer Learning with Synthetic Corpora for Spatial Role Labeling and Reasoning (2022.emnlp-main)
Copied to clipboard
| Challenge: | Existing datasets on spatial language processing are either synthetic or at small scale. |
| Approach: | They propose a dataset for transfer learning on spatial question answering and spatial role labeling that includes a larger variety of spatial relation types and spatial expressions. |
| Outcome: | The proposed dataset can be used to evaluate spatial language processing models in real-world situations. |
Can LLMs Learn to Map the World from Local Descriptions? (2026.acl-long)
Copied to clipboard
| Challenge: | Recent advances in large language models have demonstrated strong capabilities in tasks such as code generation and mathematical reasoning. |
| Approach: | They investigate whether large language models can construct coherent global spatial cognition by integrating fragmented relational descriptions. |
| Outcome: | The proposed models can generalize to unseen spatial relationships and exhibit latent representations aligned with real-world spatial distributions. |