| Challenge: | Existing work on multimodal spatial descriptions combines speech and hand gestures to form a corpus of multimodal descriptions. |
| Approach: | They present a corpus of multimodal spatial descriptions as commonly occurring in route giving tasks. |
| Outcome: | The proposed corpus of multimodal spatial descriptions is more amenable to computational analysis and useable for learning natural computer interfaces. |
Similar Papers
A Formal Analysis of Multimodal Referring Strategies Under Common Ground (2020.lrec-1)
Copied to clipboard
| Challenge: | a recent study has focused on multimodality in the CL/NLP community, but it has not been widely studied. |
| Approach: | They propose to analyze mixed-modality definite referring expressions using gestures and linguistic descriptions. |
| Outcome: | The proposed models can predict viewer judgment of referring expressions and generate more natural and informative expressions. |
Action Verb Corpus (L18-1)
Copied to clipboard
| Challenge: | a corpus of 390 simple actions is based on multimodal data of 12 humans . the dataset is annotated with orthographic transcriptions of utterances and part-of-speech tags . |
| Approach: | They present a multimodal corpus of 12 humans performing 390 simple actions . they also propose an algorithm for segmenting words into utterances and aligning visual information and speech . |
| Outcome: | The presented dataset includes 390 simple actions performed by 12 humans . it includes transcriptions of utterances, part-of-speech tags, lemmata, and hand touches . |
Encoding Gesture in Multimodal Dialogue: Creating a Corpus of Multimodal AMR (2024.lrec-main)
Copied to clipboard
| Challenge: | Abstract Meaning Representation (AMR) was designed to represent sentence meaning in English text, but recent research has explored its adaptation to broader domains, including documents, dialogues, spatial information, cross-lingual tasks, and gesture. |
| Approach: | They propose to annotate a multimodal (speech and gesture) AMR corpus in a task-based setting and capture coreference relationships across modalities. |
| Outcome: | The proposed corpus captures coreference relationships across modalities, enabling fine-grained analysis of how gesture and natural language interact. |
Multimodality for NLP-Centered Applications: Resources, Advances and Frontiers (2022.lrec-1)
Copied to clipboard
| Challenge: | resurgence of multimodal datasets has attracted significant research interest, but there is no comprehensive survey for this task. |
| Approach: | They present a survey of a multimodal dataset with different modalities according to the applications. |
| Outcome: | The proposed datasets are available online and discuss the new frontier and motivate future researches. |
Modeling Collaborative Multimodal Behavior in Group Dialogues: The MULTISIMO Corpus (L18-1)
Copied to clipboard
| Challenge: | a corpus of human-computer interactions recorded in multiple modalities is being developed to study and model collaborative aspects of multimodal behavior in groups. |
| Approach: | They propose to use a multimodal corpus to investigate collaborative aspects of multimodal behavior in groups that perform simple tasks. |
| Outcome: | The proposed corpus is designed for public release and includes survey materials, personality tests and experience assessment questionnaires filled in by all participants. |
Representation, Learning and Reasoning on Spatial Language for Downstream NLP Tasks (2020.emnlp-tutorials)
Copied to clipboard
| Challenge: | In this tutorial, we discuss the cutting-edge research results and existing challenges related to spatial language understanding including semantic annotations, existing corpora, symbolic and sub-symbolic representations, qualitative spatial reasoning, spatial common sense, deep and structured learning models. |
| Approach: | This tutorial presents cutting-edge research results and current challenges related to spatial language understanding including semantic annotations, existing corpora, symbolic and sub-symbolic representations, qualitative spatial reasoning, spatial common sense, deep and structured learning models. |
| Outcome: | This paper reviews the cutting-edge research results and current challenges related to spatial language understanding including semantic annotations, existing corpora, symbolic and sub-symbolic representations, qualitative spatial reasoning, spatial common sense, deep and structured learning models. |
Recognition of Implicit Geographic Movement in Text (2020.lrec-1)
Copied to clipboard
| Challenge: | a growing field of research is analyzing the geographic movement of humans, animals, and other entities. |
| Approach: | They created a corpus of sentences labeled as describing geographic movement or not . they used hand labeling, crowd voting and machine learning to predict more labels . |
| Outcome: | a new method uses hand labeling, crowd voting and machine learning to predict more labels. |
Decoding Language Spatial Relations to 2D Spatial Arrangements (2020.findings-emnlp)
Copied to clipboard
| Challenge: | Using a model architecture, we decode text to 2D spatial arrangements in a multi-object and multi-relationship setting. |
| Approach: | They propose a model architecture Spatial-Reasoning Bert that decodes language to 2D spatial arrangements in a multi-object and multi-relationship setting. |
| Outcome: | The proposed model can generate complete abstract scenes if paired with a clip-arts predictor and can generalize to out-of-sample data to a reasonable extent. |
Evaluation of Multilingual Ability to Use Spatial Deictic Expressions in Vision-Language Models (2026.acl-srw)
Copied to clipboard
| Challenge: | Existing studies have focused on the ability of vision-language models to utilize spatial deictic expressions, which depend on the situation of utterance. |
| Approach: | They develop a benchmark to evaluate the multilingual ability of VLMs to use spatial deictic expressions in four languages. |
| Outcome: | The proposed models use demonstratives in a different manner from humans, particularly in selecting demonstrative based on distance from the object. |
I see what you mean: Co-Speech Gestures for Reference Resolution in Multimodal Dialogue (2025.findings-acl)
Copied to clipboard
| Challenge: | Using representational co-speech gestures, face-to-face interaction participants resolve references to objects using speech and gestures. |
| Approach: | They propose a multimodal reference resolution task centred on representational gestures . they propose 'self-supervised' pre-training approach to gesture representation learning that grounds body movements in spoken language. |
| Outcome: | The proposed approach aligns with expert annotations and has significant predictive power. |