Challenge: a new study explores the relationship between gestures and language . we use contrastive learning to learn gesture embeddings .
Approach: They adapt a semi-supervised multimodal model to learn gesture embeddings using Ted talks . they show gestures are predictive of the native language of the speaker .
Outcome: The proposed model learns gesture embeddings from a multimodal dataset . it shows that gesture embeds are predictive of the native language of the speaker .

Similar Papers

No Gestures Left Behind: Learning Relationships between Spoken Language and Freeform Gestures (2020.findings-emnlp)

Copied to clipboard

Challenge: a study of spoken language and co-speech gestures shows that it is important to model the long tail of the language-gesture distribution.
Approach: They propose a method that combines adversarial learning with importance sampling to strike a balance between precision and coverage.
Outcome: The proposed method outperforms state-of-the-art methods for gesture generation.
I see what you mean: Co-Speech Gestures for Reference Resolution in Multimodal Dialogue (2025.findings-acl)

Copied to clipboard

Challenge: Using representational co-speech gestures, face-to-face interaction participants resolve references to objects using speech and gestures.
Approach: They propose a multimodal reference resolution task centred on representational gestures . they propose 'self-supervised' pre-training approach to gesture representation learning that grounds body movements in spoken language.
Outcome: The proposed approach aligns with expert annotations and has significant predictive power.
A Formal Analysis of Multimodal Referring Strategies Under Common Ground (2020.lrec-1)

Copied to clipboard

Challenge: a recent study has focused on multimodality in the CL/NLP community, but it has not been widely studied.
Approach: They propose to analyze mixed-modality definite referring expressions using gestures and linguistic descriptions.
Outcome: The proposed models can predict viewer judgment of referring expressions and generate more natural and informative expressions.
Enhancing Spoken Discourse Modeling in Language Models Using Gestural Cues (2025.acl-long)

Copied to clipboard

Challenge: linguistic research shows that non-verbal cues, such as gestures, play a crucial role in spoken discourse.
Approach: They propose to integrate gestures into language models by embedding human motion sequences into discrete gesture tokens and aligning them with text embeddables.
Outcome: The proposed model improves on spoken discourse, the authors show . the study aims to improve the accuracy of discourse markers and quantifiers .
LMs stand their Ground: Investigating the Effect of Embodiment in Figurative Language Interpretation by Language Models (2023.findings-acl)

Copied to clipboard

Challenge: Figures are based on the use of words in a way that deviates from their conventional order and meaning.
Approach: They propose to use a figurative language model to interpret embodied metaphors by using larger language models that conceptualise embodies the action of the metaphorical sentence.
Outcome: The proposed model enables interpretation of figurative language when the action of the metaphorical sentence is more embodied.
Embodied Language Learning: Opportunities, Challenges, and Future Directions (2024.findings-acl)

Copied to clipboard

Challenge: embodied language learning is a form of language understanding where the language learner is situated in the world, perceives it, and interacts with it.
Approach: They propose to use a concept of World Scopes to measure progress in language understanding research.
Outcome: The proposed framework identifies gaps and suggests future directions for language understanding research.
A Probabilistic Model for Joint Learning of Word Embeddings from Texts and Images (D18-1)

Copied to clipboard

Challenge: Existing approaches combine language and perception to infer word embeddings . however, the embeddables produced by such models do not reflect the actual word representations.
Approach: They propose a probabilistic model that integrates linguistic and perceptual inputs to explain observed word-context pairs in a text corpus.
Outcome: The proposed model achieves competitive or stronger results on tasks of assessing pairwise word similarity and image/caption retrieval compared to other state-of-the-art models.
The Interplay between Metaphors and NLP (2026.acl-tutorials)

Copied to clipboard

Challenge: This tutorial will provide an overview of the metaphor processing field.
Approach: This tutorial will provide an overview of the metaphor processing field . it will focus on recent directions opened by LLMs for metaphor interpretation .
Outcome: The tutorial will discuss the influence of various metaphor theories on the creation of annotated resources and models.
Recognizing Multimodal Entailment (2021.acl-tutorials)

Copied to clipboard

Challenge: This tutorial introduces the multimodal entailment task for detecting semantic alignments . the task requires fine-grained understanding of visual and linguistic semantics questions .
Approach: This tutorial introduces the multimodal entailment task to machine learning . it introduces a dataset for recognizing multimodal alignments .
Outcome: This tutorial introduces the multimodal entailment task . it can be useful for detecting semantic alignments when a single modality alone is not enough .
Encoding Gesture in Multimodal Dialogue: Creating a Corpus of Multimodal AMR (2024.lrec-main)

Copied to clipboard

Challenge: Abstract Meaning Representation (AMR) was designed to represent sentence meaning in English text, but recent research has explored its adaptation to broader domains, including documents, dialogues, spatial information, cross-lingual tasks, and gesture.
Approach: They propose to annotate a multimodal (speech and gesture) AMR corpus in a task-based setting and capture coreference relationships across modalities.
Outcome: The proposed corpus captures coreference relationships across modalities, enabling fine-grained analysis of how gesture and natural language interact.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations