Challenge: a key missing step is to explore whether the nonverbal information can be quantified.
Approach: They explore whether incorporating gesture representations can improve the language model’s performance . they also examine whether spontaneous gestures demonstrate entropy rate constancy (ERC) .
Outcome: The proposed model improves the performance of the mixed-modal language models against monologue video data.

Similar Papers

Gestures Are Used Rationally: Information Theoretic Evidence from Neural Sequential Models (2022.coling-1)

Copied to clipboard

Challenge: Verbal communication is companied by rich non-verbal signals, but few studies have explored the non- verbal channels with finer theoretical lens.
Approach: They extract gesture representations from monologue video data and train neural sequential models to examine their results.
Outcome: The proposed method shows that speakers use simple gestures to convey information that enhances verbal communication.
Towards Understanding the Relation between Gestures and Language (2022.coling-1)

Copied to clipboard

Challenge: a new study explores the relationship between gestures and language . we use contrastive learning to learn gesture embeddings .
Approach: They adapt a semi-supervised multimodal model to learn gesture embeddings using Ted talks . they show gestures are predictive of the native language of the speaker .
Outcome: The proposed model learns gesture embeddings from a multimodal dataset . it shows that gesture embeds are predictive of the native language of the speaker .
No Gestures Left Behind: Learning Relationships between Spoken Language and Freeform Gestures (2020.findings-emnlp)

Copied to clipboard

Challenge: a study of spoken language and co-speech gestures shows that it is important to model the long tail of the language-gesture distribution.
Approach: They propose a method that combines adversarial learning with importance sampling to strike a balance between precision and coverage.
Outcome: The proposed method outperforms state-of-the-art methods for gesture generation.
Towards Comprehensive Language Analysis for Clinically Enriched Spontaneous Dialogue (2024.lrec-main)

Copied to clipboard

Challenge: Contemporary NLP has progressed from feature-based classification to fine-tuning and prompt-based techniques . many of these techniques remain understudied in the context of real-world, clinically enriched spontaneous dialogue.
Approach: They investigate the efficacy and overall performance of a range of NLP techniques on transcribed speech from patients with schizophrenia and other disorders.
Outcome: The proposed methods are effective in analyzing transcribed speech from patients with schizophrenia and healthy controls taking a clinically-validated language test.
LLM Knows Body Language, Too: Translating Speech Voices into Human Gestures (2024.acl-long)

Copied to clipboard

Challenge: despite advances in the generation of realistic human gestures, the process often includes unintended, meaningless, or non-realistic gestures.
Approach: They propose a framework that leverages large language models to generate human gestures . the primary stage employs a transformer-based auto-encoder network to encode human gesture into discrete symbols .
Outcome: The proposed framework has demonstrated state-of-the-art performance on public TED and TED-Expressive datasets.
Enhancing Spoken Discourse Modeling in Language Models Using Gestural Cues (2025.acl-long)

Copied to clipboard

Challenge: linguistic research shows that non-verbal cues, such as gestures, play a crucial role in spoken discourse.
Approach: They propose to integrate gestures into language models by embedding human motion sequences into discrete gesture tokens and aligning them with text embeddables.
Outcome: The proposed model improves on spoken discourse, the authors show . the study aims to improve the accuracy of discourse markers and quantifiers .
SMILEE: Symmetric Multi-modal Interactions with Language-gesture Enabled (AI) Embodiment (N18-5)

Copied to clipboard

Challenge: SMILEE is a conversational agent system that interprets a user’s communicative intent from verbal utterances and non-verbal behaviors, such as gestures.
Approach: They propose to use a computer-generated avatar to embody a human-machine conversational agent system that interprets verbal utterances and non-verbal behaviors to facilitate natural symmetric multi-modal interactions.
Outcome: The proposed system interprets a user’s communicative intent from verbal utterances and non-verbal behaviors, such as gestures, and communicates with natural language and gestures through its embodiment as an avatar.
Do Multimodal Large Language Models Truly See What We Point At? Investigating Indexical, Iconic, and Symbolic Gesture Comprehension (2025.acl-short)

Copied to clipboard

Challenge: In recent years, multimodal large language models (MLLMs) excel at integrating textual, auditory, and visual information, but their ability to accurately interpret gestures remains underexplored.
Approach: They annotated five gesture type labels to 925 gesture instances from the Miraikan SC Corpus and analyzed gesture descriptions generated by state-of-the-art MLLMs, including GPT-4o.
Outcome: The proposed models lack real-world referential understanding and are inconsistent in interpreting indexical gestures.
Encoding Gesture in Multimodal Dialogue: Creating a Corpus of Multimodal AMR (2024.lrec-main)

Copied to clipboard

Challenge: Abstract Meaning Representation (AMR) was designed to represent sentence meaning in English text, but recent research has explored its adaptation to broader domains, including documents, dialogues, spatial information, cross-lingual tasks, and gesture.
Approach: They propose to annotate a multimodal (speech and gesture) AMR corpus in a task-based setting and capture coreference relationships across modalities.
Outcome: The proposed corpus captures coreference relationships across modalities, enabling fine-grained analysis of how gesture and natural language interact.
A Formal Analysis of Multimodal Referring Strategies Under Common Ground (2020.lrec-1)

Copied to clipboard

Challenge: a recent study has focused on multimodality in the CL/NLP community, but it has not been widely studied.
Approach: They propose to analyze mixed-modality definite referring expressions using gestures and linguistic descriptions.
Outcome: The proposed models can predict viewer judgment of referring expressions and generate more natural and informative expressions.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations