Gestures Are Used Rationally: Information Theoretic Evidence from Neural Sequential Models (2022.coling-1)
Copied to clipboard
| Challenge: | Verbal communication is companied by rich non-verbal signals, but few studies have explored the non- verbal channels with finer theoretical lens. |
| Approach: | They extract gesture representations from monologue video data and train neural sequential models to examine their results. |
| Outcome: | The proposed method shows that speakers use simple gestures to convey information that enhances verbal communication. |
Similar Papers
Spontaneous gestures encoded by hand positions improve language models: An Information-Theoretic motivated study (2023.findings-acl)
Copied to clipboard
| Challenge: | a key missing step is to explore whether the nonverbal information can be quantified. |
| Approach: | They explore whether incorporating gesture representations can improve the language model’s performance . they also examine whether spontaneous gestures demonstrate entropy rate constancy (ERC) . |
| Outcome: | The proposed model improves the performance of the mixed-modal language models against monologue video data. |
Does Listener Gaze in Face-to-Face Interaction Follow the Entropy Rate Constancy Principle: An Empirical Study (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies have shown that nonverbal behaviours are rich in communicative functions, such as gaze, head movements, and speech-accompanying manual gestures. |
| Approach: | They train a transformer-based neural sequence model to process gaze data extracted from video-recorded conversations and compute its information density. |
| Outcome: | The proposed model computes listeners’ gaze behaviour and the information density of speech using a pre-trained language model. |
How Much Does Nonverbal Communication Conform to Entropy Rate Constancy?: A Case Study on Listener Gaze in Interaction (2024.findings-acl)
Copied to clipboard
| Challenge: | Whether the Entropy Rate Constancy principle applies to nonverbal communication signals is still under investigation. |
| Approach: | They perform empirical analyses of video-recorded dialogue data and investigate whether listener gaze adheres to the Entropy Rate Constancy principle. |
| Outcome: | The results show that the ERC principle holds for listener gaze, and that linguistic factors syntactic complexity and turn transition potential are weakly correlated with local entropy of listener gaze. |
What Do Prosody and Text Convey? Characterizing How Meaningful Information is Distributed Across Multiple Channels (2026.acl-long)
Copied to clipboard
| Challenge: | Prosody—the melody of speech—conveys critical information often not captured by the words or text of a message. |
| Approach: | They propose an information-theoretic approach to quantify how much is conveyed by prosody that is not recoverable from text alone. |
| Outcome: | The proposed framework can quantify how much is conveyed by prosody that is not recoverable from text alone and crucially, what prosody conveys. |
Deal, or no deal (or who knows)? Forecasting Uncertainty in Conversations using Large Language Models (2024.findings-acl)
Copied to clipboard
| Challenge: | Effective interlocutors account for the uncertain goals, beliefs, and emotions of others. |
| Approach: | They propose to calibrate language models to better represent outcome uncertainty . they propose to use two methods to calibrated small open-source models . |
| Outcome: | The proposed fine-tuning strategies can calibrate smaller open-source models to beat pre-trained models 10x their size. |
Revisiting Entropy Rate Constancy in Text (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing evidence supports the uniform information density hypothesis . however, we re-evaluate the hypothesis with neural language models . |
| Approach: | They propose to use n-gram language models to argue that English documents exhibit entropy rate constancy . they re-evaluate the claims of Genzel and Charniak with neural language models . |
| Outcome: | The proposed hypothesis fails to support the proposed hypothesis with language models. |
The Effect of Efficient Messaging and Input Variability on Neural-Agent Iterated Language Learning (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies have focused on agent-based simulations of language emergence. |
| Approach: | They propose to model the trade-off between word order and inflection in natural languages by using neural network agents. |
| Outcome: | The results show that neural agents strive to maintain the utterance type distribution observed during learning, rather than developing a more efficient or systematic language. |
Towards Understanding the Relation between Gestures and Language (2022.coling-1)
Copied to clipboard
| Challenge: | a new study explores the relationship between gestures and language . we use contrastive learning to learn gesture embeddings . |
| Approach: | They adapt a semi-supervised multimodal model to learn gesture embeddings using Ted talks . they show gestures are predictive of the native language of the speaker . |
| Outcome: | The proposed model learns gesture embeddings from a multimodal dataset . it shows that gesture embeds are predictive of the native language of the speaker . |
LLM Knows Body Language, Too: Translating Speech Voices into Human Gestures (2024.acl-long)
Copied to clipboard
| Challenge: | despite advances in the generation of realistic human gestures, the process often includes unintended, meaningless, or non-realistic gestures. |
| Approach: | They propose a framework that leverages large language models to generate human gestures . the primary stage employs a transformer-based auto-encoder network to encode human gesture into discrete symbols . |
| Outcome: | The proposed framework has demonstrated state-of-the-art performance on public TED and TED-Expressive datasets. |
Unveiling the Limits of Large Language Models in Inferring Pragmatic Meaning from Non-Verbal Responses (2026.acl-long)
Copied to clipboard
| Challenge: | Existing studies have focused mainly on LLMs' comprehension of verbal behavior, with non-verbal behavior considered only in conjunction with verbal responses. |
| Approach: | They present the first systematic evaluation of LLMs’ ability to infer pragmatic meaning in dialogue consisting solely of non-verbal responses. |
| Outcome: | The proposed model fails to capture non-verbal intent and has accuracy dropping by 60% compared to verbal ones. |