Challenge: Existing studies have shown that nonverbal behaviours are rich in communicative functions, such as gaze, head movements, and speech-accompanying manual gestures.
Approach: They train a transformer-based neural sequence model to process gaze data extracted from video-recorded conversations and compute its information density.
Outcome: The proposed model computes listeners’ gaze behaviour and the information density of speech using a pre-trained language model.

Similar Papers

How Much Does Nonverbal Communication Conform to Entropy Rate Constancy?: A Case Study on Listener Gaze in Interaction (2024.findings-acl)

Copied to clipboard

Challenge: Whether the Entropy Rate Constancy principle applies to nonverbal communication signals is still under investigation.
Approach: They perform empirical analyses of video-recorded dialogue data and investigate whether listener gaze adheres to the Entropy Rate Constancy principle.
Outcome: The results show that the ERC principle holds for listener gaze, and that linguistic factors syntactic complexity and turn transition potential are weakly correlated with local entropy of listener gaze.
Revisiting Entropy Rate Constancy in Text (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing evidence supports the uniform information density hypothesis . however, we re-evaluate the hypothesis with neural language models .
Approach: They propose to use n-gram language models to argue that English documents exhibit entropy rate constancy . they re-evaluate the claims of Genzel and Charniak with neural language models .
Outcome: The proposed hypothesis fails to support the proposed hypothesis with language models.
Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language Models (2025.acl-long)

Copied to clipboard

Challenge: Large language models have demonstrated remarkable performance across a wide range of language tasks due to their remarkable ability in context modeling.
Approach: They propose to use parallel context encoding to reduce attention entropy by incorporating attention sinks and selective mechanisms to reduce irregular attention . they also propose to incorporate attention sink mechanisms into the parallel encoded context to reduce the irregular attention.
Outcome: The proposed methods lower irregular attention entropy and narrow performance gaps.
Mitigating Linguistic Artifacts in Emotion Recognition for Conversations from TV Scripts to Daily Conversations (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies on Emotion Recognition in Conversations (ERC) focus on training and testing models on the same datasets and there is no prior work on adaptability.
Approach: They propose to use contrastive learning to prioritize emotional features over a linguistic style and refining emotion predictions with pseudo-emotion intensity score to improve model's robustness and accuracy in diverse conversational contexts.
Outcome: The proposed techniques reduce reliance on linguistic artifacts found in TV transcripts and improve model’s robustness and accuracy in diverse conversational contexts.
Entropy- and Distance-Based Predictors From GPT-2 Attention Patterns Predict Reading Times Over and Above GPT-2 Surprisal (2022.emnlp-main)

Copied to clipboard

Challenge: Transformer-based large language models are trained to make predictions about the next word by aggregating representations of previous tokens through their self-attention mechanism.
Approach: They propose an entropy-based predictor that quantifies the diffuseness of self-attention and a distance-based one that captures the incremental change in attention patterns across timesteps.
Outcome: The proposed models perform better over a rigorous baseline including GPT-2 surprisal than previous models that used entropy-based predictors and distance-based ones.
Evaluating Webcam-based Gaze Data as an Alternative for Human Rationale Annotations (2024.lrec-main)

Copied to clipboard

Challenge: We compare webcam-based eye-tracking recordings with human-annotated rationales to evaluate importance scores.
Approach: They compare webcam-based eye-tracking recordings with attention-based importance scores for 4 different multilingual Transformer-based language models.
Outcome: The proposed method is comparable to human rationales in linguistic analysis.
Vision-Language Models Mistake Head Orientation for Gaze Direction: Nonverbal Conversation Cues (2026.findings-acl)

Copied to clipboard

Challenge: Where someone looks is a nonverbal communication cue that children and adults readily use.
Approach: They used 1,360 real-world photos to construct evaluation stimuli for Vision-Language Models (VLMs) they found a substantial performance gap between VLMs and humans .
Outcome: The proposed model outperforms existing models in predicting gaze direction using head orientation rather than eye appearance.
Scaling in Cognitive Modelling: a Multilingual Approach to Human Reading Times (2023.acl-short)

Copied to clipboard

Challenge: Neural language models provide conditional probability distributions over the lexicon that are predictive of human processing times.
Approach: They propose to use a transformer-based model to generate probabilistic estimates that are less predictive of early eye-tracking measurements reflecting lexical access and early semantic integration.
Outcome: The proposed models show that larger models capture late eye-tracking measurements that reflect the full integration of a word into the current language context.
Modeling Referential Gaze in Task-oriented Settings of Varying Referential Complexity (2022.findings-aacl)

Copied to clipboard

Challenge: Referential gaze is a fundamental phenomenon for psycholinguistics and human-human communication.
Approach: They propose a multimodal NLP task to predict when the gaze is referential . they train a sequential attention-based LSTM model and a transformer encoder architecture to model referential gaze and transfer gaze features to unseen situated settings .
Outcome: The proposed model can be applied to situations with different referential complexities . the proposed model is based on an attention-based LSTM model and a multivariate transformer encoder architecture .
From Human Reading to NLM Understanding: Evaluating the Role of Eye-Tracking Data in Encoder-Based Models (2025.acl-long)

Copied to clipboard

Challenge: integrating eye-tracking features into Neural Language Models does not degrade downstream task performance, enhances alignment between model attention and human attention patterns, and compresses the embedding space.
Approach: They used eye-gaze data from the Ghent Eye-Tracking Corpus to investigate how integrating knowledge of human reading behavior impacts Neural Language Models.
Outcome: The proposed approach does not degrade downstream task performance, enhances alignment between model attention and human attention patterns, and compresses the embedding space.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations