Does Listener Gaze in Face-to-Face Interaction Follow the Entropy Rate Constancy Principle: An Empirical Study (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies have shown that nonverbal behaviours are rich in communicative functions, such as gaze, head movements, and speech-accompanying manual gestures. |
| Approach: | They train a transformer-based neural sequence model to process gaze data extracted from video-recorded conversations and compute its information density. |
| Outcome: | The proposed model computes listeners’ gaze behaviour and the information density of speech using a pre-trained language model. |
Similar Papers
How Much Does Nonverbal Communication Conform to Entropy Rate Constancy?: A Case Study on Listener Gaze in Interaction (2024.findings-acl)
Copied to clipboard
| Challenge: | Whether the Entropy Rate Constancy principle applies to nonverbal communication signals is still under investigation. |
| Approach: | They perform empirical analyses of video-recorded dialogue data and investigate whether listener gaze adheres to the Entropy Rate Constancy principle. |
| Outcome: | The results show that the ERC principle holds for listener gaze, and that linguistic factors syntactic complexity and turn transition potential are weakly correlated with local entropy of listener gaze. |
Revisiting Entropy Rate Constancy in Text (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing evidence supports the uniform information density hypothesis . however, we re-evaluate the hypothesis with neural language models . |
| Approach: | They propose to use n-gram language models to argue that English documents exhibit entropy rate constancy . they re-evaluate the claims of Genzel and Charniak with neural language models . |
| Outcome: | The proposed hypothesis fails to support the proposed hypothesis with language models. |
Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language Models (2025.acl-long)
Copied to clipboard
Zhisong Zhang, Yan Wang, Xinting Huang, Tianqing Fang, Hongming Zhang, Chenlong Deng, Shuaiyi Li, Dong Yu
| Challenge: | Large language models have demonstrated remarkable performance across a wide range of language tasks due to their remarkable ability in context modeling. |
| Approach: | They propose to use parallel context encoding to reduce attention entropy by incorporating attention sinks and selective mechanisms to reduce irregular attention . they also propose to incorporate attention sink mechanisms into the parallel encoded context to reduce the irregular attention. |
| Outcome: | The proposed methods lower irregular attention entropy and narrow performance gaps. |
Mitigating Linguistic Artifacts in Emotion Recognition for Conversations from TV Scripts to Daily Conversations (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing studies on Emotion Recognition in Conversations (ERC) focus on training and testing models on the same datasets and there is no prior work on adaptability. |
| Approach: | They propose to use contrastive learning to prioritize emotional features over a linguistic style and refining emotion predictions with pseudo-emotion intensity score to improve model's robustness and accuracy in diverse conversational contexts. |
| Outcome: | The proposed techniques reduce reliance on linguistic artifacts found in TV transcripts and improve model’s robustness and accuracy in diverse conversational contexts. |
Entropy- and Distance-Based Predictors From GPT-2 Attention Patterns Predict Reading Times Over and Above GPT-2 Surprisal (2022.emnlp-main)
Copied to clipboard
| Challenge: | Transformer-based large language models are trained to make predictions about the next word by aggregating representations of previous tokens through their self-attention mechanism. |
| Approach: | They propose an entropy-based predictor that quantifies the diffuseness of self-attention and a distance-based one that captures the incremental change in attention patterns across timesteps. |
| Outcome: | The proposed models perform better over a rigorous baseline including GPT-2 surprisal than previous models that used entropy-based predictors and distance-based ones. |
Evaluating Webcam-based Gaze Data as an Alternative for Human Rationale Annotations (2024.lrec-main)
Copied to clipboard
| Challenge: | We compare webcam-based eye-tracking recordings with human-annotated rationales to evaluate importance scores. |
| Approach: | They compare webcam-based eye-tracking recordings with attention-based importance scores for 4 different multilingual Transformer-based language models. |
| Outcome: | The proposed method is comparable to human rationales in linguistic analysis. |
Vision-Language Models Mistake Head Orientation for Gaze Direction: Nonverbal Conversation Cues (2026.findings-acl)
Copied to clipboard
Zory Zhang, Pinyuan Feng, Bingyang Wang, Tianwei Zhao, Suyang Yu, Qingying Gao, Hokin Deng, Ziqiao Ma, Yijiang Li, Dezhi Luo
| Challenge: | Where someone looks is a nonverbal communication cue that children and adults readily use. |
| Approach: | They used 1,360 real-world photos to construct evaluation stimuli for Vision-Language Models (VLMs) they found a substantial performance gap between VLMs and humans . |
| Outcome: | The proposed model outperforms existing models in predicting gaze direction using head orientation rather than eye appearance. |
Scaling in Cognitive Modelling: a Multilingual Approach to Human Reading Times (2023.acl-short)
Copied to clipboard
| Challenge: | Neural language models provide conditional probability distributions over the lexicon that are predictive of human processing times. |
| Approach: | They propose to use a transformer-based model to generate probabilistic estimates that are less predictive of early eye-tracking measurements reflecting lexical access and early semantic integration. |
| Outcome: | The proposed models show that larger models capture late eye-tracking measurements that reflect the full integration of a word into the current language context. |
Modeling Referential Gaze in Task-oriented Settings of Varying Referential Complexity (2022.findings-aacl)
Copied to clipboard
| Challenge: | Referential gaze is a fundamental phenomenon for psycholinguistics and human-human communication. |
| Approach: | They propose a multimodal NLP task to predict when the gaze is referential . they train a sequential attention-based LSTM model and a transformer encoder architecture to model referential gaze and transfer gaze features to unseen situated settings . |
| Outcome: | The proposed model can be applied to situations with different referential complexities . the proposed model is based on an attention-based LSTM model and a multivariate transformer encoder architecture . |
From Human Reading to NLM Understanding: Evaluating the Role of Eye-Tracking Data in Encoder-Based Models (2025.acl-long)
Copied to clipboard
| Challenge: | integrating eye-tracking features into Neural Language Models does not degrade downstream task performance, enhances alignment between model attention and human attention patterns, and compresses the embedding space. |
| Approach: | They used eye-gaze data from the Ghent Eye-Tracking Corpus to investigate how integrating knowledge of human reading behavior impacts Neural Language Models. |
| Outcome: | The proposed approach does not degrade downstream task performance, enhances alignment between model attention and human attention patterns, and compresses the embedding space. |