Challenge: Context-dependent predictive processes have been proposed as a core component of the human cognitive system.
Approach: They extract surprisal estimates from mBERT and assess their predictive power on the MECO corpus, a cross-linguistic dataset of eye movement behavior in reading.
Outcome: The proposed model is based on a cross-linguistic dataset of eye movement behavior in reading.

Similar Papers

The Linearity of the Effect of Surprisal on Reading Times across Languages (2023.findings-emnlp)

Copied to clipboard

Challenge: a large amount of insight into human language processing can be gleaned by studying word-by-word processing difficulty.
Approach: They extend the study by examining eyetracking corpora of seven languages . they find evidence for superlinearity in some languages, but highly sensitive to language models .
Outcome: The study extends existing studies on english to Danish, Dutch, English, German, Japanese, Mandarin, and Russian.
On the Role of Context in Reading Time Prediction (2024.emnlp-main)

Copied to clipboard

Challenge: a new perspective on how readers integrate context during reading time prediction is presented . a recent study shows that the proportion of variance in reading times explained by context is smaller when context is represented by the orthogonalized predictor.
Approach: They propose a technique where they project surprisal onto the orthogonal complement of frequency.
Outcome: The proposed method shows that the proportion of variance in reading times explained by context is smaller when context is represented by the orthogonalized predictor.
Why Does Surprisal From Larger Transformer-Based Language Models Provide a Poorer Fit to Human Reading Times? (2023.tacl-1)

Copied to clipboard

Challenge: Existing studies have shown that larger pre-trained language models with more parameters and lower perplexity are less predictive of human reading times.
Approach: They propose to use a transformer-based model with more parameters and lower perplexity to investigate why these models are less predictive of human reading times.
Outcome: The results show that the larger models with more parameters and lower perplexity are less predictive of human reading times and eye-gaze durations collected during naturalistic reading.
Surprisal from Larger Transformer-based Language Models Predicts fMRI Data More Poorly (2026.eacl-short)

Copied to clipboard

Challenge: Recent work has observed an inverse scaling relationship between Transformers’ per-word estimated probability and the predictive power of their surprisal estimates on reading times.
Approach: They conducted a more comprehensive evaluation using surprisal estimates from 17 pre-trained LMs on two functional magnetic resonance imaging datasets.
Outcome: Recent work shows that surprisal from larger Transformer-based models is less predictive of reading times, resolving the inconclusive results and indicating that this trend is not specific to latency-based measures.
Transformer-Based Language Model Surprisal Predicts Human Reading Times Best with About Two Billion Training Tokens (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent studies have drawn conflicting conclusions about the relationship between the quality of a language model and the ability of its surprisal estimates to predict human reading times.
Approach: They propose to evaluate surprisal estimates from Transformer-based language model variants that vary systematically in the amount of training data and model capacity on their ability to predict human reading times.
Outcome: The proposed model variants with contemporary model capacities provide the best fit after seeing about two billion training tokens, while smaller models show a ‘tipping point’ at convergence after the decrease in language model perplexity .
Using surprisal and fMRI to map the neural bases of broad and local contextual prediction during natural language comprehension (2021.findings-acl)

Copied to clipboard

Challenge: a prior work using surprisal only considered within-sentence context, using n-grams, neural language models, or syntactic structure as conditioning context.
Approach: They extend the surprisal approach to use broader topical context . they identify distinct patterns of neural activation for lexical surprised and topical surpresed .
Outcome: The proposed method captures effects of local and topical contexts on processing . it shows that local and broad contextual cues recruit different brain regions .
Surprisal Predicts Code-Switching in Chinese-English Bilingual Text (2020.emnlp-main)

Copied to clipboard

Challenge: a new study examines the propensity of bilinguals to switch languages . word surprisal and word entropy are important predictors of code-switching .
Approach: They propose high cognitive effort as a reason for code-switching . they use a computational model of surprisal and word entropy to model code-changing .
Outcome: The proposed model shows that word surprisal, but not entropy, is a significant predictor . sentence length is also a predictor, which has been related to sentence complexity .
Temperature-scaling surprisal estimates improve fit to human reading times – but does it do so for the “right reasons”? (2024.acl-long)

Copied to clipboard

Challenge: a wide body of evidence shows that human language processing difficulty is predicted by the information-theoretic measure surprisal, a word’s negative log probability in context.
Approach: They propose to use large language models to predict the surprisal of a word's negative log probability in context to test their predictive power.
Outcome: The proposed model can be significantly more accurate than humans because it has more data.
The Inverse Scaling Effect of Pre-Trained Language Model Surprisal Is Not Due to Data Leakage (2025.findings-acl)

Copied to clipboard

Challenge: Language models (LMs) have been shown to flexibly capture many linguistic regularities from raw text, but the source stimuli of reading time datasets are often naturalistic text that are available online.
Approach: They propose to replicate the negative relationship between language model size and the fit of surprisal to reading times using models trained on ‘leakage-free’ data that overlaps only minimally with the reading time corpora.
Outcome: The proposed models show that language models trained on 'leakage-free' data are not driven by data leakage.
Language models emulate certain cognitive profiles: An investigation of how predictability measures interact with individual differences (2024.findings-acl)

Copied to clipboard

Challenge: incorporating cognitive capacities increases predictive power of surprisal and entropy measures on reading data, whereas high performance in the psychometric tests is associated with lower sensitivity to predictability effects.
Approach: They examine the predictive power (PP) of surprisal and entropy estimated from generative language models (LMs) on reading data from individuals who also completed a wide range of psychometric tests.
Outcome: The LMs' predictive power is based on cognitive capacities and high performance in psychometric tests is associated with lower sensitivity to predictability effects.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations