The Inverse Scaling Effect of Pre-Trained Language Model Surprisal Is Not Due to Data Leakage (2025.findings-acl)
Copied to clipboard
| Challenge: | Language models (LMs) have been shown to flexibly capture many linguistic regularities from raw text, but the source stimuli of reading time datasets are often naturalistic text that are available online. |
| Approach: | They propose to replicate the negative relationship between language model size and the fit of surprisal to reading times using models trained on ‘leakage-free’ data that overlaps only minimally with the reading time corpora. |
| Outcome: | The proposed models show that language models trained on 'leakage-free' data are not driven by data leakage. |
Similar Papers
Why Does Surprisal From Larger Transformer-Based Language Models Provide a Poorer Fit to Human Reading Times? (2023.tacl-1)
Copied to clipboard
| Challenge: | Existing studies have shown that larger pre-trained language models with more parameters and lower perplexity are less predictive of human reading times. |
| Approach: | They propose to use a transformer-based model with more parameters and lower perplexity to investigate why these models are less predictive of human reading times. |
| Outcome: | The results show that the larger models with more parameters and lower perplexity are less predictive of human reading times and eye-gaze durations collected during naturalistic reading. |
Frequency Explains the Inverse Correlation of Large Language Models’ Size, Training Data Amount, and Surprisal’s Fit to Reading Times (2024.eacl-long)
Copied to clipboard
| Challenge: | Recent studies have shown that as Transformer-based language models become larger and are trained on very large amounts of data, the fit of their surprisal estimates to naturalistic human reading times degrades. |
| Approach: | They present a series of analyses showing that word frequency is a key explanatory factor underlying these two trends. |
| Outcome: | The results show that word frequency is a key explanatory factor underlying these two trends. |
Transformer-Based Language Model Surprisal Predicts Human Reading Times Best with About Two Billion Training Tokens (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Recent studies have drawn conflicting conclusions about the relationship between the quality of a language model and the ability of its surprisal estimates to predict human reading times. |
| Approach: | They propose to evaluate surprisal estimates from Transformer-based language model variants that vary systematically in the amount of training data and model capacity on their ability to predict human reading times. |
| Outcome: | The proposed model variants with contemporary model capacities provide the best fit after seeing about two billion training tokens, while smaller models show a ‘tipping point’ at convergence after the decrease in language model perplexity . |
Surprisal from Larger Transformer-based Language Models Predicts fMRI Data More Poorly (2026.eacl-short)
Copied to clipboard
| Challenge: | Recent work has observed an inverse scaling relationship between Transformers’ per-word estimated probability and the predictive power of their surprisal estimates on reading times. |
| Approach: | They conducted a more comprehensive evaluation using surprisal estimates from 17 pre-trained LMs on two functional magnetic resonance imaging datasets. |
| Outcome: | Recent work shows that surprisal from larger Transformer-based models is less predictive of reading times, resolving the inconclusive results and indicating that this trend is not specific to latency-based measures. |
Temperature-scaling surprisal estimates improve fit to human reading times – but does it do so for the “right reasons”? (2024.acl-long)
Copied to clipboard
| Challenge: | a wide body of evidence shows that human language processing difficulty is predicted by the information-theoretic measure surprisal, a word’s negative log probability in context. |
| Approach: | They propose to use large language models to predict the surprisal of a word's negative log probability in context to test their predictive power. |
| Outcome: | The proposed model can be significantly more accurate than humans because it has more data. |
The Impact of Token Granularity on the Predictive Power of Language Model Surprisal (2025.acl-long)
Copied to clipboard
| Challenge: | Word-by-word language model surprisal is often used to model the incremental processing of human readers, but has been overlooked in cognitive modeling due to the granularity of subword tokens. |
| Approach: | They propose to manipulate token granularity to account for processing difficulty of naturalistic text and garden-path constructions. |
| Outcome: | The proposed model can account for the processing difficulty of naturalistic text and garden-path constructions by using tokens defined by a vocabulary size of 8,000. |
The Effects of Surprisal across Languages: Results from Native and Non-native Reading (2022.findings-aacl)
Copied to clipboard
| Challenge: | Context-dependent predictive processes have been proposed as a core component of the human cognitive system. |
| Approach: | They extract surprisal estimates from mBERT and assess their predictive power on the MECO corpus, a cross-linguistic dataset of eye movement behavior in reading. |
| Outcome: | The proposed model is based on a cross-linguistic dataset of eye movement behavior in reading. |
The Linearity of the Effect of Surprisal on Reading Times across Languages (2023.findings-emnlp)
Copied to clipboard
| Challenge: | a large amount of insight into human language processing can be gleaned by studying word-by-word processing difficulty. |
| Approach: | They extend the study by examining eyetracking corpora of seven languages . they find evidence for superlinearity in some languages, but highly sensitive to language models . |
| Outcome: | The study extends existing studies on english to Danish, Dutch, English, German, Japanese, Mandarin, and Russian. |
Emergent Inabilities? Inverse Scaling Over the Course of Pretraining (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Recent research has found that increased number of model parameters and increased size of the training dataset positively influence model performance. |
| Approach: | They investigate whether language models' performance on specific tasks can decrease over the course of training. |
| Outcome: | The proposed model size-based scaling is found on 8 tasks on which Pythia 12B shows decreased performance over the course of training. |
An Existence Proof for Neural Language Models That Can Explain Garden-Path Effects via Surprisal (2026.acl-long)
Copied to clipboard
| Challenge: | Surprisal theory claims that difficulty of sentences increases linearly with surprise . a neural LM that can explain garden-path effects cannot be built, says a new study . |
| Approach: | They propose to fine-tune neural LMs to better align surprisal-based reading-time estimates with actual reading times. |
| Outcome: | a new study shows that fine-tuned neural LMs do not overfit on held-out items . the results show that they improve predictive power for human reading times . |