Challenge: Word-by-word language model surprisal is often used to model the incremental processing of human readers, but has been overlooked in cognitive modeling due to the granularity of subword tokens.
Approach: They propose to manipulate token granularity to account for processing difficulty of naturalistic text and garden-path constructions.
Outcome: The proposed model can account for the processing difficulty of naturalistic text and garden-path constructions by using tokens defined by a vocabulary size of 8,000.

Similar Papers

On the Proper Treatment of Units in Surprisal Theory (2026.acl-long)

Copied to clipboard

Challenge: empirical work often leaves the notion of a unit underspecified . empirical work has sought to characterize the processing difficulty comprehenders experience .
Approach: They propose a framework for reasoning about surprisal over arbitrary unit inventories . they argue that surprises should be explicit and treat tokenization as implementation detail .
Outcome: The proposed framework disentangles the models' definitions and the regions of interest and treats tokenization as an implementation detail rather than a scientific primitive.
The Inverse Scaling Effect of Pre-Trained Language Model Surprisal Is Not Due to Data Leakage (2025.findings-acl)

Copied to clipboard

Challenge: Language models (LMs) have been shown to flexibly capture many linguistic regularities from raw text, but the source stimuli of reading time datasets are often naturalistic text that are available online.
Approach: They propose to replicate the negative relationship between language model size and the fit of surprisal to reading times using models trained on ‘leakage-free’ data that overlaps only minimally with the reading time corpora.
Outcome: The proposed models show that language models trained on 'leakage-free' data are not driven by data leakage.
On the Proper Treatment of Tokenization in Psycholinguistics (2024.emnlp-main)

Copied to clipboard

Challenge: Language models are used in computational psycholinguistics to test theories that relate the surprisal of a region of interest to its cognitive cost experienced by readers.
Approach: They propose to marginalize token-level language models into character-level ones before they are used in psycholinguistic studies.
Outcome: The proposed model over token strings is better than character-level model, the authors show . the proposed model marginalizes token-level models into character-based models before they are used in psycholinguistic studies.
An Existence Proof for Neural Language Models That Can Explain Garden-Path Effects via Surprisal (2026.acl-long)

Copied to clipboard

Challenge: Surprisal theory claims that difficulty of sentences increases linearly with surprise . a neural LM that can explain garden-path effects cannot be built, says a new study .
Approach: They propose to fine-tune neural LMs to better align surprisal-based reading-time estimates with actual reading times.
Outcome: a new study shows that fine-tuned neural LMs do not overfit on held-out items . the results show that they improve predictive power for human reading times .
Temperature-scaling surprisal estimates improve fit to human reading times – but does it do so for the “right reasons”? (2024.acl-long)

Copied to clipboard

Challenge: a wide body of evidence shows that human language processing difficulty is predicted by the information-theoretic measure surprisal, a word’s negative log probability in context.
Approach: They propose to use large language models to predict the surprisal of a word's negative log probability in context to test their predictive power.
Outcome: The proposed model can be significantly more accurate than humans because it has more data.
Transformer-Based Language Model Surprisal Predicts Human Reading Times Best with About Two Billion Training Tokens (2023.findings-emnlp)

Copied to clipboard

Challenge: Recent studies have drawn conflicting conclusions about the relationship between the quality of a language model and the ability of its surprisal estimates to predict human reading times.
Approach: They propose to evaluate surprisal estimates from Transformer-based language model variants that vary systematically in the amount of training data and model capacity on their ability to predict human reading times.
Outcome: The proposed model variants with contemporary model capacities provide the best fit after seeing about two billion training tokens, while smaller models show a ‘tipping point’ at convergence after the decrease in language model perplexity .
Words, Subwords, and Morphemes: What Really Matters in the Surprisal-Reading Time Relationship? (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing studies using LLMs on psycholinguistic data have gone unverified . a growing body of research is using word-level prediction as a computational proxy .
Approach: They compare morphological, morphologic, and BPE tokenization estimates with reading time data.
Outcome: The proposed method could be used to evaluate morphological prediction.
Why Does Surprisal From Larger Transformer-Based Language Models Provide a Poorer Fit to Human Reading Times? (2023.tacl-1)

Copied to clipboard

Challenge: Existing studies have shown that larger pre-trained language models with more parameters and lower perplexity are less predictive of human reading times.
Approach: They propose to use a transformer-based model with more parameters and lower perplexity to investigate why these models are less predictive of human reading times.
Outcome: The results show that the larger models with more parameters and lower perplexity are less predictive of human reading times and eye-gaze durations collected during naturalistic reading.
The Effects of Surprisal across Languages: Results from Native and Non-native Reading (2022.findings-aacl)

Copied to clipboard

Challenge: Context-dependent predictive processes have been proposed as a core component of the human cognitive system.
Approach: They extract surprisal estimates from mBERT and assess their predictive power on the MECO corpus, a cross-linguistic dataset of eye movement behavior in reading.
Outcome: The proposed model is based on a cross-linguistic dataset of eye movement behavior in reading.
Surprisal and Metaphor Novelty Judgments: Moderate Correlations and Divergent Scaling Effects Revealed by Corpus-Based and Synthetic Datasets (2026.eacl-long)

Copied to clipboard

Challenge: Novel metaphor comprehension involves complex semantic processes and linguistic creativity.
Approach: They propose a cloze-style surprisal method that conditions on full-sentence context.
Outcome: The proposed method shows that LM surprisal yields moderate correlations with scores/labels of metaphor novelty.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations