Using surprisal and fMRI to map the neural bases of broad and local contextual prediction during natural language comprehension (2021.findings-acl)
Copied to clipboard
| Challenge: | a prior work using surprisal only considered within-sentence context, using n-grams, neural language models, or syntactic structure as conditioning context. |
| Approach: | They extend the surprisal approach to use broader topical context . they identify distinct patterns of neural activation for lexical surprised and topical surpresed . |
| Outcome: | The proposed method captures effects of local and topical contexts on processing . it shows that local and broad contextual cues recruit different brain regions . |
Similar Papers
On the Role of Context in Reading Time Prediction (2024.emnlp-main)
Copied to clipboard
| Challenge: | a new perspective on how readers integrate context during reading time prediction is presented . a recent study shows that the proportion of variance in reading times explained by context is smaller when context is represented by the orthogonalized predictor. |
| Approach: | They propose a technique where they project surprisal onto the orthogonal complement of frequency. |
| Outcome: | The proposed method shows that the proportion of variance in reading times explained by context is smaller when context is represented by the orthogonalized predictor. |
Coreference-aware Surprisal Predicts Brain Response (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing studies have shown that coreference resolution is a key component of language processing and has been used to manipulate variables of interest. |
| Approach: | They propose to enable the parser to process subword information that might better approximate human morphological knowledge and extend evaluation of coreference effects from self-paced reading to human brain imaging data. |
| Outcome: | The proposed model enables the parser to process subword information that might better approximate human morphological knowledge and extends evaluation of coreference effects from self-paced reading to human brain imaging data. |
An Existence Proof for Neural Language Models That Can Explain Garden-Path Effects via Surprisal (2026.acl-long)
Copied to clipboard
| Challenge: | Surprisal theory claims that difficulty of sentences increases linearly with surprise . a neural LM that can explain garden-path effects cannot be built, says a new study . |
| Approach: | They propose to fine-tune neural LMs to better align surprisal-based reading-time estimates with actual reading times. |
| Outcome: | a new study shows that fine-tuned neural LMs do not overfit on held-out items . the results show that they improve predictive power for human reading times . |
The Effects of Surprisal across Languages: Results from Native and Non-native Reading (2022.findings-aacl)
Copied to clipboard
| Challenge: | Context-dependent predictive processes have been proposed as a core component of the human cognitive system. |
| Approach: | They extract surprisal estimates from mBERT and assess their predictive power on the MECO corpus, a cross-linguistic dataset of eye movement behavior in reading. |
| Outcome: | The proposed model is based on a cross-linguistic dataset of eye movement behavior in reading. |
Extracting structure from an LLM - how to improve on surprisal-based models of Human Language Processing (2025.coling-main)
Copied to clipboard
| Challenge: | Existing computational models capture prediction and reanalysis using Large Language Models (LLMs) and a statistical measure known as ‘surprisal’. |
| Approach: | They propose to extract structural information from Large Language Models and a statistical measure known as ‘surprisal’ to integrate it with their learnt statistics. |
| Outcome: | The proposed model achieved higher correlation with human reading times and better predicted the garden path effect and could distinguish between sentence types with different levels of difficulty. |
Topicalization in Language Models: A Case Study on Japanese (2022.coling-1)
Copied to clipboard
| Challenge: | a recent study has shown that neural language models can capture discourse-level preferences in text generation . a particular aspect of discourse is the topic-comment structure . |
| Approach: | They analyze whether neural language models can capture discourse-level preferences in text generation . they use Japanese language and crowdsourced human topicalization judgment data . |
| Outcome: | The proposed model can capture human-like generalizations in discourse-level linguistic aspects. |
Surprisal from Larger Transformer-based Language Models Predicts fMRI Data More Poorly (2026.eacl-short)
Copied to clipboard
| Challenge: | Recent work has observed an inverse scaling relationship between Transformers’ per-word estimated probability and the predictive power of their surprisal estimates on reading times. |
| Approach: | They conducted a more comprehensive evaluation using surprisal estimates from 17 pre-trained LMs on two functional magnetic resonance imaging datasets. |
| Outcome: | Recent work shows that surprisal from larger Transformer-based models is less predictive of reading times, resolving the inconclusive results and indicating that this trend is not specific to latency-based measures. |
The Linearity of the Effect of Surprisal on Reading Times across Languages (2023.findings-emnlp)
Copied to clipboard
| Challenge: | a large amount of insight into human language processing can be gleaned by studying word-by-word processing difficulty. |
| Approach: | They extend the study by examining eyetracking corpora of seven languages . they find evidence for superlinearity in some languages, but highly sensitive to language models . |
| Outcome: | The study extends existing studies on english to Danish, Dutch, English, German, Japanese, Mandarin, and Russian. |
The Impact of Token Granularity on the Predictive Power of Language Model Surprisal (2025.acl-long)
Copied to clipboard
| Challenge: | Word-by-word language model surprisal is often used to model the incremental processing of human readers, but has been overlooked in cognitive modeling due to the granularity of subword tokens. |
| Approach: | They propose to manipulate token granularity to account for processing difficulty of naturalistic text and garden-path constructions. |
| Outcome: | The proposed model can account for the processing difficulty of naturalistic text and garden-path constructions by using tokens defined by a vocabulary size of 8,000. |
On the Proper Treatment of Units in Surprisal Theory (2026.acl-long)
Copied to clipboard
| Challenge: | empirical work often leaves the notion of a unit underspecified . empirical work has sought to characterize the processing difficulty comprehenders experience . |
| Approach: | They propose a framework for reasoning about surprisal over arbitrary unit inventories . they argue that surprises should be explicit and treat tokenization as implementation detail . |
| Outcome: | The proposed framework disentangles the models' definitions and the regions of interest and treats tokenization as an implementation detail rather than a scientific primitive. |