Variance of Average Surprisal: A Better Predictor for Quality of Grammar from Unsupervised PCFG Induction (P19-1)
Copied to clipboard
| Challenge: | In unsupervised grammar induction, data likelihood is only weakly correlated with parsing accuracy, especially at convergence after multiple runs. |
| Approach: | They propose to use VAS instead of data likelihood to find better grammars by examining linguistically-motivated constraints related to syntax. |
| Outcome: | The proposed model is better suited for word order typology classification than data likelihood. |
Similar Papers
Temperature-scaling surprisal estimates improve fit to human reading times – but does it do so for the “right reasons”? (2024.acl-long)
Copied to clipboard
| Challenge: | a wide body of evidence shows that human language processing difficulty is predicted by the information-theoretic measure surprisal, a word’s negative log probability in context. |
| Approach: | They propose to use large language models to predict the surprisal of a word's negative log probability in context to test their predictive power. |
| Outcome: | The proposed model can be significantly more accurate than humans because it has more data. |
Language Model Quality Correlates with Psychometric Predictive Power in Multiple Languages (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies have found that higher quality language models provide more powerful predictors of human reading behavior, but empirical support for the QP hypothesis is mixed. |
| Approach: | They propose to test the quality–power hypothesis by using surprisal language models to test their ability to predict eye tracking data. |
| Outcome: | The proposed model is based on a set of language models with a 'quality-power' hypothesis. |
Extracting structure from an LLM - how to improve on surprisal-based models of Human Language Processing (2025.coling-main)
Copied to clipboard
| Challenge: | Existing computational models capture prediction and reanalysis using Large Language Models (LLMs) and a statistical measure known as ‘surprisal’. |
| Approach: | They propose to extract structural information from Large Language Models and a statistical measure known as ‘surprisal’ to integrate it with their learnt statistics. |
| Outcome: | The proposed model achieved higher correlation with human reading times and better predicted the garden path effect and could distinguish between sentence types with different levels of difficulty. |
Frequency Explains the Inverse Correlation of Large Language Models’ Size, Training Data Amount, and Surprisal’s Fit to Reading Times (2024.eacl-long)
Copied to clipboard
| Challenge: | Recent studies have shown that as Transformer-based language models become larger and are trained on very large amounts of data, the fit of their surprisal estimates to naturalistic human reading times degrades. |
| Approach: | They present a series of analyses showing that word frequency is a key explanatory factor underlying these two trends. |
| Outcome: | The results show that word frequency is a key explanatory factor underlying these two trends. |
Character-based PCFG Induction for Modeling the Syntactic Acquisition of Morphologically Rich Languages (2021.findings-emnlp)
Copied to clipboard
| Challenge: | Existing models for syntactic acquisition are word-based and do not inspect functional affixes. |
| Approach: | They propose a computer-based induction model that allows a clean ablation of the influence of subword information in grammar induction. |
| Outcome: | The proposed model is more accurate in morphologically richer languages with subword information than word-based models. |
The Impact of Token Granularity on the Predictive Power of Language Model Surprisal (2025.acl-long)
Copied to clipboard
| Challenge: | Word-by-word language model surprisal is often used to model the incremental processing of human readers, but has been overlooked in cognitive modeling due to the granularity of subword tokens. |
| Approach: | They propose to manipulate token granularity to account for processing difficulty of naturalistic text and garden-path constructions. |
| Outcome: | The proposed model can account for the processing difficulty of naturalistic text and garden-path constructions by using tokens defined by a vocabulary size of 8,000. |
An Existence Proof for Neural Language Models That Can Explain Garden-Path Effects via Surprisal (2026.acl-long)
Copied to clipboard
| Challenge: | Surprisal theory claims that difficulty of sentences increases linearly with surprise . a neural LM that can explain garden-path effects cannot be built, says a new study . |
| Approach: | They propose to fine-tune neural LMs to better align surprisal-based reading-time estimates with actual reading times. |
| Outcome: | a new study shows that fine-tuned neural LMs do not overfit on held-out items . the results show that they improve predictive power for human reading times . |
Language models emulate certain cognitive profiles: An investigation of how predictability measures interact with individual differences (2024.findings-acl)
Copied to clipboard
| Challenge: | incorporating cognitive capacities increases predictive power of surprisal and entropy measures on reading data, whereas high performance in the psychometric tests is associated with lower sensitivity to predictability effects. |
| Approach: | They examine the predictive power (PP) of surprisal and entropy estimated from generative language models (LMs) on reading data from individuals who also completed a wide range of psychometric tests. |
| Outcome: | The LMs' predictive power is based on cognitive capacities and high performance in psychometric tests is associated with lower sensitivity to predictability effects. |
Surprisal from Larger Transformer-based Language Models Predicts fMRI Data More Poorly (2026.eacl-short)
Copied to clipboard
| Challenge: | Recent work has observed an inverse scaling relationship between Transformers’ per-word estimated probability and the predictive power of their surprisal estimates on reading times. |
| Approach: | They conducted a more comprehensive evaluation using surprisal estimates from 17 pre-trained LMs on two functional magnetic resonance imaging datasets. |
| Outcome: | Recent work shows that surprisal from larger Transformer-based models is less predictive of reading times, resolving the inconclusive results and indicating that this trend is not specific to latency-based measures. |
Towards a Similarity-adjusted Surprisal Theory (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies have shown that surprisal theory ignores the possibility of similarity between words and treats them as distinct entities. |
| Approach: | They propose a new measure of comprehension effort called information value that accounts for communicative equivalences between possible continuations. |
| Outcome: | The proposed measure of comprehension effort is based on the diversity index of the diversity of communicative units. |