Challenge: In unsupervised grammar induction, data likelihood is only weakly correlated with parsing accuracy, especially at convergence after multiple runs.
Approach: They propose to use VAS instead of data likelihood to find better grammars by examining linguistically-motivated constraints related to syntax.
Outcome: The proposed model is better suited for word order typology classification than data likelihood.

Similar Papers

Temperature-scaling surprisal estimates improve fit to human reading times – but does it do so for the “right reasons”? (2024.acl-long)

Copied to clipboard

Challenge: a wide body of evidence shows that human language processing difficulty is predicted by the information-theoretic measure surprisal, a word’s negative log probability in context.
Approach: They propose to use large language models to predict the surprisal of a word's negative log probability in context to test their predictive power.
Outcome: The proposed model can be significantly more accurate than humans because it has more data.
Language Model Quality Correlates with Psychometric Predictive Power in Multiple Languages (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies have found that higher quality language models provide more powerful predictors of human reading behavior, but empirical support for the QP hypothesis is mixed.
Approach: They propose to test the quality–power hypothesis by using surprisal language models to test their ability to predict eye tracking data.
Outcome: The proposed model is based on a set of language models with a 'quality-power' hypothesis.
Extracting structure from an LLM - how to improve on surprisal-based models of Human Language Processing (2025.coling-main)

Copied to clipboard

Challenge: Existing computational models capture prediction and reanalysis using Large Language Models (LLMs) and a statistical measure known as ‘surprisal’.
Approach: They propose to extract structural information from Large Language Models and a statistical measure known as ‘surprisal’ to integrate it with their learnt statistics.
Outcome: The proposed model achieved higher correlation with human reading times and better predicted the garden path effect and could distinguish between sentence types with different levels of difficulty.
Frequency Explains the Inverse Correlation of Large Language Models’ Size, Training Data Amount, and Surprisal’s Fit to Reading Times (2024.eacl-long)

Copied to clipboard

Challenge: Recent studies have shown that as Transformer-based language models become larger and are trained on very large amounts of data, the fit of their surprisal estimates to naturalistic human reading times degrades.
Approach: They present a series of analyses showing that word frequency is a key explanatory factor underlying these two trends.
Outcome: The results show that word frequency is a key explanatory factor underlying these two trends.
Character-based PCFG Induction for Modeling the Syntactic Acquisition of Morphologically Rich Languages (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing models for syntactic acquisition are word-based and do not inspect functional affixes.
Approach: They propose a computer-based induction model that allows a clean ablation of the influence of subword information in grammar induction.
Outcome: The proposed model is more accurate in morphologically richer languages with subword information than word-based models.
The Impact of Token Granularity on the Predictive Power of Language Model Surprisal (2025.acl-long)

Copied to clipboard

Challenge: Word-by-word language model surprisal is often used to model the incremental processing of human readers, but has been overlooked in cognitive modeling due to the granularity of subword tokens.
Approach: They propose to manipulate token granularity to account for processing difficulty of naturalistic text and garden-path constructions.
Outcome: The proposed model can account for the processing difficulty of naturalistic text and garden-path constructions by using tokens defined by a vocabulary size of 8,000.
An Existence Proof for Neural Language Models That Can Explain Garden-Path Effects via Surprisal (2026.acl-long)

Copied to clipboard

Challenge: Surprisal theory claims that difficulty of sentences increases linearly with surprise . a neural LM that can explain garden-path effects cannot be built, says a new study .
Approach: They propose to fine-tune neural LMs to better align surprisal-based reading-time estimates with actual reading times.
Outcome: a new study shows that fine-tuned neural LMs do not overfit on held-out items . the results show that they improve predictive power for human reading times .
Language models emulate certain cognitive profiles: An investigation of how predictability measures interact with individual differences (2024.findings-acl)

Copied to clipboard

Challenge: incorporating cognitive capacities increases predictive power of surprisal and entropy measures on reading data, whereas high performance in the psychometric tests is associated with lower sensitivity to predictability effects.
Approach: They examine the predictive power (PP) of surprisal and entropy estimated from generative language models (LMs) on reading data from individuals who also completed a wide range of psychometric tests.
Outcome: The LMs' predictive power is based on cognitive capacities and high performance in psychometric tests is associated with lower sensitivity to predictability effects.
Surprisal from Larger Transformer-based Language Models Predicts fMRI Data More Poorly (2026.eacl-short)

Copied to clipboard

Challenge: Recent work has observed an inverse scaling relationship between Transformers’ per-word estimated probability and the predictive power of their surprisal estimates on reading times.
Approach: They conducted a more comprehensive evaluation using surprisal estimates from 17 pre-trained LMs on two functional magnetic resonance imaging datasets.
Outcome: Recent work shows that surprisal from larger Transformer-based models is less predictive of reading times, resolving the inconclusive results and indicating that this trend is not specific to latency-based measures.
Towards a Similarity-adjusted Surprisal Theory (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies have shown that surprisal theory ignores the possibility of similarity between words and treats them as distinct entities.
Approach: They propose a new measure of comprehension effort called information value that accounts for communicative equivalences between possible continuations.
Outcome: The proposed measure of comprehension effort is based on the diversity index of the diversity of communicative units.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations