Papers by Ethan Wilcox

21 papers
Language Model Quality Correlates with Psychometric Predictive Power in Multiple Languages (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies have found that higher quality language models provide more powerful predictors of human reading behavior, but empirical support for the QP hypothesis is mixed.
Approach: They propose to test the quality–power hypothesis by using surprisal language models to test their ability to predict eye tracking data.
Outcome: The proposed model is based on a set of language models with a 'quality-power' hypothesis.
Structural Supervision Improves Few-Shot Learning and Syntactic Generalization in Neural Language Models (2020.emnlp-main)

Copied to clipboard

Challenge: Existing studies have not investigated the relationship between a token's frequency in the training corpus and syntactic properties models learn about it.
Approach: They develop controlled experiments that probe models’ syntactic nominal number and verbal argument structure generalizations for tokens seen as few as two times during training.
Outcome: The proposed models can make syntactic generalizations for tokens seen as few as two times during training and transfer them to transformed contexts.
Unpacking Let Alone: Human-Scale Models Generalize to a Rare Construction in Form but not Meaning (2025.emnlp-main)

Copied to clipboard

Challenge: Recent evidence suggests that language models with human-scale pretraining data may possess a similar generalization ability by generalizing from frequent to rare constructions.
Approach: They construct a synthetic benchmark that targets syntactic and semantic properties of the English Let-Alone construction and compare it with a human-scale transformer language model.
Outcome: The proposed model can generalize from frequent to rare constructions, but human-scale models do not make correct generalizations about Let-Alone’s meaning.
Revisiting the Optimality of Word Lengths (2023.emnlp-main)

Copied to clipboard

Challenge: Existing theories of communication cost are based on the length of an utterance, but they are not based upon the frequency of the utterant.
Approach: They propose a method to minimize CCH's cost by comparing language lengths to surprisal's expectation and variance-to-mean ratios.
Outcome: The proposed method does not minimize CCH’s cost, but rather a lower bound, which we term CCH-lower.
SyntaxGym: An Online Platform for Targeted Evaluation of Language Models (2020.acl-demos)

Copied to clipboard

Challenge: SyntaxGym is an online platform and open-source framework for targeted syntactic evaluation of neural network language models.
Approach: They propose to make targeted syntactic evaluations accessible to both experts in NLP and linguistics and reproducible across computing environments.
Outcome: The proposed framework is reproducible across computing environments and standardized following the norms of psycholinguistic experimental design.
Surprise! Uniform Information Density Isn’t the Whole Story: Predicting Surprisal Contours in Long-form Discourse (2024.emnlp-main)

Copied to clipboard

Challenge: Uniform Information Density (UID) hypothesis posits that speakers tend to distribute information evenly across linguistic units to achieve efficient communication.
Approach: They propose a functional pressure that speakers modulate information rate based on location within a hierarchically-structured model of discourse.
Outcome: The proposed hypothesis posits that speakers tend to distribute information evenly across linguistic units to achieve efficient communication.
The time scale of redundancy between prosody and linguistic context (2025.acl-long)

Copied to clipboard

Challenge: Prior work has shown that the information carried by prosodic features is substantially redundant with that carried by the surrounding words.
Approach: They examine the time scale of this relationship, studying how it varies with the length of past and future contexts.
Outcome: The results show that prosody features show some redundancy with future words, but only with a short scale of 1-2 words, consistent with reports of incremental short-term planning in language production.
Reverse-Engineering the Reader (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies have sought to determine to what extent language models can serve as useful models of human cognition by aligning them to human psychometric data.
Approach: They propose a method to fine-tune a language model to implicitly optimize parameters of a linear regressor that directly predicts humans’ reading times of in-context linguistic units.
Outcome: The proposed technique improves language models’ psychometric predictive power but also its perplexity on held-out test data.
On the Role of Context in Reading Time Prediction (2024.emnlp-main)

Copied to clipboard

Challenge: a new perspective on how readers integrate context during reading time prediction is presented . a recent study shows that the proportion of variance in reading times explained by context is smaller when context is represented by the orthogonalized predictor.
Approach: They propose a technique where they project surprisal onto the orthogonal complement of frequency.
Outcome: The proposed method shows that the proportion of variance in reading times explained by context is smaller when context is represented by the orthogonalized predictor.
Structural Supervision Improves Learning of Non-Local Grammatical Dependencies (N19-1)

Copied to clipboard

Challenge: State-of-the-art LSTM language models learn sequential contingencies with some success . LS models fail to learn other non-local grammatical dependencies, however .
Approach: They compare LSTM language models with RNNGs to examine grammatical dependencies . they find that hierarchical supervision improves learning of non-local dependencies.
Outcome: The proposed model outperforms the existing model on non-local dependencies and learns many of the Island Constraints on the filler-gap dependency.
The Harmonic Structure of Information Contours (2025.acl-long)

Copied to clipboard

Challenge: Language typically does not maintain a uniform information rate, but it fluctuates around a global average . a new study suggests periodicity may be a factor in information rate oscillations .
Approach: They propose a hypothesis that language does not maintain a uniform information rate . they apply harmonic regression and introduce a new extension to detect periodicity .
Outcome: The proposed method reveals that language oscillates at periodic intervals across frequencies . it also offers a framework for uncovering structural pressures at various levels of linguistic granularity.
Using Information Theory to Characterize Prosodic Typology: The Case of Tone, Pitch-Accent and Stress-Accent (2025.acl-long)

Copied to clipboard

Challenge: lexical identity and prosody are well-studied parameters of linguistic variation, but they are difficult to predict in tonal languages.
Approach: They propose to characterize the relationship between lexical identity and prosody using information theory to estimate mutual information between the text and pitch curves.
Outcome: The proposed hypothesis supports perspectives that view linguistic typology as gradient, rather than categorical.
A Targeted Assessment of Incremental Processing in Neural Language Models and Humans (2021.acl-long)

Copied to clipboard

Challenge: Using by-word reaction time data, we compare incremental processing in humans and neural language models across a range of structural phenomena.
Approach: They propose to scale up incremental processing in humans and language models by collecting by-word reaction time data for 16 different syntactic test suites.
Outcome: The proposed model outputs match human and model accuracy scores, but underpredict the difference in magnitude of incremental processing difficulty between grammatical and ungrammatically-spaced sentences.
Anything Goes? A Crosslinguistic Study of (Im)possible Language Learning in LMs (2025.acl-long)

Copied to clipboard

Challenge: LMs are highly flexible learners, capable of acquiring linguistic patterns beyond those learnable by humans.
Approach: They train LMs to model impossible and typologically unattested languages . they find that the model does not achieve perfect separation between attested and unattest languages - suggesting some human-like inductive biases .
Outcome: The proposed model can largely distinguish attested from impossible languages, but does not achieve perfect separation between them and their impossible counterparts.
Quantifying the redundancy between prosody and text (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies suggest partial redundancy between prosody and linguistic information.
Approach: They use large language models to estimate how much information is redundant between prosody and the words themselves.
Outcome: The proposed model can predict prosodic features across prosodic features, including intensity, duration, pauses, and pitch contours.
Neural language models as psycholinguistic subjects: Representations of syntactic state (N19-1)

Copied to clipboard

Challenge: a recent study examines the extent to which neural network language models reflect incremental representations of syntactic state . we examine neural network model behavior on sentences chosen to probe specific aspects of the learned representations .
Approach: They employ experimental methodologies developed in psycholinguistics to study syntactic representation in the human mind.
Outcome: The proposed models are trained on large datasets and only sensitive to subtle cues . the results raise questions about the accuracy of the models and their performance .
Modeling Bottom-up Information Quality during Language Processing (2025.emnlp-main)

Copied to clipboard

Challenge: Contemporary theories of language processing model language processing as integrating both top-down expectations and bottom-up inputs.
Approach: They propose an information-theoretic operationalization for the “quality” of bottom-up information as the mutual information between visual information and word identity.
Outcome: The proposed model compares reading times in English and Chinese in which words' information quality has been reduced by occluding their top or bottom half with full words.
A Systematic Assessment of Syntactic Generalization in Neural Language Models (2020.acl-main)

Copied to clipboard

Challenge: Existing work on syntactic knowledge models has not provided a clear picture of the properties required to produce proper syntaktic generalizations.
Approach: They propose to evaluate syntactic knowledge of language models by varying model architectures . they find substantial differences in syntaktic generalization performance by model architecture .
Outcome: The proposed model architectures outperform other architectures on a set of 34 English-language syntactic test suites.
Language Models Grow Less Humanlike beyond Phase Transition (2025.acl-long)

Copied to clipboard

Challenge: Existing studies have shown that LMs' alignment with human reading behavior improves during pretraining up to a tipping point, beyond which it plateaus or degrades.
Approach: They hypothesize that a pretraining phase transition is responsible for the tipping point in PPP and that phase transitions alter the subsequent learning dynamics of the model, such that further training keeps damaging PPP.
Outcome: The proposed model is able to produce attention patterns that contribute to the degradation of PPP, but it is not capable of producing attention patterns.
Representation of Constituents in Neural Language Models: Coordination Phrase as a Case Study (D19-1)

Copied to clipboard

Challenge: Existing studies have focused on the ability of neural models to compute and employ phrase-level features attached to a set of words, such as subject number or whquestion words.
Approach: They examine whether models can represent constituent-level features, using coordinated noun phrases as a case study.
Outcome: The proposed model can combine gender and gender features to drive downstream expectations, while having less success with gender agreement.
On the Efficacy of Sampling Adapters (2023.acl-long)

Copied to clipboard

Challenge: Using sampling adapters can improve the quality of the generated text.
Approach: They propose a framework for understanding sampling adapters and propose 'sampling adapters' they argue that the shift enforced by them can be viewed as a trade-off between precision and recall .
Outcome: The proposed framework can be used to improve the quality of language models by modifying their distributions to improve their precision and recall.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations