Papers by Ethan Wilcox
Language Model Quality Correlates with Psychometric Predictive Power in Multiple Languages (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies have found that higher quality language models provide more powerful predictors of human reading behavior, but empirical support for the QP hypothesis is mixed. |
| Approach: | They propose to test the quality–power hypothesis by using surprisal language models to test their ability to predict eye tracking data. |
| Outcome: | The proposed model is based on a set of language models with a 'quality-power' hypothesis. |
Structural Supervision Improves Few-Shot Learning and Syntactic Generalization in Neural Language Models (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies have not investigated the relationship between a token's frequency in the training corpus and syntactic properties models learn about it. |
| Approach: | They develop controlled experiments that probe models’ syntactic nominal number and verbal argument structure generalizations for tokens seen as few as two times during training. |
| Outcome: | The proposed models can make syntactic generalizations for tokens seen as few as two times during training and transfer them to transformed contexts. |
Unpacking Let Alone: Human-Scale Models Generalize to a Rare Construction in Form but not Meaning (2025.emnlp-main)
Copied to clipboard
| Challenge: | Recent evidence suggests that language models with human-scale pretraining data may possess a similar generalization ability by generalizing from frequent to rare constructions. |
| Approach: | They construct a synthetic benchmark that targets syntactic and semantic properties of the English Let-Alone construction and compare it with a human-scale transformer language model. |
| Outcome: | The proposed model can generalize from frequent to rare constructions, but human-scale models do not make correct generalizations about Let-Alone’s meaning. |
Revisiting the Optimality of Word Lengths (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing theories of communication cost are based on the length of an utterance, but they are not based upon the frequency of the utterant. |
| Approach: | They propose a method to minimize CCH's cost by comparing language lengths to surprisal's expectation and variance-to-mean ratios. |
| Outcome: | The proposed method does not minimize CCH’s cost, but rather a lower bound, which we term CCH-lower. |
SyntaxGym: An Online Platform for Targeted Evaluation of Language Models (2020.acl-demos)
Copied to clipboard
| Challenge: | SyntaxGym is an online platform and open-source framework for targeted syntactic evaluation of neural network language models. |
| Approach: | They propose to make targeted syntactic evaluations accessible to both experts in NLP and linguistics and reproducible across computing environments. |
| Outcome: | The proposed framework is reproducible across computing environments and standardized following the norms of psycholinguistic experimental design. |
Surprise! Uniform Information Density Isn’t the Whole Story: Predicting Surprisal Contours in Long-form Discourse (2024.emnlp-main)
Copied to clipboard
| Challenge: | Uniform Information Density (UID) hypothesis posits that speakers tend to distribute information evenly across linguistic units to achieve efficient communication. |
| Approach: | They propose a functional pressure that speakers modulate information rate based on location within a hierarchically-structured model of discourse. |
| Outcome: | The proposed hypothesis posits that speakers tend to distribute information evenly across linguistic units to achieve efficient communication. |
The time scale of redundancy between prosody and linguistic context (2025.acl-long)
Copied to clipboard
Tamar I Regev, Chiebuka Ohams, Shaylee Xie, Lukas Wolf, Evelina Fedorenko, Alex Warstadt, Ethan Wilcox, Tiago Pimentel
| Challenge: | Prior work has shown that the information carried by prosodic features is substantially redundant with that carried by the surrounding words. |
| Approach: | They examine the time scale of this relationship, studying how it varies with the length of past and future contexts. |
| Outcome: | The results show that prosody features show some redundancy with future words, but only with a short scale of 1-2 words, consistent with reports of incremental short-term planning in language production. |
Reverse-Engineering the Reader (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing studies have sought to determine to what extent language models can serve as useful models of human cognition by aligning them to human psychometric data. |
| Approach: | They propose a method to fine-tune a language model to implicitly optimize parameters of a linear regressor that directly predicts humans’ reading times of in-context linguistic units. |
| Outcome: | The proposed technique improves language models’ psychometric predictive power but also its perplexity on held-out test data. |
On the Role of Context in Reading Time Prediction (2024.emnlp-main)
Copied to clipboard
| Challenge: | a new perspective on how readers integrate context during reading time prediction is presented . a recent study shows that the proportion of variance in reading times explained by context is smaller when context is represented by the orthogonalized predictor. |
| Approach: | They propose a technique where they project surprisal onto the orthogonal complement of frequency. |
| Outcome: | The proposed method shows that the proportion of variance in reading times explained by context is smaller when context is represented by the orthogonalized predictor. |
Structural Supervision Improves Learning of Non-Local Grammatical Dependencies (N19-1)
Copied to clipboard
| Challenge: | State-of-the-art LSTM language models learn sequential contingencies with some success . LS models fail to learn other non-local grammatical dependencies, however . |
| Approach: | They compare LSTM language models with RNNGs to examine grammatical dependencies . they find that hierarchical supervision improves learning of non-local dependencies. |
| Outcome: | The proposed model outperforms the existing model on non-local dependencies and learns many of the Island Constraints on the filler-gap dependency. |
The Harmonic Structure of Information Contours (2025.acl-long)
Copied to clipboard
Eleftheria Tsipidi, Samuel Kiegeland, Franz Nowak, Tianyang Xu, Ethan Wilcox, Alex Warstadt, Ryan Cotterell, Mario Giulianelli
| Challenge: | Language typically does not maintain a uniform information rate, but it fluctuates around a global average . a new study suggests periodicity may be a factor in information rate oscillations . |
| Approach: | They propose a hypothesis that language does not maintain a uniform information rate . they apply harmonic regression and introduce a new extension to detect periodicity . |
| Outcome: | The proposed method reveals that language oscillates at periodic intervals across frequencies . it also offers a framework for uncovering structural pressures at various levels of linguistic granularity. |
Using Information Theory to Characterize Prosodic Typology: The Case of Tone, Pitch-Accent and Stress-Accent (2025.acl-long)
Copied to clipboard
| Challenge: | lexical identity and prosody are well-studied parameters of linguistic variation, but they are difficult to predict in tonal languages. |
| Approach: | They propose to characterize the relationship between lexical identity and prosody using information theory to estimate mutual information between the text and pitch curves. |
| Outcome: | The proposed hypothesis supports perspectives that view linguistic typology as gradient, rather than categorical. |
A Targeted Assessment of Incremental Processing in Neural Language Models and Humans (2021.acl-long)
Copied to clipboard
| Challenge: | Using by-word reaction time data, we compare incremental processing in humans and neural language models across a range of structural phenomena. |
| Approach: | They propose to scale up incremental processing in humans and language models by collecting by-word reaction time data for 16 different syntactic test suites. |
| Outcome: | The proposed model outputs match human and model accuracy scores, but underpredict the difference in magnitude of incremental processing difficulty between grammatical and ungrammatically-spaced sentences. |
Anything Goes? A Crosslinguistic Study of (Im)possible Language Learning in LMs (2025.acl-long)
Copied to clipboard
| Challenge: | LMs are highly flexible learners, capable of acquiring linguistic patterns beyond those learnable by humans. |
| Approach: | They train LMs to model impossible and typologically unattested languages . they find that the model does not achieve perfect separation between attested and unattest languages - suggesting some human-like inductive biases . |
| Outcome: | The proposed model can largely distinguish attested from impossible languages, but does not achieve perfect separation between them and their impossible counterparts. |
Quantifying the redundancy between prosody and text (2023.emnlp-main)
Copied to clipboard
Lukas Wolf, Tiago Pimentel, Evelina Fedorenko, Ryan Cotterell, Alex Warstadt, Ethan Wilcox, Tamar Regev
| Challenge: | Existing studies suggest partial redundancy between prosody and linguistic information. |
| Approach: | They use large language models to estimate how much information is redundant between prosody and the words themselves. |
| Outcome: | The proposed model can predict prosodic features across prosodic features, including intensity, duration, pauses, and pitch contours. |
Neural language models as psycholinguistic subjects: Representations of syntactic state (N19-1)
Copied to clipboard
| Challenge: | a recent study examines the extent to which neural network language models reflect incremental representations of syntactic state . we examine neural network model behavior on sentences chosen to probe specific aspects of the learned representations . |
| Approach: | They employ experimental methodologies developed in psycholinguistics to study syntactic representation in the human mind. |
| Outcome: | The proposed models are trained on large datasets and only sensitive to subtle cues . the results raise questions about the accuracy of the models and their performance . |
Modeling Bottom-up Information Quality during Language Processing (2025.emnlp-main)
Copied to clipboard
| Challenge: | Contemporary theories of language processing model language processing as integrating both top-down expectations and bottom-up inputs. |
| Approach: | They propose an information-theoretic operationalization for the “quality” of bottom-up information as the mutual information between visual information and word identity. |
| Outcome: | The proposed model compares reading times in English and Chinese in which words' information quality has been reduced by occluding their top or bottom half with full words. |
A Systematic Assessment of Syntactic Generalization in Neural Language Models (2020.acl-main)
Copied to clipboard
| Challenge: | Existing work on syntactic knowledge models has not provided a clear picture of the properties required to produce proper syntaktic generalizations. |
| Approach: | They propose to evaluate syntactic knowledge of language models by varying model architectures . they find substantial differences in syntaktic generalization performance by model architecture . |
| Outcome: | The proposed model architectures outperform other architectures on a set of 34 English-language syntactic test suites. |
Language Models Grow Less Humanlike beyond Phase Transition (2025.acl-long)
Copied to clipboard
| Challenge: | Existing studies have shown that LMs' alignment with human reading behavior improves during pretraining up to a tipping point, beyond which it plateaus or degrades. |
| Approach: | They hypothesize that a pretraining phase transition is responsible for the tipping point in PPP and that phase transitions alter the subsequent learning dynamics of the model, such that further training keeps damaging PPP. |
| Outcome: | The proposed model is able to produce attention patterns that contribute to the degradation of PPP, but it is not capable of producing attention patterns. |
Representation of Constituents in Neural Language Models: Coordination Phrase as a Case Study (D19-1)
Copied to clipboard
| Challenge: | Existing studies have focused on the ability of neural models to compute and employ phrase-level features attached to a set of words, such as subject number or whquestion words. |
| Approach: | They examine whether models can represent constituent-level features, using coordinated noun phrases as a case study. |
| Outcome: | The proposed model can combine gender and gender features to drive downstream expectations, while having less success with gender agreement. |
On the Efficacy of Sampling Adapters (2023.acl-long)
Copied to clipboard
| Challenge: | Using sampling adapters can improve the quality of the generated text. |
| Approach: | They propose a framework for understanding sampling adapters and propose 'sampling adapters' they argue that the shift enforced by them can be viewed as a trade-off between precision and recall . |
| Outcome: | The proposed framework can be used to improve the quality of language models by modifying their distributions to improve their precision and recall. |