Papers by Jonathan Brennan
Large Language Models as Neurolinguistic Subjects: Discrepancy between Performance and Competence (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing evaluations of Large Language Models (LLMs) reflect statistical rules that may not accurately represent LLMs’ true linguistic competence. |
| Approach: | They propose a method that combines minimal pair and diagnostic probing to analyze activation patterns across model layers. |
| Outcome: | The proposed method combines minimal pair and diagnostic probing to analyze activation patterns across model layers. |
Text Genre and Training Data Size in Human-like Parsing (D19-1)
Copied to clipboard
| Challenge: | Using domain-specific training, NLP systems work better, but only when the training examples come from the same textual genre. |
| Approach: | They relate the states of a neural phrase-structure parser to electrophysiological measures from human participants. |
| Outcome: | The proposed model is well-matched to the training data from human participants, but only when the training examples come from the same genre. |
Finding syntax in human encephalography with beam search (P18-1)
Copied to clipboard
| Challenge: | RNNGs are generative models of (tree , string ) pairs that evaluate derivational choices . a non-syntactic neural language model yields no reliable effects . |
| Approach: | They propose to combine a probabilistic generative grammar with a parsing procedure that uses it to manage syntactic derivations as it advances from one word to the next. |
| Outcome: | The proposed model derives two amplitude effects when used against human encephalography data. |
The Alice Datasets: fMRI & EEG Observations of Natural Language Comprehension (2020.lrec-1)
Copied to clipboard
| Challenge: | "naturalistic" stimuli are now offering a new way to study language comprehension in the brain, in synergy with natural language processing tools. |
| Approach: | They propose to use a set of datasets from a story in English to test new linguistic and computational hypotheses about natural language comprehension in the brain. |
| Outcome: | The Alice Datasets are a set of datasets based on magnetic resonance and electrophysiological data, collected while participants heard a story in English. |