Papers by Brian Dillon
UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu (2026.findings-acl)
Copied to clipboard
| Challenge: | Evaluating how large language models capture grammatical structure of low-resource languages remains underexplored. |
| Approach: | They evaluate a set of 5,696 minimal pairs that contrast grammatical acceptability across ten core syntactic and morpho-syntactical phenomena in Urdu. |
| Outcome: | The proposed framework compares multilingual models with the proprietary model . the proposed framework achieves the highest average accuracy on regular phenomena . |
Memory efficiency and resource-rational encoding in sentence processing (2026.acl-long)
Copied to clipboard
| Challenge: | Existing studies have shown that language models need to be constrained in their use of working memory for context, the analogue to human working memory (WM). |
| Approach: | They propose to inject noise into hidden representations of Transformer-based LMs to capture constraint on memory encoding. |
| Outcome: | The proposed model improves alignment with human reading times and makes them more compressed and categorical. |