Papers by Leon Bergen
Word Frequency Does Not Predict Grammatical Knowledge in Language Models (2020.emnlp-main)
Copied to clipboard
| Challenge: | Neural language models learn the grammatical properties of natural languages to varying degrees of accuracy. |
| Approach: | They focus on subject-verb agreement and reflexive anaphora to investigate whether there are systematic sources of variation in the language models’ accuracy. |
| Outcome: | The proposed model can learn grammatical properties from training data. |
Predicting Reference: What do Language Models Learn about Discourse Models? (2020.emnlp-main)
Copied to clipboard
| Challenge: | a growing literature that probes neural language models to assess their latent acquisition of grammatical knowledge has not investigated their acquisition of discourse modeling ability. |
| Approach: | They draw on a psycholinguistic literature that has established how different contexts affect referential biases concerning who is likely to be referred to next. |
| Outcome: | The proposed models do not resemble human language users, the authors show . their models capture the linguistic knowledge required to perform discourse modeling . |
Constraint-based Learning of Phonological Processes (D19-1)
Copied to clipboard
| Challenge: | Phonological processes govern the way speech sounds in natural languages change depending on context . a novel approach to learning phonological processes from related utterances is proposed . |
| Approach: | They propose an unsupervised approach to learning phonological processes from related utterances . they encode the problem into Boolean constraints that enable data efficiency and fast inference . |
| Outcome: | The proposed approach achieves high accuracy at interactive speeds on phonology problems and datasets. |
IR2: Information Regularization for Information Retrieval (2024.lrec-main)
Copied to clipboard
| Challenge: | Effective information retrieval (IR) in settings with limited training data remains a challenging task. |
| Approach: | They propose a technique for reducing overfitting during synthetic data generation . they use DORIS-MAE, ArguAna, and WhatsThatBook as examples . |
| Outcome: | The proposed technique outperforms previous methods and reduces cost by 50% on three recent IR tasks characterized by complex queries. |
Measuring Risk of Bias in Biomedical Reports: The RoBBR Benchmark (2025.emnlp-main)
Copied to clipboard
Jianyou Wang, Weili Cao, Longtian Bao, Youze Zheng, Gil Pasternak, Kaicheng Wang, Xiaoyue Wang, Ramamohan Paturi, Leon Bergen
| Challenge: | Systematic reviews should take into account the quality of available evidence, placing more weight on studies that use a valid methodology. |
| Approach: | They propose to use a risk-of-bias framework to assess the methodological strength of biomedical papers by combining expert reviewers' judgments with research paper sentences. |
| Outcome: | The proposed system measures the methodological strength of biomedical papers using the risk-of-bias framework used for systematic reviews. |
Speakers enhance contextually confusable words (2020.acl-main)
Copied to clipboard
| Challenge: | Recent work has found that natural languages are shaped by pressures for efficient communication. |
| Approach: | They develop a measure of contextual confusability during word recognition based on psychoacoustic data and apply it to naturalistic speech corpora. |
| Outcome: | The proposed measure of confusability suggests that speakers alter productions to make contextually more confused words easier to understand. |