Papers by Leon Bergen

6 papers
Word Frequency Does Not Predict Grammatical Knowledge in Language Models (2020.emnlp-main)

Copied to clipboard

Challenge: Neural language models learn the grammatical properties of natural languages to varying degrees of accuracy.
Approach: They focus on subject-verb agreement and reflexive anaphora to investigate whether there are systematic sources of variation in the language models’ accuracy.
Outcome: The proposed model can learn grammatical properties from training data.
Predicting Reference: What do Language Models Learn about Discourse Models? (2020.emnlp-main)

Copied to clipboard

Challenge: a growing literature that probes neural language models to assess their latent acquisition of grammatical knowledge has not investigated their acquisition of discourse modeling ability.
Approach: They draw on a psycholinguistic literature that has established how different contexts affect referential biases concerning who is likely to be referred to next.
Outcome: The proposed models do not resemble human language users, the authors show . their models capture the linguistic knowledge required to perform discourse modeling .
Constraint-based Learning of Phonological Processes (D19-1)

Copied to clipboard

Challenge: Phonological processes govern the way speech sounds in natural languages change depending on context . a novel approach to learning phonological processes from related utterances is proposed .
Approach: They propose an unsupervised approach to learning phonological processes from related utterances . they encode the problem into Boolean constraints that enable data efficiency and fast inference .
Outcome: The proposed approach achieves high accuracy at interactive speeds on phonology problems and datasets.
IR2: Information Regularization for Information Retrieval (2024.lrec-main)

Copied to clipboard

Challenge: Effective information retrieval (IR) in settings with limited training data remains a challenging task.
Approach: They propose a technique for reducing overfitting during synthetic data generation . they use DORIS-MAE, ArguAna, and WhatsThatBook as examples .
Outcome: The proposed technique outperforms previous methods and reduces cost by 50% on three recent IR tasks characterized by complex queries.
Measuring Risk of Bias in Biomedical Reports: The RoBBR Benchmark (2025.emnlp-main)

Copied to clipboard

Challenge: Systematic reviews should take into account the quality of available evidence, placing more weight on studies that use a valid methodology.
Approach: They propose to use a risk-of-bias framework to assess the methodological strength of biomedical papers by combining expert reviewers' judgments with research paper sentences.
Outcome: The proposed system measures the methodological strength of biomedical papers using the risk-of-bias framework used for systematic reviews.
Speakers enhance contextually confusable words (2020.acl-main)

Copied to clipboard

Challenge: Recent work has found that natural languages are shaped by pressures for efficient communication.
Approach: They develop a measure of contextual confusability during word recognition based on psychoacoustic data and apply it to naturalistic speech corpora.
Outcome: The proposed measure of confusability suggests that speakers alter productions to make contextually more confused words easier to understand.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations