Papers by Jonathan Dunn

7 papers
Pre-Trained Language Models Represent Some Geographic Populations Better than Others (2024.lrec-main)

Copied to clipboard

Challenge: Existing studies have focused on measuring the degree to which pre-trained language models capture purely linguistic knowledge and reasoning abilities and world knowledge.
Approach: They use geography to demarcate different populations around the world and comparable corpora to measure how well two families of LLMs perform across these different populations.
Outcome: The results show that pre-trained models perform better for some populations than others.
Validating and Exploring Large Geographic Corpora (2024.lrec-main)

Copied to clipboard

Challenge: a paper examines the impact of corpus creation decisions on multi-lingual web corpora . the goal is to understand the impact on downstream corporata with a focus on under-represented languages and populations.
Approach: This paper evaluates the impact of corpus creation decisions on multi-lingual web corpora . three cleaning methods are used to improve the quality of sub-corpora in the common crawl . the goal is to understand the impact on downstream corporan with a focus on under-represented languages .
Outcome: The results show that the validity of sub-corpora is improved with each stage of cleaning but that this improvement is unevenly distributed across languages and populations.
Geographically-Balanced Gigaword Corpora for 50 Language Varieties (2020.lrec-1)

Copied to clipboard

Challenge: Existing corpora of text corpors over-represent inner-circle varieties from the US and UK . this paper uses country-level population demographics to correct implicit geographic and demographic biases .
Approach: They propose to use country-level population demographics to correct geographic biases . they use a population-based sampling technique to remove geographic bias from gigaword corpora .
Outcome: The proposed corpus family removes geographic biases by comparing the population-based sampling with the baseline corpus.
Language Identification for Austronesian Languages (2022.lrec-1)

Copied to clipboard

Challenge: This paper provides language identification models for low- and under-resourced languages in the Pacific region with a focus on previously unavailable Austronesian languages.
Approach: They compare a classifier based on skip-gram embeddings with other methods . they then increase the number of non-Austronesian languages to 800 to evaluate their performance .
Outcome: The proposed model improves on the previous methods for low- and under-resourced languages in the Pacific region.
Predicting Embedding Reliability in Low-Resource Settings Using Corpus Similarity Measures (2022.lrec-1)

Copied to clipboard

Challenge: a paper aims to evaluate embedding similarity, stability and reliability in low-resource settings . it uses corpus similarity measures before training to predict properties of embeddables .
Approach: They use corpus similarity measures before training to predict properties of embeddings . they then apply the same measures to low-resource settings by modelling reliability . authors hope to use this method to evaluate low-source languages with limited corpus size .
Outcome: The paper shows that it is possible to predict downstream embedding similarity using upstream corpus similarity measures . the main finding is that the measures remain robust on small amounts of training data .
Stability of Syntactic Dialect Classification over Space and Time (2022.coling-1)

Copied to clipboard

Challenge: a paper examines the degree to which dialect classifiers remain stable over time . it finds that the models remain robust over time with a fixed decay rate .
Approach: They construct a test set for 12 dialects of English that spans three years at monthly intervals with a fixed spatial distribution across 1,120 cities.
Outcome: The proposed model can reveal linguistic variation over space and time.
Geographically-Informed Language Identification (2024.lrec-main)

Copied to clipboard

Challenge: a paper develops a method to identify languages based on geographic origin of text . the model is based in regions where languages are widely spoken and may occur anywhere .
Approach: They propose to incorporate geographic information into a language identification model to ensure coverage of linguae francae regardless of location.
Outcome: The proposed model includes 31 widely-spoken international languages . the model improves on social media data and improves performance on 916 languages compared to baseline models .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations