Papers by Jonathan Dunn
Pre-Trained Language Models Represent Some Geographic Populations Better than Others (2024.lrec-main)
Copied to clipboard
| Challenge: | Existing studies have focused on measuring the degree to which pre-trained language models capture purely linguistic knowledge and reasoning abilities and world knowledge. |
| Approach: | They use geography to demarcate different populations around the world and comparable corpora to measure how well two families of LLMs perform across these different populations. |
| Outcome: | The results show that pre-trained models perform better for some populations than others. |
Validating and Exploring Large Geographic Corpora (2024.lrec-main)
Copied to clipboard
| Challenge: | a paper examines the impact of corpus creation decisions on multi-lingual web corpora . the goal is to understand the impact on downstream corporata with a focus on under-represented languages and populations. |
| Approach: | This paper evaluates the impact of corpus creation decisions on multi-lingual web corpora . three cleaning methods are used to improve the quality of sub-corpora in the common crawl . the goal is to understand the impact on downstream corporan with a focus on under-represented languages . |
| Outcome: | The results show that the validity of sub-corpora is improved with each stage of cleaning but that this improvement is unevenly distributed across languages and populations. |
Geographically-Balanced Gigaword Corpora for 50 Language Varieties (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing corpora of text corpors over-represent inner-circle varieties from the US and UK . this paper uses country-level population demographics to correct implicit geographic and demographic biases . |
| Approach: | They propose to use country-level population demographics to correct geographic biases . they use a population-based sampling technique to remove geographic bias from gigaword corpora . |
| Outcome: | The proposed corpus family removes geographic biases by comparing the population-based sampling with the baseline corpus. |
Language Identification for Austronesian Languages (2022.lrec-1)
Copied to clipboard
| Challenge: | This paper provides language identification models for low- and under-resourced languages in the Pacific region with a focus on previously unavailable Austronesian languages. |
| Approach: | They compare a classifier based on skip-gram embeddings with other methods . they then increase the number of non-Austronesian languages to 800 to evaluate their performance . |
| Outcome: | The proposed model improves on the previous methods for low- and under-resourced languages in the Pacific region. |
Predicting Embedding Reliability in Low-Resource Settings Using Corpus Similarity Measures (2022.lrec-1)
Copied to clipboard
| Challenge: | a paper aims to evaluate embedding similarity, stability and reliability in low-resource settings . it uses corpus similarity measures before training to predict properties of embeddables . |
| Approach: | They use corpus similarity measures before training to predict properties of embeddings . they then apply the same measures to low-resource settings by modelling reliability . authors hope to use this method to evaluate low-source languages with limited corpus size . |
| Outcome: | The paper shows that it is possible to predict downstream embedding similarity using upstream corpus similarity measures . the main finding is that the measures remain robust on small amounts of training data . |
Stability of Syntactic Dialect Classification over Space and Time (2022.coling-1)
Copied to clipboard
| Challenge: | a paper examines the degree to which dialect classifiers remain stable over time . it finds that the models remain robust over time with a fixed decay rate . |
| Approach: | They construct a test set for 12 dialects of English that spans three years at monthly intervals with a fixed spatial distribution across 1,120 cities. |
| Outcome: | The proposed model can reveal linguistic variation over space and time. |
Geographically-Informed Language Identification (2024.lrec-main)
Copied to clipboard
| Challenge: | a paper develops a method to identify languages based on geographic origin of text . the model is based in regions where languages are widely spoken and may occur anywhere . |
| Approach: | They propose to incorporate geographic information into a language identification model to ensure coverage of linguae francae regardless of location. |
| Outcome: | The proposed model includes 31 widely-spoken international languages . the model improves on social media data and improves performance on 916 languages compared to baseline models . |