Papers by David Yarowsky
Deciphering and Characterizing Out-of-Vocabulary Words for Morphologically Rich Languages (2022.coling-1)
Copied to clipboard
| Challenge: | a detailed empirical case study of out-of-vocabulary words in modern text is presented . unfamiliar words cause trouble for machine processing or comprehension of text, authors say . |
| Approach: | They propose a detailed empirical case study of the nature of out-of-vocabulary words encountered in modern text in a moderate-resource language such as Bulgarian . they apply a multi-faceted distributional analysis of the underlying word-formation processes to characterize the residual vocabulary . |
| Outcome: | The proposed method can be used to aid in compositional translation, parsing, language modeling, and other NLP tasks. |
Measuring the Similarity of Grammatical Gender Systems by Comparing Partitions (2020.emnlp-main)
Copied to clipboard
| Challenge: | A grammatical gender system divides a lexicon into a small number of fixed categories with fixed usage across speakers. |
| Approach: | They propose to define gender systems extensionally to reduce comparisons to cluster evaluation by comparing pairwise overlaps between gender systems. |
| Outcome: | The proposed measures are based on a phylogenetic tree over extant Indo-European languages. |
Wiktionary Normalization of Translations and Morphological Information (2020.coling-main)
Copied to clipboard
| Challenge: | We extend the Yawipa Wiktionary Parser to extract and normalize translations from etymology glosses and morphological form-of relations. |
| Approach: | They extend Yawipa to extract and normalize translations from etymology glosses . they propose a method to identify typos in translation annotations based on extracted morphological data . |
| Outcome: | The proposed method improves on a standard attention baseline by using copy attention. |
Neural Transduction for Multilingual Lexical Translation (2020.coling-main)
Copied to clipboard
| Challenge: | a method for completing multilingual translation dictionaries is proposed . a 27% relative improvement in whole-word accuracy is achieved when multilingual data is unavailable . |
| Approach: | They propose a method for completing multilingual translation dictionaries using multilingual inputs and multilingual decoding objective. |
| Outcome: | The proposed method can synthesize new word forms in multilingual translation dictionaries . it can perform in settings where correct translations have not been observed in text . |
Creating Large-Scale Multilingual Cognate Tables (L18-1)
Copied to clipboard
| Challenge: | Low-resource languages often suffer from a lack of high-coverage lexical resources. |
| Approach: | They propose a method to generate cognate tables by clustering words from existing lexical resources. |
| Outcome: | The proposed method outperforms baselines on the Romance and Turkic language families. |
Sequence Models for Computational Etymology of Borrowings (2021.findings-acl)
Copied to clipboard
| Challenge: | a computational model of word borrowing can be useful for lexicon expansion and language preservation. |
| Approach: | They propose to use neural sequence models to model word borrowings from a donor word to an incorporated word. |
| Outcome: | The proposed model beats baseline models in both directions, with the quantity of data strongly influencing performance. |
The Johns Hopkins University Bible Corpus: 1600+ Tongues for Typological Exploration (2020.lrec-1)
Copied to clipboard
Arya D. McCarthy, Rachel Wicks, Dylan Lewis, Aaron Mueller, Winston Wu, Oliver Adams, Garrett Nicolai, Matt Post, David Yarowsky
| Challenge: | Our corpus spans 1611 diverse written languages, with constituents of more than 90 language families. |
| Approach: | They propose to scrape and merge online resources and merge them with existing corpora to create a verse-parallel scheme for all translations. |
| Outcome: | The results show that the Bible provides high coverage of core vocabulary. |
Massively Multilingual Adversarial Speech Recognition (N19-1)
Copied to clipboard
| Challenge: | Prior work in multilingual and cross-lingual speech recognition has been limited to a subset of the world's most-spoken languages. |
| Approach: | They propose to use phonemes and phonemes as pretraining objectives to encourage language-independent representations. |
| Outcome: | The proposed model is able to learn language-independent representations of speech using multilingual training. |
Evaluating Large Language Models along Dimensions of Language Variation: A Systematik Invesdigatiom uv Cross-lingual Generalization (2024.emnlp-main)
Copied to clipboard
| Challenge: | Xue et al., 2021) show that large language models suffer from performance degradation on unseen closely-related languages and dialects relative to their high-resource language neighbour (HRLN). |
| Approach: | They propose to model phonological, morphological, and lexical distance as Bayesian noise processes to synthesize artificial languages that are controllably distant from the HRLN. |
| Outcome: | The proposed model offers insights on model robustness to isolated and composed linguistic phenomena and the impact of task and HRL characteristics on PD. |
Computational Etymology and Word Emergence (2020.lrec-1)
Copied to clipboard
| Challenge: | etymology is the study of words' origins. |
| Approach: | They develop an extensible Wiktionary parser that predicts the etymology of a word across the full range of ethymological types and languages in Wiktionaries. |
| Outcome: | The proposed parser predicts the etymology of a word across the full range of ethymologies and languages in Wiktionary, and shows the application of tymatics in modeling this phenomenon. |
UniMorph 2.0: Universal Morphology (L18-1)
Copied to clipboard
Christo Kirov, Ryan Cotterell, John Sylak-Glassman, Géraldine Walther, Ekaterina Vylomova, Patrick Xia, Manaal Faruqui, Sabrina J. Mielke, Arya McCarthy, Sandra Kübler, David Yarowsky, Jason Eisner, Mans Hulden
| Challenge: | The Universal Morphology project is a collaborative effort to improve how NLP handles complex morphology across the world's languages. |
| Approach: | They propose to use a universal tagset to annotate morphological data using a schema that includes a lemma and a bundle of morphology features. |
| Outcome: | The project releases annotated morphological data using a universal tagset, the UniMorph schema. |
A Comparative Study of Extremely Low-Resource Transliteration of the World’s Languages (L18-1)
Copied to clipboard
| Challenge: | a phrase-based MT system performs better than other methods for transliterating Bible names . combining data and training a single neural system yields significant gains . |
| Approach: | They compare several machine translation methods for transliterating Bible names . they find a phrase-based MT system performs better than other methods . |
| Outcome: | The phrase-based MT system performs better than other methods, the study finds . but, the single-language system outperforms the phrase-backed MT systems . |
Multilingual Dictionary Based Construction of Core Vocabulary (2020.lrec-1)
Copied to clipboard
| Challenge: | Existing methods for core vocabulary lists for multiple applications are lacking coverage in sparse dictionaries . we propose a new method for definition and construction of core vocabulary sets based on coverage in dictionary dictionaria . |
| Approach: | They propose a functional definition and construction method for core vocabulary sets based on relative coverage of a target concept in bilingual dictionaries. |
| Outcome: | The proposed method achieves high overlap with existing vocabulary lists . it uses a cognate prediction method to recover missing coverage of the vocabulary . |
DialUp! Modeling the Language Continuum by Adapting Models to Dialects and Dialects to Models (2025.acl-long)
Copied to clipboard
Niyati Bafna, Emily Chang, Nathaniel Romney Robinson, David R. Mortensen, Kenton Murray, David Yarowsky, Hale Sirin
| Challenge: | Recent advances in MT quality and language coverage have shown that language varieties with low baseline performance are more likely to benefit from these approaches. |
| Approach: | They propose a training-time technique for adapting a pretrained model to dialectal data and an inference-time intervention adapting dialectal datasets to the model expertise. |
| Outcome: | The proposed model shows significant performance gains for several dialects from four language families, and modest gains for two other language families. |
On the Robustness of Cognate Generation Models (2022.lrec-1)
Copied to clipboard
| Challenge: | We examine different types of noise generated by human errors and how these noisy inputs affect the performance of cognate generation models. |
| Approach: | They evaluate two popular neural cognate generation models’ robustness to human-plausible noise. |
| Outcome: | The proposed models are robust to deletion, duplication, swapping, keyboard errors, and a new type of error, phonological errors. |
Modeling Color Terminology Across Thousands of Languages (D19-1)
Copied to clipboard
| Challenge: | Existing studies on what constitutes a "basic" color term and its acquisition sequence are flawed . a pan-lingual approach may reveal general color trends more reliably than smaller datasets. |
| Approach: | They propose to operationalize and critique the Berlin and Kay color term hypotheses . they use 14 empirically-grounded computational linguistic metrics to analyze cross-linguistic data . |
| Outcome: | The proposed measures correlate strongly with the Berlin and Kay color term partition and their hypothesized universal acquisition sequence. |
An Analysis of Massively Multilingual Neural Machine Translation for Low-Resource Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | In this study, we explore massively multilingual low-resource neural machine translation. |
| Approach: | They propose to use Bible translations to train models with up to 1,107 source languages and create multilingual corpora varying the number and relatedness of source languages. |
| Outcome: | The proposed approach is highly language-specific and can be tailored to the source language and its typology. |
UniMorph 4.0: Universal Morphology (2022.lrec-1)
Copied to clipboard
Khuyagbaatar Batsuren, Omer Goldman, Salam Khalifa, Nizar Habash, Witold Kieraś, Gábor Bella, Brian Leonard, Garrett Nicolai, Kyle Gorman, Yustinus Ghanggo Ate, Maria Ryskina, Sabrina Mielke, Elena Budianskaya, Charbel El-Khaissi, Tiago Pimentel, Michael Gasser, William Abbott Lane, Mohit Raj, Matt Coler, Jaime Rafael Montoya Samame, Delio Siticonatzi Camaiteri, Esaú Zumaeta Rojas, Didier López Francis, Arturo Oncevay, Juan López Bautista, Gema Celeste Silva Villegas, Lucas Torroba Hennigen, Adam Ek, David Guriel, Peter Dirix, Jean-Philippe Bernardy, Andrey Scherbakov, Aziyana Bayyr-ool, Antonios Anastasopoulos, Roberto Zariquiey, Karina Sheifer, Sofya Ganieva, Hilaria Cruz, Ritván Karahóǧa, Stella Markantonatou, George Pavlidis, Matvey Plugaryov, Elena Klyachko, Ali Salehi, Candy Angulo, Jatayu Baxi, Andrew Krizhanovsky, Natalia Krizhanovskaya, Elizabeth Salesky, Clara Vania, Sardana Ivanova, Jennifer White, Rowan Hall Maudslay, Josef Valvoda, Ran Zmigrod, Paula Czarnowska, Irene Nikkarinen, Aelita Salchak, Brijesh Bhatt, Christopher Straughn, Zoey Liu, Jonathan North Washington, Yuval Pinter, Duygu Ataman, Marcin Wolinski, Totok Suhardijanto, Anna Yablonskaya, Niklas Stoehr, Hossep Dolatian, Zahroh Nuriah, Shyam Ratan, Francis M. Tyers, Edoardo M. Ponti, Grant Aiton, Aryaman Arora, Richard J. Hatcher, Ritesh Kumar, Jeremiah Young, Daria Rodionova, Anastasia Yemelina, Taras Andrushko, Igor Marchenko, Polina Mashkovtseva, Alexandra Serova, Emily Prud’hommeaux, Maria Nepomniashchaya, Fausto Giunchiglia, Eleanor Chodroff, Mans Hulden, Miikka Silfverberg, Arya D. McCarthy, David Yarowsky, Ryan Cotterell, Reut Tsarfaty, Ekaterina Vylomova
| Challenge: | The Universal Morphology project provides broad-coverage instantiated morphological inflection tables for hundreds of diverse languages. |
| Approach: | They propose a language-independent feature schema for rich morphological annotation and a type-level resource of annotated data in diverse languages realizing that schema. |
| Outcome: | The proposed schema has added 66 new languages, including 24 endangered languages. |
Massively Translingual Compound Analysis and Translation Discovery (L18-1)
Copied to clipboard
| Challenge: | Morphological compounding is one of the most common and productive methods of word formation across the world's languages. |
| Approach: | They propose a model for compounding using bilingual dictionaries and no annotated training data . they also release a massively multilingual dataset of compound words and their decompositions . |
| Outcome: | The proposed model generates novel translations of English concepts on a multilingual dataset . the model can be applied to a wide range of languages and is highly reproducible. |
Fine-grained Morphosyntactic Analysis and Generation Tools for More Than One Thousand Languages (2020.lrec-1)
Copied to clipboard
| Challenge: | Using morphosyntactic tools, we train and distribute tools for approximately one thousand languages. |
| Approach: | They train and distribute morphosyntactic tools for approximately one thousand languages. |
| Outcome: | The results show that the tools generalize well across rare and common forms alike. |
Learning Morphosyntactic Analyzers from the Bible via Iterative Annotation Projection across 26 Languages (P19-1)
Copied to clipboard
| Challenge: | Currently, computational tools for low-resource languages are limited by a lack of supervised training data. |
| Approach: | They propose to use English taggers and parsers to project morphological information onto translations of the Bible in 26 different test languages. |
| Outcome: | The proposed method reduces lemmatization and morphological analysis over a strong initial system. |
UniMorph 3.0: Universal Morphology (2020.lrec-1)
Copied to clipboard
Arya D. McCarthy, Christo Kirov, Matteo Grella, Amrit Nidhi, Patrick Xia, Kyle Gorman, Ekaterina Vylomova, Sabrina J. Mielke, Garrett Nicolai, Miikka Silfverberg, Timofey Arkhangelskiy, Nataly Krizhanovsky, Andrew Krizhanovsky, Elena Klyachko, Alexey Sorokin, John Mansfield, Valts Ernštreits, Yuval Pinter, Cassandra L. Jacobs, Ryan Cotterell, Mans Hulden, David Yarowsky
| Challenge: | Explicit modeling of morphology has demonstrable benefits for language modeling, speech recognition, word embedding and keyword search. |
| Approach: | They propose a language-independent feature schema for rich morphological annotation and a type-level resource for annotated data in diverse languages. |
| Outcome: | The proposed schema has been improved to make it more complete and correct, and adds 66 new languages and parts of speech for 12 languages. |
Creating a Translation Matrix of the Bible’s Names Across 591 Languages (L18-1)
Copied to clipboard
| Challenge: | In low-resource languages, the Bible is the only significant bilingual, or even monolingual, text available . standard word alignment tools can be noisy, making downstream tasks difficult . a novel resource of 1129 aligned Bible person and place names is developed . |
| Approach: | They propose to use Bible person and place names as a tool for translation and transliteration . they use weighted edit distance, machine translation-based transliterations and affixal induction and transformation models to improve the Bible's output. |
| Outcome: | The proposed model outperforms a widely used word aligner on 97% of test words on multilingual named-entity alignment and translation across 591 languages. |