Papers by David Mortensen
Zero-shot Learning for Grapheme to Phoneme Conversion with Language Ensemble (2022.findings-acl)
Copied to clipboard
| Challenge: | Existing work focuses on low-resource and endangered languages with limited training sets. |
| Approach: | They propose a hypothesis set for any unseen target language and combine it with a confusion network to propose 'the most likely hypothesis' they test the approach on over 600 unseened languages and demonstrate it significantly outperforms baselines. |
| Outcome: | The proposed model outperforms baselines on over 600 unseen languages. |
Zero-Shot Cross-Lingual NER Using Phonemic Representations for Low-Resource Languages (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing zero-shot cross-lingual NER approaches require substantial prior knowledge of the target language, which is impractical for low-resource languages. |
| Approach: | They propose a phonemic representation based on the International Phonetic Alphabet (IPA) to bridge the gap between representations of different languages. |
| Outcome: | The proposed method outperforms baseline models in low-resource languages with highest average F1 score and lowest standard deviation. |
Wav2Gloss: Generating Interlinear Glossed Text from Speech (2024.acl-long)
Copied to clipboard
Taiqi He, Kwanghee Choi, Lindia Tjuatja, Nathaniel Robinson, Jiatong Shi, Shinji Watanabe, Graham Neubig, David Mortensen, Lori Levin
| Challenge: | Interlinear Glossed Text (IGT) is a form of linguistic annotation that can support documentation and resource creation for endangered languages. |
| Approach: | They propose a task in which these four annotation components are extracted automatically from speech and introduce a dataset to lay the groundwork for future research on IGT generation from speech. |
| Outcome: | The proposed dataset provides the first dataset to lay the groundwork for future research on IGT generation from speech, including end-to-end versus cascaded, monolingual versus multilingual, and single-task versus multiple-task approaches. |
Learning the Ordering of Coordinate Compounds and Elaborate Expressions in Hmong, Lahu, and Chinese (2022.naacl-main)
Copied to clipboard
| Challenge: | phonological hierarchies that predict coordinate constructions are often phonetically “natural” . a neural sequence labeling model can learn elaborate expressions in Hmong without using phonology information. |
| Approach: | They propose that coordinate compounds and elaborate expressions can be learned empirically by phonological hierarchies and a neural sequence labeling model can learn the ordering of elaborate expression in Hmong without using phonology. |
| Outcome: | The proposed models beat strong baselines for all three languages and learn hierarchies similar to those proposed by Mortensen. |
Linear Script Representations in Speech Foundation Models Enable Zero-Shot Transliteration (2026.findings-acl)
Copied to clipboard
Ryan Soh-Eun Shim, Kwanghee Choi, Kalvin Chang, Ming-Hao Hsu, Florian Eichin, Zhizheng Wu, Alane Suhr, Michael A. Hedderich, David Harwath, David R. Mortensen, Barbara Plank
| Challenge: | We show that script information is linearly encoded in the activation space of multilingual speech models . modifying activations at inference time induces script change even in unconventional pairings . |
| Approach: | They propose to add script vectors to activations at test time to induce script change . they also show that script information is linearly encoded in the activation space of multilingual speech models . |
| Outcome: | The proposed approach can induce script change even in unconventional language-script pairings. |
XferBench: a Data-Driven Benchmark for Emergent Language (2024.naacl-long)
Copied to clipboard
| Challenge: | Existing methods to teach models to "language" are full of bias, toxicity, and potential intellectual property violations. |
| Approach: | They propose a benchmark for evaluating the overall quality of emergent languages using data-driven methods. |
| Outcome: | The proposed benchmark is based on utterances from the emergent language and is validated using human, synthetic, and emergentic language baselines. |
[b] = [d] - [t] + [p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing studies on how self-supervised speech models encode rich phonetic information have not explored how they are structured. |
| Approach: | They conduct a comprehensive analysis of the underlying structure of S3M representations with particular attention to phonological vectors. |
| Outcome: | The proposed model encodes phonologically interpretable and compositional vectors, demonstrating phonology vector arithmetic. |
Do All Languages Cost the Same? Tokenization in the Era of Commercial Language Models (2023.emnlp-main)
Copied to clipboard
Orevaoghene Ahia, Sachin Kumar, Hila Gonen, Jungo Kasai, David Mortensen, Noah Smith, Yulia Tsvetkov
| Challenge: | Language models have evolved from being research prototypes to commercialized products offered as web APIs. |
| Approach: | They conduct a systematic analysis of the cost and utility of OpenAI’s language model API on multilingual benchmarks in 22 typologically diverse languages. |
| Outcome: | The proposed language model API performs poorly on multiple languages and speakers of a large number of languages are overcharged while obtaining poorer results. |
Counting the Bugs in ChatGPT’s Wugs: A Multilingual Investigation into the Morphological Capabilities of a Large Language Model (2023.emnlp-main)
Copied to clipboard
Leonie Weissweiler, Valentin Hofmann, Anjali Kantharuban, Anna Cai, Ritam Dutt, Amey Hengle, Anubha Kabra, Atharva Kulkarni, Abhishek Vijayakumar, Haofei Yu, Hinrich Schuetze, Kemal Oflazer, David Mortensen
| Challenge: | Existing studies on large language models (LLMs) ignore the remarkable ability of humans to generalize and focus only on English. |
| Approach: | They conduct the first rigorous analysis of the morphological capabilities of ChatGPT in four typologically varied languages. |
| Outcome: | The proposed model massively underperforms purpose-built systems, particularly in English. |
Parser combinators for Tigrinya and Oromo morphology (L18-1)
Copied to clipboard
Patrick Littell, Tom McCoy, Na-Rae Han, Shruti Rijhwani, Zaid Sheikh, David Mortensen, Teruko Mitamura, Lori Levin
| Challenge: | morphological parsers for two Afroasiatic languages are developed using a parser-combinator paradigm . the paradigm allows rapid development and ease of integration with other systems, but at a cost of non-optimal theoretical efficiency. |
| Approach: | They propose a rule-based morphological parser paradigm for Tigrinya and Oromo languages . they use a parsers-combinator paradigm instead of a finite-state paradigm . |
| Outcome: | The proposed paradigm allows rapid development and ease of integration with other systems, but at cost of non-optimal theoretical efficiency. |
DialUp! Modeling the Language Continuum by Adapting Models to Dialects and Dialects to Models (2025.acl-long)
Copied to clipboard
Niyati Bafna, Emily Chang, Nathaniel Romney Robinson, David R. Mortensen, Kenton Murray, David Yarowsky, Hale Sirin
| Challenge: | Recent advances in MT quality and language coverage have shown that language varieties with low baseline performance are more likely to benefit from these approaches. |
| Approach: | They propose a training-time technique for adapting a pretrained model to dialectal data and an inference-time intervention adapting dialectal datasets to the model expertise. |
| Outcome: | The proposed model shows significant performance gains for several dialects from four language families, and modest gains for two other language families. |
Calibrated Seq2seq Models for Efficient and Generalizable Ultra-fine Entity Typing (2023.findings-emnlp)
Copied to clipboard
| Challenge: | CASENT predicts ultra-fine entities mentioned in text into types with calibrated confidence scores. |
| Approach: | They propose a model that predicts ultra-fine entities with calibrated confidence scores for entity typing. |
| Outcome: | The proposed model outperforms existing models in terms of F1 score and calibration error while achieving 50 times faster inference speed. |
Quantifying Cognitive Factors in Lexical Decline (2021.tacl-1)
Copied to clipboard
| Challenge: | Existing studies on lexical decline suggest that cognitive and linguistic factors play a role in the survival of words and their success in the linguistic ecosystem. |
| Approach: | They propose a variety of psycholinguistic factors that are predictive of lexical decline, in which words greatly decrease in frequency over time. |
| Outcome: | The proposed factors show significant differences in the expected direction between each curated set of declining words and their matched stable words. |