Papers by David Mortensen

13 papers
Zero-shot Learning for Grapheme to Phoneme Conversion with Language Ensemble (2022.findings-acl)

Copied to clipboard

Challenge: Existing work focuses on low-resource and endangered languages with limited training sets.
Approach: They propose a hypothesis set for any unseen target language and combine it with a confusion network to propose 'the most likely hypothesis' they test the approach on over 600 unseened languages and demonstrate it significantly outperforms baselines.
Outcome: The proposed model outperforms baselines on over 600 unseen languages.
Zero-Shot Cross-Lingual NER Using Phonemic Representations for Low-Resource Languages (2024.emnlp-main)

Copied to clipboard

Challenge: Existing zero-shot cross-lingual NER approaches require substantial prior knowledge of the target language, which is impractical for low-resource languages.
Approach: They propose a phonemic representation based on the International Phonetic Alphabet (IPA) to bridge the gap between representations of different languages.
Outcome: The proposed method outperforms baseline models in low-resource languages with highest average F1 score and lowest standard deviation.
Wav2Gloss: Generating Interlinear Glossed Text from Speech (2024.acl-long)

Copied to clipboard

Challenge: Interlinear Glossed Text (IGT) is a form of linguistic annotation that can support documentation and resource creation for endangered languages.
Approach: They propose a task in which these four annotation components are extracted automatically from speech and introduce a dataset to lay the groundwork for future research on IGT generation from speech.
Outcome: The proposed dataset provides the first dataset to lay the groundwork for future research on IGT generation from speech, including end-to-end versus cascaded, monolingual versus multilingual, and single-task versus multiple-task approaches.
Learning the Ordering of Coordinate Compounds and Elaborate Expressions in Hmong, Lahu, and Chinese (2022.naacl-main)

Copied to clipboard

Challenge: phonological hierarchies that predict coordinate constructions are often phonetically “natural” . a neural sequence labeling model can learn elaborate expressions in Hmong without using phonology information.
Approach: They propose that coordinate compounds and elaborate expressions can be learned empirically by phonological hierarchies and a neural sequence labeling model can learn the ordering of elaborate expression in Hmong without using phonology.
Outcome: The proposed models beat strong baselines for all three languages and learn hierarchies similar to those proposed by Mortensen.
Linear Script Representations in Speech Foundation Models Enable Zero-Shot Transliteration (2026.findings-acl)

Copied to clipboard

Challenge: We show that script information is linearly encoded in the activation space of multilingual speech models . modifying activations at inference time induces script change even in unconventional pairings .
Approach: They propose to add script vectors to activations at test time to induce script change . they also show that script information is linearly encoded in the activation space of multilingual speech models .
Outcome: The proposed approach can induce script change even in unconventional language-script pairings.
XferBench: a Data-Driven Benchmark for Emergent Language (2024.naacl-long)

Copied to clipboard

Challenge: Existing methods to teach models to "language" are full of bias, toxicity, and potential intellectual property violations.
Approach: They propose a benchmark for evaluating the overall quality of emergent languages using data-driven methods.
Outcome: The proposed benchmark is based on utterances from the emergent language and is validated using human, synthetic, and emergentic language baselines.
[b] = [d] - [t] + [p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic (2026.findings-acl)

Copied to clipboard

Challenge: Existing studies on how self-supervised speech models encode rich phonetic information have not explored how they are structured.
Approach: They conduct a comprehensive analysis of the underlying structure of S3M representations with particular attention to phonological vectors.
Outcome: The proposed model encodes phonologically interpretable and compositional vectors, demonstrating phonology vector arithmetic.
Do All Languages Cost the Same? Tokenization in the Era of Commercial Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: Language models have evolved from being research prototypes to commercialized products offered as web APIs.
Approach: They conduct a systematic analysis of the cost and utility of OpenAI’s language model API on multilingual benchmarks in 22 typologically diverse languages.
Outcome: The proposed language model API performs poorly on multiple languages and speakers of a large number of languages are overcharged while obtaining poorer results.
Counting the Bugs in ChatGPT’s Wugs: A Multilingual Investigation into the Morphological Capabilities of a Large Language Model (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies on large language models (LLMs) ignore the remarkable ability of humans to generalize and focus only on English.
Approach: They conduct the first rigorous analysis of the morphological capabilities of ChatGPT in four typologically varied languages.
Outcome: The proposed model massively underperforms purpose-built systems, particularly in English.
Parser combinators for Tigrinya and Oromo morphology (L18-1)

Copied to clipboard

Challenge: morphological parsers for two Afroasiatic languages are developed using a parser-combinator paradigm . the paradigm allows rapid development and ease of integration with other systems, but at a cost of non-optimal theoretical efficiency.
Approach: They propose a rule-based morphological parser paradigm for Tigrinya and Oromo languages . they use a parsers-combinator paradigm instead of a finite-state paradigm .
Outcome: The proposed paradigm allows rapid development and ease of integration with other systems, but at cost of non-optimal theoretical efficiency.
DialUp! Modeling the Language Continuum by Adapting Models to Dialects and Dialects to Models (2025.acl-long)

Copied to clipboard

Challenge: Recent advances in MT quality and language coverage have shown that language varieties with low baseline performance are more likely to benefit from these approaches.
Approach: They propose a training-time technique for adapting a pretrained model to dialectal data and an inference-time intervention adapting dialectal datasets to the model expertise.
Outcome: The proposed model shows significant performance gains for several dialects from four language families, and modest gains for two other language families.
Calibrated Seq2seq Models for Efficient and Generalizable Ultra-fine Entity Typing (2023.findings-emnlp)

Copied to clipboard

Challenge: CASENT predicts ultra-fine entities mentioned in text into types with calibrated confidence scores.
Approach: They propose a model that predicts ultra-fine entities with calibrated confidence scores for entity typing.
Outcome: The proposed model outperforms existing models in terms of F1 score and calibration error while achieving 50 times faster inference speed.
Quantifying Cognitive Factors in Lexical Decline (2021.tacl-1)

Copied to clipboard

Challenge: Existing studies on lexical decline suggest that cognitive and linguistic factors play a role in the survival of words and their success in the linguistic ecosystem.
Approach: They propose a variety of psycholinguistic factors that are predictive of lexical decline, in which words greatly decrease in frequency over time.
Outcome: The proposed factors show significant differences in the expected direction between each curated set of declining words and their matched stable words.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations