Papers by Winston Wu

24 papers
Deciphering and Characterizing Out-of-Vocabulary Words for Morphologically Rich Languages (2022.coling-1)

Copied to clipboard

Challenge: a detailed empirical case study of out-of-vocabulary words in modern text is presented . unfamiliar words cause trouble for machine processing or comprehension of text, authors say .
Approach: They propose a detailed empirical case study of the nature of out-of-vocabulary words encountered in modern text in a moderate-resource language such as Bulgarian . they apply a multi-faceted distributional analysis of the underlying word-formation processes to characterize the residual vocabulary .
Outcome: The proposed method can be used to aid in compositional translation, parsing, language modeling, and other NLP tasks.
MOKA: Moral Knowledge Augmentation for Moral Event Extraction (2024.naacl-long)

Copied to clipboard

Challenge: Existing methods for discerning moral values are limited due to lack of context, lack of moral reasoning capabilities and complexity of moral stances.
Approach: They propose a framework for moral event extraction using moral words and moral scenarios.
Outcome: The proposed framework outperforms baselines across three moral event understanding tasks.
Wiktionary Normalization of Translations and Morphological Information (2020.coling-main)

Copied to clipboard

Challenge: We extend the Yawipa Wiktionary Parser to extract and normalize translations from etymology glosses and morphological form-of relations.
Approach: They extend Yawipa to extract and normalize translations from etymology glosses . they propose a method to identify typos in translation annotations based on extracted morphological data .
Outcome: The proposed method improves on a standard attention baseline by using copy attention.
Neural Transduction for Multilingual Lexical Translation (2020.coling-main)

Copied to clipboard

Challenge: a method for completing multilingual translation dictionaries is proposed . a 27% relative improvement in whole-word accuracy is achieved when multilingual data is unavailable .
Approach: They propose a method for completing multilingual translation dictionaries using multilingual inputs and multilingual decoding objective.
Outcome: The proposed method can synthesize new word forms in multilingual translation dictionaries . it can perform in settings where correct translations have not been observed in text .
Creating Large-Scale Multilingual Cognate Tables (L18-1)

Copied to clipboard

Challenge: Low-resource languages often suffer from a lack of high-coverage lexical resources.
Approach: They propose a method to generate cognate tables by clustering words from existing lexical resources.
Outcome: The proposed method outperforms baselines on the Romance and Turkic language families.
Late Fusion with Triplet Margin Objective for Multimodal Ideology Prediction and Analysis (2022.emnlp-main)

Copied to clipboard

Challenge: Prior work on ideology prediction has focused on single modalities, i.e., text or images.
Approach: They propose a task where a model predicts binary or five-point scale ideological leanings given a text-image pair with political content.
Outcome: The proposed model outperforms the state-of-the-art model by almost 4% and a strong multimodal baseline with no pretraining by over 3%.
Cross-Cultural Analysis of Human Values, Morals, and Biases in Folk Tales (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies on folk tales focus on European tales, ignoring large swaths of the world's diverse cultures.
Approach: They compile a corpus of over 1,900 folk tales originating from 27 diverse cultures across six continents and employ lexicon-based correlation analyses to examine human values, morals, and gender biases.
Outcome: The results show that folk tales are influenced by cultural norms and cultural values and are well-known for their morals and values.
Sequence Models for Computational Etymology of Borrowings (2021.findings-acl)

Copied to clipboard

Challenge: a computational model of word borrowing can be useful for lexicon expansion and language preservation.
Approach: They propose to use neural sequence models to model word borrowings from a donor word to an incorporated word.
Outcome: The proposed model beats baseline models in both directions, with the quantity of data strongly influencing performance.
The Johns Hopkins University Bible Corpus: 1600+ Tongues for Typological Exploration (2020.lrec-1)

Copied to clipboard

Challenge: Our corpus spans 1611 diverse written languages, with constituents of more than 90 language families.
Approach: They propose to scrape and merge online resources and merge them with existing corpora to create a verse-parallel scheme for all translations.
Outcome: The results show that the Bible provides high coverage of core vocabulary.
Statistical and Neural Methods for Hawaiian Orthography Modernization (2025.emnlp-main)

Copied to clipboard

Challenge: Hawaiian orthography employs two distinct spelling systems, both of which are used by communities of speakers today.
Approach: They develop models that convert between the ‘okina letter and kahak diacritic, which represent glottal stops and long vowels, respectively.
Outcome: The proposed models outperform neural seq2seq models and LLMs in a low-resource setting, highlighting the potential for traditional machine learning approaches in . low-cost environments.
Computational Etymology and Word Emergence (2020.lrec-1)

Copied to clipboard

Challenge: etymology is the study of words' origins.
Approach: They develop an extensible Wiktionary parser that predicts the etymology of a word across the full range of ethymological types and languages in Wiktionaries.
Outcome: The proposed parser predicts the etymology of a word across the full range of ethymologies and languages in Wiktionary, and shows the application of tymatics in modeling this phenomenon.
A Comparative Study of Extremely Low-Resource Transliteration of the World’s Languages (L18-1)

Copied to clipboard

Challenge: a phrase-based MT system performs better than other methods for transliterating Bible names . combining data and training a single neural system yields significant gains .
Approach: They compare several machine translation methods for transliterating Bible names . they find a phrase-based MT system performs better than other methods .
Outcome: The phrase-based MT system performs better than other methods, the study finds . but, the single-language system outperforms the phrase-backed MT systems .
Multilingual Dictionary Based Construction of Core Vocabulary (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods for core vocabulary lists for multiple applications are lacking coverage in sparse dictionaries . we propose a new method for definition and construction of core vocabulary sets based on coverage in dictionary dictionaria .
Approach: They propose a functional definition and construction method for core vocabulary sets based on relative coverage of a target concept in bilingual dictionaries.
Outcome: The proposed method achieves high overlap with existing vocabulary lists . it uses a cognate prediction method to recover missing coverage of the vocabulary .
Crossing the Aisle: Unveiling Partisan and Counter-Partisan Events in News Reporting (2023.findings-emnlp)

Copied to clipboard

Challenge: Prior work in NLP has only studied media bias via linguistic style and word usage.
Approach: They annotate a dataset containing 8,511 (counter-)partisan event annotations in 304 news articles from ideologically diverse media outlets.
Outcome: The proposed dataset contains 8,511 (counter-)partisan event annotations in 304 news articles from ideologically diverse media outlets.
On the Robustness of Cognate Generation Models (2022.lrec-1)

Copied to clipboard

Challenge: We examine different types of noise generated by human errors and how these noisy inputs affect the performance of cognate generation models.
Approach: They evaluate two popular neural cognate generation models’ robustness to human-plausible noise.
Outcome: The proposed models are robust to deletion, duplication, swapping, keyboard errors, and a new type of error, phonological errors.
Analyzing Occupational Distribution Representation in Japanese Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Recent advances in large language models have enabled users to generate fluent and seemingly convincing text, but they have uneven performance in different languages, which is associated with undesirable societal biases toward marginalized populations.
Approach: They develop three Japanese language prompts to probe LLMs’ understanding of Japanese names and their association between gender and occupations.
Outcome: The proposed models can associate Japanese names with correct gendered occupations when using constrained decoding, but with sampling or greedy decoding they prefer a small set of stereotypically genderes.
Modeling Color Terminology Across Thousands of Languages (D19-1)

Copied to clipboard

Challenge: Existing studies on what constitutes a "basic" color term and its acquisition sequence are flawed . a pan-lingual approach may reveal general color trends more reliably than smaller datasets.
Approach: They propose to operationalize and critique the Berlin and Kay color term hypotheses . they use 14 empirically-grounded computational linguistic metrics to analyze cross-linguistic data .
Outcome: The proposed measures correlate strongly with the Berlin and Kay color term partition and their hypothesized universal acquisition sequence.
An Analysis of Massively Multilingual Neural Machine Translation for Low-Resource Languages (2020.lrec-1)

Copied to clipboard

Challenge: In this study, we explore massively multilingual low-resource neural machine translation.
Approach: They propose to use Bible translations to train models with up to 1,107 source languages and create multilingual corpora varying the number and relatedness of source languages.
Outcome: The proposed approach is highly language-specific and can be tailored to the source language and its typology.
Massively Translingual Compound Analysis and Translation Discovery (L18-1)

Copied to clipboard

Challenge: Morphological compounding is one of the most common and productive methods of word formation across the world's languages.
Approach: They propose a model for compounding using bilingual dictionaries and no annotated training data . they also release a massively multilingual dataset of compound words and their decompositions .
Outcome: The proposed model generates novel translations of English concepts on a multilingual dataset . the model can be applied to a wide range of languages and is highly reproducible.
Evaluating Neural Model Robustness for Machine Comprehension (2021.eacl-main)

Copied to clipboard

Challenge: evaluating model robustness to adversarial attacks can provide deeper understanding of how deep neural networks work and what kind of linguistic information is actually captured by neural networks.
Approach: They propose a method for strategic sentence-level perturbations to evaluate model robustness to adversarial attacks using character and word perturbations.
Outcome: The proposed model improves model performance during adversarial attacks by using ensembles and predicts errors in adversarials.
Fine-grained Morphosyntactic Analysis and Generation Tools for More Than One Thousand Languages (2020.lrec-1)

Copied to clipboard

Challenge: Using morphosyntactic tools, we train and distribute tools for approximately one thousand languages.
Approach: They train and distribute morphosyntactic tools for approximately one thousand languages.
Outcome: The results show that the tools generalize well across rare and common forms alike.
Has It All Been Solved? Open NLP Research Questions Not Solved by Large Language Models (2024.lrec-main)

Copied to clipboard

Challenge: Recent advances in large language models have led to misleading public discourse that “it’s all been solved.”
Approach: They identify 14 research areas encompassing 45 research directions that require new research and are not directly solvable by LLMs.
Outcome: The research areas identified are 45 research directions that require new research and are not directly solvable by LLMs.
You Are What You Annotate: Towards Better Models through Annotator Representations (2023.findings-emnlp)

Copied to clipboard

Challenge: Annotator disagreement is ubiquitous in natural language processing tasks.
Approach: They propose to model annotators' idiosyncrasies and account for their idioms by creating representations for each annotator and their annotations.
Outcome: The proposed model improves on an existing dataset with eight annotators with inherent disagreements while increasing model size by 1%.
Creating a Translation Matrix of the Bible’s Names Across 591 Languages (L18-1)

Copied to clipboard

Challenge: In low-resource languages, the Bible is the only significant bilingual, or even monolingual, text available . standard word alignment tools can be noisy, making downstream tasks difficult . a novel resource of 1129 aligned Bible person and place names is developed .
Approach: They propose to use Bible person and place names as a tool for translation and transliteration . they use weighted edit distance, machine translation-based transliterations and affixal induction and transformation models to improve the Bible's output.
Outcome: The proposed model outperforms a widely used word aligner on 97% of test words on multilingual named-entity alignment and translation across 591 languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations