Papers by David Wu

31 papers
CONQRR: Conversational Query Rewriting for Retrieval with Reinforcement Learning (2022.emnlp-main)

Copied to clipboard

Challenge: Existing models for conversational question answering require specific retrievers to understand user questions.
Approach: They develop a query rewriting model CONQRR that rewrites a conversational question into a standalone question.
Outcome: The proposed model achieves state-of-the-art on an open-domain conversational question answering dataset and is effective for two different off-the shelf retrievers.
Deciphering and Characterizing Out-of-Vocabulary Words for Morphologically Rich Languages (2022.coling-1)

Copied to clipboard

Challenge: a detailed empirical case study of out-of-vocabulary words in modern text is presented . unfamiliar words cause trouble for machine processing or comprehension of text, authors say .
Approach: They propose a detailed empirical case study of the nature of out-of-vocabulary words encountered in modern text in a moderate-resource language such as Bulgarian . they apply a multi-faceted distributional analysis of the underlying word-formation processes to characterize the residual vocabulary .
Outcome: The proposed method can be used to aid in compositional translation, parsing, language modeling, and other NLP tasks.
Wiktionary Normalization of Translations and Morphological Information (2020.coling-main)

Copied to clipboard

Challenge: We extend the Yawipa Wiktionary Parser to extract and normalize translations from etymology glosses and morphological form-of relations.
Approach: They extend Yawipa to extract and normalize translations from etymology glosses . they propose a method to identify typos in translation annotations based on extracted morphological data .
Outcome: The proposed method improves on a standard attention baseline by using copy attention.
How Far can 100 Samples Go? Unlocking Zero-Shot Translation with Tiny Multi-Parallel Data (2024.findings-acl)

Copied to clipboard

Challenge: a common solution to zero-shot translation is to add as many related translation directions as possible to the training corpus.
Approach: They show that a small amount of multi-parallel data can achieve significant zero-shot improvements . they say that the resulting non-English performance is close to the complete translation upper bound .
Outcome: The proposed model achieves +21.7 ChrF++ non-English translation improvements on EC30 dataset . the resulting non- English performance exceeds M2M100 by an average of 5.9 ChrF+ .
Neural Transduction for Multilingual Lexical Translation (2020.coling-main)

Copied to clipboard

Challenge: a method for completing multilingual translation dictionaries is proposed . a 27% relative improvement in whole-word accuracy is achieved when multilingual data is unavailable .
Approach: They propose a method for completing multilingual translation dictionaries using multilingual inputs and multilingual decoding objective.
Outcome: The proposed method can synthesize new word forms in multilingual translation dictionaries . it can perform in settings where correct translations have not been observed in text .
Image Difference Captioning via Adversarial Preference Optimization (2025.emnlp-main)

Copied to clipboard

Challenge: Existing supervised approaches to image difference captioning overfit to dataset-specific language patterns and fail to capture accurate preferences.
Approach: They propose an adversarial direct preference optimization framework that aligns captioning policy with pairwise difference preferences via Direct Preference Optimization.
Outcome: The proposed approach outperforms baselines on benchmark IDC datasets in generating fine-grained and accurate difference descriptions.
Beyond Label Attention: Transparency in Language Models for Automated Medical Coding via Dictionary Learning (2024.emnlp-main)

Copied to clipboard

Challenge: Current efforts in interpretability of medical coding rely heavily on label attention mechanisms, which often leads to the highlighting of extraneous tokens irrelevant to the ICD code.
Approach: They propose to leverage dictionary learning to extract sparsely activated representations from dense language models embedded in superposition to facilitate accurate interpretability.
Outcome: The proposed model extracts sparsely activated representations from dense language models in superposition, even when the highlighted tokens are medically irrelevant.
Creating Large-Scale Multilingual Cognate Tables (L18-1)

Copied to clipboard

Challenge: Low-resource languages often suffer from a lack of high-coverage lexical resources.
Approach: They propose a method to generate cognate tables by clustering words from existing lexical resources.
Outcome: The proposed method outperforms baselines on the Romance and Turkic language families.
Linear Script Representations in Speech Foundation Models Enable Zero-Shot Transliteration (2026.findings-acl)

Copied to clipboard

Challenge: We show that script information is linearly encoded in the activation space of multilingual speech models . modifying activations at inference time induces script change even in unconventional pairings .
Approach: They propose to add script vectors to activations at test time to induce script change . they also show that script information is linearly encoded in the activation space of multilingual speech models .
Outcome: The proposed approach can induce script change even in unconventional language-script pairings.
ACT2: A multi-disciplinary semi-structured dataset for importance and purpose classification of citations (2022.lrec-1)

Copied to clipboard

Challenge: Existing methods for classifying citations rely on bibliometric measures to consider the semantics of citation.
Approach: They propose to use a Citation Context Classification (3C) shared task dataset to classify citations according to their purpose and importance.
Outcome: The proposed model can be used to link research works to graphs and enable efficient knowledge discovery.
Sequence Models for Computational Etymology of Borrowings (2021.findings-acl)

Copied to clipboard

Challenge: a computational model of word borrowing can be useful for lexicon expansion and language preservation.
Approach: They propose to use neural sequence models to model word borrowings from a donor word to an incorporated word.
Outcome: The proposed model beats baseline models in both directions, with the quantity of data strongly influencing performance.
Detecting Latin in Historical Books with Large Language Models: A Multimodal Benchmark (2026.eacl-long)

Copied to clipboard

Challenge: This paper presents a novel task of extracting low-resourced and noisy Latin fragments from mixed-language historical documents with varied layouts.
Approach: They propose to extract low-resourced and noisy Latin fragments from mixed-language historical documents with varied layouts using a multimodal dataset.
Outcome: The proposed model lacks a functional comprehension of Latin, but reliable detection is achievable with zero-shot models.
The Johns Hopkins University Bible Corpus: 1600+ Tongues for Typological Exploration (2020.lrec-1)

Copied to clipboard

Challenge: Our corpus spans 1611 diverse written languages, with constituents of more than 90 language families.
Approach: They propose to scrape and merge online resources and merge them with existing corpora to create a verse-parallel scheme for all translations.
Outcome: The results show that the Bible provides high coverage of core vocabulary.
Akan Cinematic Emotions (ACE): A Multimodal Multi-party Dataset for Emotion Recognition in Movie Dialogues (2025.findings-acl)

Copied to clipboard

Challenge: Akan Cinematic Emotions (AkaCE) is the first multimodal emotion dialogue dataset for an African language . it contains 385 emotion-labeled dialogues and 6162 utterances across audio, visual, and textual modalities, along with word-level prosodic prominence annotations.
Approach: They propose to use AkaCE to analyze African cinematic emotions using word-level prosodic prominence annotations.
Outcome: The Akan Cinematic Emotions (AkaCE) dataset addresses the significant lack of resources for low-resource languages in emotion recognition research.
CReSE: Benchmark Data and Automatic Evaluation Framework for Recommending Eligibility Criteria from Clinical Trial Information (2024.findings-eacl)

Copied to clipboard

Challenge: Eligibility criteria (EC) are defined as a set of conditions an individual must meet to participate in a clinical trial.
Approach: They propose to recommend EC based on clinical trial information, including trial titles, and introduce an automatic evaluation framework to assess clinical validity of the EC recommendation model.
Outcome: The proposed model outperforms existing language models pre-trained on the biomedical domain in EC clustering.
Computational Etymology and Word Emergence (2020.lrec-1)

Copied to clipboard

Challenge: etymology is the study of words' origins.
Approach: They develop an extensible Wiktionary parser that predicts the etymology of a word across the full range of ethymological types and languages in Wiktionaries.
Outcome: The proposed parser predicts the etymology of a word across the full range of ethymologies and languages in Wiktionary, and shows the application of tymatics in modeling this phenomenon.
A Comparative Study of Extremely Low-Resource Transliteration of the World’s Languages (L18-1)

Copied to clipboard

Challenge: a phrase-based MT system performs better than other methods for transliterating Bible names . combining data and training a single neural system yields significant gains .
Approach: They compare several machine translation methods for transliterating Bible names . they find a phrase-based MT system performs better than other methods .
Outcome: The phrase-based MT system performs better than other methods, the study finds . but, the single-language system outperforms the phrase-backed MT systems .
Multilingual Dictionary Based Construction of Core Vocabulary (2020.lrec-1)

Copied to clipboard

Challenge: Existing methods for core vocabulary lists for multiple applications are lacking coverage in sparse dictionaries . we propose a new method for definition and construction of core vocabulary sets based on coverage in dictionary dictionaria .
Approach: They propose a functional definition and construction method for core vocabulary sets based on relative coverage of a target concept in bilingual dictionaries.
Outcome: The proposed method achieves high overlap with existing vocabulary lists . it uses a cognate prediction method to recover missing coverage of the vocabulary .
Fine-grained Artificial Neurons in Audio-transformers for Disentangling Neural Auditory Encoding (2023.findings-acl)

Copied to clipboard

Challenge: Existing studies treat each transformer encoding layer as a single artificial neuron . layer-level embeddings aggregate multiple types of contextual attention captured by multiple head modules .
Approach: They propose to embed each transformer encoding layer as a single artificial neuron . they propose to couple those ANs with their biological-neuron counterparts in the human brain .
Outcome: The proposed models can be used to link representations to brain activity, the authors say . their results show that the proposed models carry meaningful neurolinguistic information .
Academics Can Contribute to Domain-Specialized Language Models (2024.emnlp-main)

Copied to clipboard

Challenge: Commercially available models dominate academic leaderboards, focusing on creating and adapting general-purpose models . however, general- purpose models often underperform in specialized domains, and domain-specific models yield superior results.
Approach: They advocate for a renewed focus on developing and evaluating domain- and task-specific models . they advocate for an adapted or adapted model that can be used to improve academic leaderboard standings .
Outcome: The proposed model can do well on professional and linguistic examinations, college-level knowledge questions, and collections of reasoning tasks.
On the Robustness of Cognate Generation Models (2022.lrec-1)

Copied to clipboard

Challenge: We examine different types of noise generated by human errors and how these noisy inputs affect the performance of cognate generation models.
Approach: They evaluate two popular neural cognate generation models’ robustness to human-plausible noise.
Outcome: The proposed models are robust to deletion, duplication, swapping, keyboard errors, and a new type of error, phonological errors.
Efficient Citer: Tuning Large Language Models for Enhanced Answer Quality and Verification (2024.findings-naacl)

Copied to clipboard

Challenge: Existing models with explicit citations lack the ability to verify information generated by these models.
Approach: They construct a citation training dataset and fine-tune two models to address the challenge of explicit citations efficiently.
Outcome: The proposed models surpass ChatGPT and exhibit exceptional out-of-domain generalization in both human and automatic evaluation.
MobileLLM-Flash: Latency-Guided On-Device LLM Design for Industry Scale Deployment (2026.acl-industry)

Copied to clipboard

Challenge: MobileLLM-Flash is a family of foundation models for efficient on-device use with strong capabilities.
Approach: They propose a method for designing on-device large language models under mobile latency constraints using hardware-in-the-loop architecture search.
Outcome: The proposed model is amenable to industry-scale deployment and is compatible with mobile runtimes like Executorch.
Modeling Color Terminology Across Thousands of Languages (D19-1)

Copied to clipboard

Challenge: Existing studies on what constitutes a "basic" color term and its acquisition sequence are flawed . a pan-lingual approach may reveal general color trends more reliably than smaller datasets.
Approach: They propose to operationalize and critique the Berlin and Kay color term hypotheses . they use 14 empirically-grounded computational linguistic metrics to analyze cross-linguistic data .
Outcome: The proposed measures correlate strongly with the Berlin and Kay color term partition and their hypothesized universal acquisition sequence.
An Analysis of Massively Multilingual Neural Machine Translation for Low-Resource Languages (2020.lrec-1)

Copied to clipboard

Challenge: In this study, we explore massively multilingual low-resource neural machine translation.
Approach: They propose to use Bible translations to train models with up to 1,107 source languages and create multilingual corpora varying the number and relatedness of source languages.
Outcome: The proposed approach is highly language-specific and can be tailored to the source language and its typology.
Walk in Others’ Shoes with a Single Glance: Human-Centric Visual Grounding with Top-View Perspective Transformation (2025.acl-long)

Copied to clipboard

Challenge: Existing VLMs are insensitive to information differences induced by slight perspective changes.
Approach: They propose a visual perspective-taking task that requires robots to interpret human-centric instructions and identify corresponding objects from robot perspectives.
Outcome: The proposed method improves performance by up to 18% and generalizes effectively to robotic and dynamic scenarios.
Massively Translingual Compound Analysis and Translation Discovery (L18-1)

Copied to clipboard

Challenge: Morphological compounding is one of the most common and productive methods of word formation across the world's languages.
Approach: They propose a model for compounding using bilingual dictionaries and no annotated training data . they also release a massively multilingual dataset of compound words and their decompositions .
Outcome: The proposed model generates novel translations of English concepts on a multilingual dataset . the model can be applied to a wide range of languages and is highly reproducible.
MixCE: Training Autoregressive Language Models by Mixing Forward and Reverse Cross-Entropies (2023.acl-long)

Copied to clipboard

Challenge: Autoregressive language models are trained by minimizing the cross-entropy of the model distribution Q relative to the data distribution P. However, these systems still struggle in many openended generation settings, where they are asked to produce a long text following a short prompt.
Approach: They propose to combine forward and reverse cross-entropy to train autoregressive language models by minimizing the cross-Entropy of the model distribution Q relative to the data distribution P.
Outcome: The proposed model overgeneralizes and produces non-human-like text without complex decoding strategies.
Fine-grained Morphosyntactic Analysis and Generation Tools for More Than One Thousand Languages (2020.lrec-1)

Copied to clipboard

Challenge: Using morphosyntactic tools, we train and distribute tools for approximately one thousand languages.
Approach: They train and distribute morphosyntactic tools for approximately one thousand languages.
Outcome: The results show that the tools generalize well across rare and common forms alike.
Scaling Collaborative Effort with Agents (2026.findings-acl)

Copied to clipboard

Challenge: Current evaluations of agents focus on producing high-quality, final outputs in one shot, failing to account for the inherently iterative nature of many real-world problems.
Approach: They propose a framework that captures how an agent’s utility grows with increasing user involvement.
Outcome: The proposed framework captures how an agent’s utility grows with increasing user involvement, revealing a missing ingredient in agent design: the ability to sustain engagement and scaffold user understanding.
Creating a Translation Matrix of the Bible’s Names Across 591 Languages (L18-1)

Copied to clipboard

Challenge: In low-resource languages, the Bible is the only significant bilingual, or even monolingual, text available . standard word alignment tools can be noisy, making downstream tasks difficult . a novel resource of 1129 aligned Bible person and place names is developed .
Approach: They propose to use Bible person and place names as a tool for translation and transliteration . they use weighted edit distance, machine translation-based transliterations and affixal induction and transformation models to improve the Bible's output.
Outcome: The proposed model outperforms a widely used word aligner on 97% of test words on multilingual named-entity alignment and translation across 591 languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations