Papers by David Tan

10 papers
CHIME: LLM-Assisted Hierarchical Organization of Scientific Studies for Literature Review Support (2024.findings-acl)

Copied to clipboard

Challenge: Literature review requires researchers to synthesize a large amount of information.
Approach: They propose to use LLMs to generate hierarchical organizations from a set of studies . they use a human-in-the-loop process to correct errors in LLM-generated hierarchies .
Outcome: The proposed model improves assignment of studies to categories by 12.6 F1 points.
How Far can 100 Samples Go? Unlocking Zero-Shot Translation with Tiny Multi-Parallel Data (2024.findings-acl)

Copied to clipboard

Challenge: a common solution to zero-shot translation is to add as many related translation directions as possible to the training corpus.
Approach: They show that a small amount of multi-parallel data can achieve significant zero-shot improvements . they say that the resulting non-English performance is close to the complete translation upper bound .
Outcome: The proposed model achieves +21.7 ChrF++ non-English translation improvements on EC30 dataset . the resulting non- English performance exceeds M2M100 by an average of 5.9 ChrF+ .
What Gets Echoed? Understanding the “Pointers” in Explanations of Persuasive Arguments (D19-1)

Copied to clipboard

Challenge: Explanations are central to everyday life, and are a topic of growing interest in the AI community.
Approach: They propose a word-level prediction task to investigate how explanations selectively reuse information from what is being explained.
Outcome: The proposed features have strong predictive power on the echoing of a word in an explanation, and enhance neural methods of generating explanations.
SParC: Cross-Domain Semantic Parsing in Context (P19-1)

Copied to clipboard

Challenge: Xu et al., 2017): a dataset for cross-domain semantic parsing in context with 4,298 question sequences.
Approach: They present a dataset for cross-domainSemanticParsing inContext that consists of 4,298 coherent question sequences.
Outcome: The proposed dataset demonstrates that it has greater semantic diversity and can be generalized to unseen domains due to its cross-domain nature and the unseened databases at test time.
High Quality Rather than High Model Probability: Minimum Bayes Risk Decoding with Neural Metrics (2022.tacl-1)

Copied to clipboard

Challenge: Neural machine translations are ranked below human translations in professional evaluations .
Approach: They apply minimum bayes risk decoding to optimize different metrics of translation quality . they show that model estimates and translation quality only vaguely correlate .
Outcome: The proposed method improves human translations with different models and metric.
Inferring Events from Time Series using Language Models (2026.acl-long)

Copied to clipboard

Challenge: Prior work on reasoning about time series in conjunction with natural language has largely overlooked event descriptions and focused on tasks involving just numeric data like trend analysis or anomaly detection.
Approach: They propose a method for generating tasks that test a model’s ability to reason about events associated with time series data based on sports data and develop a benchmarking method.
Outcome: The proposed method can infer unobserved events from time series data, even when providing minimal context.
When Flores Bloomz Wrong: Cross-Direction Contamination in Machine Translation Evaluation (2026.eacl-short)

Copied to clipboard

Challenge: Large language models (LLMs) can be benchmark-contaminated, resulting in inflated scores that mask memorization as generalization.
Approach: They use the FLORES-200 translation benchmark as a diagnostic to investigate cross-direction data contamination.
Outcome: The proposed model can be cross-directional, boosting performance in unseen translation directions due to target-side memorization.
Experts, Errors, and Context: A Large-Scale Study of Human Evaluation for Machine Translation (2021.tacl-1)

Copied to clipboard

Challenge: a large study of machine translation systems shows poor evaluation procedures can lead to erroneous conclusions.
Approach: They propose an evaluation methodology grounded in explicit error analysis based on the Multidimensional Quality Metrics framework.
Outcome: The proposed evaluation methodology outperforms crowd workers in two languages . it shows that human-based metrics outperformed crowd workers .
Adaptive Data Flywheel: Applying MAPE Control Loops to AI Agent Improvement (2026.eacl-industry)

Copied to clipboard

Challenge: NVInfo AI is a generative AI agent that can be deployed in production without full-scale retraining or infrastructure overhauls.
Approach: They propose to implement a retrieval-augmented generation (RAG)-driven data flywheel in NVInfo AI, a mixture-of-experts knowledge assistant, for 30,000 employees.
Outcome: The proposed system addresses failures in retrieval-augmented generation pipelines and enables continuous learning.
ROBBIE: Robust Bias Evaluation of Large Generative Language Models (2023.emnlp-main)

Copied to clipboard

Challenge: generative large language models (LLMs) are becoming more performant and prevalent . we need tools to measure and improve their fairness, authors say .
Approach: They propose to compare 6 different prompt-based bias and toxicity metrics across 12 demographic axes and 5 families of generative large language models.
Outcome: The proposed model can be tested on more datasets to better characterize and mitigate biases . the study compared 6 prompt-based bias and toxicity metrics across 12 demographic axes and 5 families of generative large language models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations