Papers by David M. Chan

5 papers
Do What? Teaching Vision-Language-Action Models to Reject the Impossible (2025.findings-emnlp)

Copied to clipboard

Challenge: Recent studies show that VLAs can recognize, interpret, and respond to false-premise instructions.
Approach: They propose a framework that detects when an instruction cannot be executed due to a false premise and engages in language-based clarification or correction.
Outcome: The proposed framework detects when an instruction cannot be executed due to a false premise and engages in language-based clarification or correction.
Distribution Aware Metrics for Conditional Natural Language Generation (2024.lrec-main)

Copied to clipboard

Challenge: Existing metrics for conditional natural language generation rely on pairwise comparisons between a single generated text and the best-matching reference.
Approach: They propose a family of meta-metrics that build on existing pairwise distance functions to evaluate conditional natural language generation models.
Outcome: The proposed method evaluates the ability of a model to generate text matching diversity in references in visual description and summarization.
Multi-Stage Multi-Modal Pre-Training for Automatic Speech Recognition (2024.lrec-main)

Copied to clipboard

Challenge: Existing methods for pre-training for automatic speech recognition (ASR) focus on single-stage pre-train followed by fine-tuning on downstream task.
Approach: They propose a multi-modal pre-training method that combines unsupervised pre-training with translation-based supervised mid-training.
Outcome: The proposed method improves WERs by 38.45% over baselines on both Librispeech and SUPERB.
Puzzled by Puzzles: When Vision-Language Models Can’t Take a Hint (2025.emnlp-main)

Copied to clipboard

Challenge: rebus puzzles encode language through imagery, spatial arrangement, and symbolic substitution.
Approach: They construct a benchmark of rebus puzzles in english language to test their ability to interpret and solve them.
Outcome: The proposed model performs well on a set of english-language rebus puzzles.
Enough Coin Flips Can Make LLMs Act Bayesian (2025.acl-long)

Copied to clipboard

Challenge: Large language models exhibit the ability to generalize given few-shot examples in their input prompt, an emergent capability known as in-context learning.
Approach: They investigate whether large language models use in-context learning to generalize given few-shot examples in their input prompt.
Outcome: The proposed model can generalize given few-shot examples in their input prompt, an emergent capability known as in-context learning.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations