Papers by Stefan Lee

9 papers
An Evaluation Dataset for Intent Classification and Out-of-Scope Prediction (D19-1)

Copied to clipboard

Challenge: Task-oriented dialog systems need to know when a query falls outside their range of supported intents.
Approach: They propose a dataset that includes queries that are out-of-scope and 150 intent classes over 10 domains.
Outcome: The proposed dataset includes queries that are out-of-scope, i.e., queries that do not fall into any of the system’s supported intents.
On the Sub-layer Functionalities of Transformer Decoder (2020.findings-emnlp)

Copied to clipboard

Challenge: Existing efforts to interpret the encoder of Transformer-based encoder-decoder architectures for neural machine translation have focused on assessing the encoded representations or interpreting the multi-head self-attentions.
Approach: They propose to use Transformer-based encoder-decoder architectures to analyze how information is propagated through each module of each decoder layer.
Outcome: The proposed model can be dropped with minimal loss of performance on three translation datasets and can be used to train and inference faster.
Outlier Detection for Improved Data Quality and Diversity in Dialog Systems (N19-1)

Copied to clipboard

Challenge: Existing methods to detect outliers in text have been neglected in NLP . outlier detection is a problem in dialog systems where text is often no more than a few sentences in length.
Approach: They propose a technique that uses sentence embeddings to detect outliers in short texts using neural sentence embeds and distance-based outlier detection.
Outcome: The proposed technique detects outliers in a corpus of short texts while generating highly diverse corpora that produce more robust intent classification and slot-filling models.
Sunny and Dark Outside?! Improving Answer Consistency in VQA through Entailed Question Generation (D19-1)

Copied to clipboard

Challenge: interacting with a model for Visual Question Answering (VQA) quickly reveals that these models lack consistency.
Approach: They propose a dataset, ConVQA, and metrics that enable quantitative evaluation of consistency in VQA.
Outcome: The proposed data augmentation module improves the consistency of VQA models on the Con-VQA dataset and is a strong baseline for future research.
Improving Multilingual Translation by Representation and Gradient Regularization (2021.emnlp-main)

Copied to clipboard

Challenge: Multilingual Neural Machine Translation models often produce low quality translations, often failing to produce outputs in the right target language.
Approach: They propose a joint approach to regularize NMT models at both representation-level and gradient-level to reduce off-target translation occurrences and improve zero-shot translation performance.
Outcome: The proposed approach reduces off-target translation occurrences and improves zero-shot translation performance by +5.59 and +10.38 BLEU on WMT and OPUS datasets.
Enhancing Zero-Shot Chain-of-Thought Reasoning in Large Language Models through Logic (2024.lrec-main)

Copied to clipboard

Challenge: Experimental evaluations of large language models demonstrate the efficacy of enhanced reasoning by logic.
Approach: They propose a framework that uses symbolic logic to verify and rectify reasoning steps by steps.
Outcome: The proposed framework improves the zero-shot chain-of-thought reasoning ability of large language models by verifying and rectifying the reasoning steps step by step.
Where Are You? Localization from Embodied Dialog (2020.emnlp-main)

Copied to clipboard

Challenge: Observer and Locator perform a cooperative localization task in a 3D environment.
Approach: They propose a dataset of 6k dialogs in which two humans complete a cooperative localization task.
Outcome: The proposed model achieves 32.7% success at identifying the Observer’s location within 3m in unseen buildings, vs. 70.4% for human Locators.
Knowing the Facts but Choosing the Shortcut: Understanding How Large Language Models Compare Entities (2026.eacl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are increasingly used for knowledge-based reasoning tasks, yet understanding when they rely on genuine knowledge versus superficial heuristics remains challenging.
Approach: They propose to ask LLMs to compare numerical attributes to find out which country has the highest population, France or Germany.
Outcome: The proposed model comparisons show that heuristics override principled reasoning for larger models, while smaller models show no discrimination.
Language-Informed Beam Search Decoding for Multilingual Machine Translation (2024.findings-acl)

Copied to clipboard

Challenge: Beam search decoding is the de-facto method for decoding auto-regressive Neural Machine Translation (NMT) models, but decoding multilingual NMT models produces off-target translations .
Approach: They propose a general decoding algorithm incorporating an off-the-shelf Language Identification (LiD) model into beam search decoding to reduce off-target translations.
Outcome: The proposed language-informed beam search improves +1.1 BLEU and +0.9 BLUE on WMT and OPUS datasets and reduces off-target rates from 22.9% to 7.7% and 65.8% to 25.3% respectively.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations