Papers by Chris Quirk
Multilingual Whispers: Generating Paraphrases with Translation (D19-55)
Copied to clipboard
| Challenge: | Humans naturally paraphrase, but they can generate approximately the same meaning with a different surface realization. |
| Approach: | They compare translation-based paraphrase gathering using human, automatic, or hybrid techniques to monolingual paraphrasing by experts and non-experts. |
| Outcome: | The proposed methods outperform human translation systems in a variety of translation tasks. |
Microsoft Icecaps: An Open-Source Toolkit for Conversation Modeling (P19-3)
Copied to clipboard
Vighnesh Leonardo Shiv, Chris Quirk, Anshuman Suri, Xiang Gao, Khuram Shahid, Nithya Govindarajan, Yizhe Zhang, Jianfeng Gao, Michel Galley, Chris Brockett, Tulasi Menon, Bill Dolan
| Challenge: | upcoming open-source natural language processing repository aims to train conversational agents for multi-turn situations. |
| Approach: | They present the Intelligent Conversation Engine: Code and Pre-trained Systems (ICECAPS) the framework wraps TensorFlow functionality in a modular component-based architecture. |
| Outcome: | The Intelligent Conversation Engine: Code and Pre-trained Systems (ICECAPS) is an open-source natural language processing repository. |
When does text prediction benefit from additional context? An exploration of contextual signals for chat and email messages (2021.naacl-industry)
Copied to clipboard
| Challenge: | Prior-message context provides the greatest lift in Teams (chat) scenario. |
| Approach: | They compare prior-message context with email and chat messages from Microsoft Teams and Outlook. |
| Outcome: | The proposed model outperforms existing models on two of the largest commercial communication platforms: Microsoft Teams and Outlook. |
Probing Factually Grounded Content Transfer with Factual Ablation (2022.findings-acl)
Copied to clipboard
| Challenge: | Despite recent success, large neural models often generate factually incorrect text . lack of a standard evaluation for factuality complicates factual grounded generation . |
| Approach: | They propose a method to measure factual consistency by presenting two evaluation sets . large pretrained models have shown impressive effectiveness at longstanding tasks . |
| Outcome: | The proposed method improves over strong baselines by presenting two evaluation sets. |
Towards Content Transfer through Grounded Text Generation (N19-1)
Copied to clipboard
| Challenge: | Recent work in neural natural language generation has attracted significant interest in controlling the form of text, such as style, persona, and wordiness. |
| Approach: | They propose a task where the task is to generate a next sentence in a document that fits its context and is grounded in . external textual source such as a news story. |
| Outcome: | The proposed task is based on 640k Wikipedia referenced sentences paired with the source articles to show significant improvements against baselines. |
SimulatorArena: Are User Simulators Reliable Proxies for Multi-Turn Evaluation of AI Assistants? (2025.emnlp-main)
Copied to clipboard
Yao Dou, Michel Galley, Baolin Peng, Chris Kedzie, Weixin Cai, Alan Ritter, Chris Quirk, Wei Xu, Jianfeng Gao
| Challenge: | Large language models (LLMs) are increasingly used in interactive applications, and human evaluation remains the gold standard for assessing their performance in multi-turn conversations. |
| Approach: | They propose to use large language models to simulate users for automatic assistant evaluation. |
| Outcome: | The proposed model outperforms human evaluations on two interactive tasks and achieves Spearman’s of 0.7 on both tasks. |
Confidence Modeling for Neural Semantic Parsing (P18-1)
Copied to clipboard
| Challenge: | Experimental results show that neural semantic parsers are difficult to interpret due to their complexity. |
| Approach: | They propose to use confidence models to estimate predictions for neural semantic parsers . they outline three major causes of uncertainty and use metrics to quantify them . |
| Outcome: | The proposed model outperforms a widely used method that relies on posterior probability and improves interpretation quality. |
Text Editing by Command (2021.naacl-main)
Copied to clipboard
| Challenge: | Recent work has focused on making such models more controllable and factually grounded. |
| Approach: | They propose a novel interactive text generation setting in which the user interacts with the system by issuing commands to edit existing text. |
| Outcome: | The proposed model outperforms baseline models and obtains positive results in automatic and human evaluations. |