Papers by Sewon Min
Multi-hop Reading Comprehension through Question Decomposition and Rescoring (P19-1)
Copied to clipboard
| Challenge: | Existing systems for multi-hop reading comprehension decompose compositional questions into simpler sub-questions . authors propose a system that learns to break compositional multi- hop questions into simple singlehop sub-question . |
| Approach: | They propose a system that decomposes a compositional question into simpler sub-questions . they propose recast subquestion generation as a span prediction problem . |
| Outcome: | The proposed system generates as effective as human-authored sub-questions using 400 examples . it also provides explainable evidence for its decision making in the form of sub-questions . |
FaVIQ: FAct Verification from Information-seeking Questions (2022.acl-long)
Copied to clipboard
| Challenge: | Existing fact verification datasets with crowdsourced claims introduce subtle biases that are difficult to control for. |
| Approach: | They construct a large-scale fact verification dataset with ambiguous questions . they use a corpus of 188k claims to construct false and true claims . |
| Outcome: | The proposed dataset outperforms models trained on the dataset FEVER or in-domain data by up to 17% absolute. |
Efficient and Robust Question Answering from Minimal Context over Documents (P18-1)
Copied to clipboard
| Challenge: | Recent work shows that neural QA models are sensitive to adversarial inputs. |
| Approach: | They propose a sentence selector to select the minimal set of sentences to feed into a QA model. |
| Outcome: | The proposed system reduces training time and inference time by up to 13 times . it is comparable to or better than the state-of-the-art on SQuAD, NewsQA, TriviaQA and SQu AD-Open . |
Measuring and Narrowing the Compositionality Gap in Language Models (2023.findings-emnlp)
Copied to clipboard
| Challenge: | a language model can correctly answer all sub-problems but not generate the overall solution. |
| Approach: | They propose a method that asks itself and then answers follow-up questions to narrow the compositionality gap by reasoning explicitly instead of implicitly. |
| Outcome: | The proposed method improves on chain of thought by asking itself and answering follow-up questions. |
InSCIt: Information-Seeking Conversations with Mixed-Initiative Interactions (2023.tacl-1)
Copied to clipboard
Zeqiu Wu, Ryu Parish, Hao Cheng, Sewon Min, Prithviraj Ammanabrolu, Mari Ostendorf, Hannaneh Hajishirzi
| Challenge: | In information-seeking conversations, a user may ask questions that are under-specified or unanswerable. |
| Approach: | They present a dataset for information-seeking conversations with mixed-initiative interactions . they use Wikipedia to search for answers and provide relevant information . |
| Outcome: | The proposed system significantly underperforms humans in two of the most recent studies. |
Efficient One-Pass End-to-End Entity Linking for Questions (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing models for entity linking are limited to entity disambiguation and require mention boundaries to be given in the input. |
| Approach: | They propose a fast end-to-end entity linking model that uses a biencoder to jointly detect mentions and link in one pass. |
| Outcome: | The proposed model outperforms the current state of the art on WebQSP and GraphQuestions with extended annotations that cover multiple entities per question. |
UNIFIEDQA: Crossing Format Boundaries with a Single QA System (2020.findings-emnlp)
Copied to clipboard
Daniel Khashabi, Sewon Min, Tushar Khot, Ashish Sabharwal, Oyvind Tafjord, Peter Clark, Hannaneh Hajishirzi
| Challenge: | Question answering (QA) tasks have been posed using a variety of formats . a new study aims to develop specialized QA models that can be used to train QA systems . |
| Approach: | They build a pre-trained question answering model that performs well across 19 QA datasets . they argue that format-specialized models can limit the ability to teach reasoning . |
| Outcome: | a new model that trains on QA datasets performs on par with 8 models trained on individual datasets . a single model that trained on UNIFIEDQA performs well on 19 QA data . |
CREPE: Open-Domain Question Answering with False Presuppositions (2023.acl-long)
Copied to clipboard
| Challenge: | Existing question answering datasets assume all questions have well defined answers. |
| Approach: | They propose a QA dataset containing a distribution of false presuppositions . they find that 25% of questions contain false presumptions . |
| Outcome: | The proposed model finds that 25% of questions contain false presuppositions . the model can find presuffpositions moderately well, but struggle when predicting correctness . |
OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens (2025.acl-demo)
Copied to clipboard
Jiacheng Liu, Taylor Blanton, Yanai Elazar, Sewon Min, Yen-Sung Chen, Arnavi Chheda-Kothary, Huy Tran, Byron Bischoff, Eric Marsh, Michael Schmitz, Cassidy Trier, Aaron Sarnat, Jenna James, Jon Borchardt, Bailey Kuehl, Evie Yu-Yen Cheng, Karen Farley, Taira Anderson, David Albright, Carissa Schoenick, Luca Soldaini, Dirk Groeneveld, Rock Yuren Pang, Pang Wei Koh, Noah A. Smith, Sophie Lebrecht, Yejin Choi, Hannaneh Hajishirzi, Ali Farhadi, Jesse Dodge
| Challenge: | tracing language models' outputs back to training data is a problem because they are trained on text corpora with trillions of tokens . existing methods for tracers have not been scaled to work within this multi-trillion-token setting . |
| Approach: | They propose a system that traces language models' outputs verbatim back to training data . OLMOTRACE retrieves documents from the model's training data that contain exact matches . |
| Outcome: | The proposed system can find verbatim matches between LM output and training data . it can be used to explore fact checking, hallucination, and creativity of language models . |
Nonparametric Masked Language Modeling (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing language models (LMs) predict tokens with a softmax over a finite vocabulary, which can make it difficult to predict rare tokens or phrases. |
| Approach: | They introduce a nonparametric masked language model that replaces a softmax with a distribution over every phrase in a reference corpus and uses an in-batch approximation to train it. |
| Outcome: | The proposed model outperforms larger parametric models on 16 tasks including classification, fact probing and question answering. |
Towards Understanding Chain-of-Thought Prompting: An Empirical Study of What Matters (2023.acl-long)
Copied to clipboard
| Challenge: | Chain-of-Thought (CoT) prompting can dramatically improve the multi-step reasoning abilities of large language models (LLMs). |
| Approach: | They propose to use Chain-of-Thought (CoT) prompting to encourage the LLM to generate intermediate rationales for solving a problem by providing a series of reasoning steps in the demonstrations. |
| Outcome: | The proposed model can generate coherent lines of reasoning even with invalid demonstrations while still generating coherent lines during inference. |
MetaICL: Learning to Learn In Context (2022.naacl-main)
Copied to clipboard
| Challenge: | Large language models can do in-context learning by conditioning on a few training examples with no parameter updates or task-specific templates. |
| Approach: | They propose a meta-training framework where a pretrained language model is tuned to do in-context learning on a large set of training tasks. |
| Outcome: | The proposed framework outperforms baseline models on 142 NLP datasets and a range of target tasks with domain shifts. |
Z-ICL: Zero-Shot In-Context Learning with Pseudo-Demonstrations (2023.acl-long)
Copied to clipboard
| Challenge: | Existing methods for zero-shot learning are based on in-context training, but performance drops when no demonstrations are available. |
| Approach: | They propose a new method that constructs pseudo-demonstrations for a given test input using a raw text corpus and applies techniques to reduce copying. |
| Outcome: | The proposed method outperforms previous zero-shot methods on nine classification datasets and is on par with in-context learning with labeled training data in the few-shot setting. |
RECONSIDER: Improved Re-Ranking using Span-Focused Cross-Attention for Open Domain Question Answering (2021.naacl-main)
Copied to clipboard
| Challenge: | State-of-the-art Machine Reading Comprehension (MRC) models for Open-domain Question Answering (QA) achieve high recall amongst top few predictions, but low overall accuracy, motivating the need for answer re-ranking. |
| Approach: | They propose a method to make answer re-ranking successful for span-extraction tasks even beyond large pre-training. |
| Outcome: | The proposed approach achieves 45.5% Exact Match accuracy on Natural Questions and 61.7% on TriviaQA. |
Exploring The Landscape of Distributional Robustness for Question Answering Models (2022.findings-emnlp)
Copied to clipboard
Anas Awadalla, Mitchell Wortsman, Gabriel Ilharco, Sewon Min, Ian Magnusson, Hannaneh Hajishirzi, Ludwig Schmidt
| Challenge: | Existing methods for predicting distributional robustness fail to generalize reliably in a variety of test conditions. |
| Approach: | They conduct a large empirical evaluation to investigate the landscape of distributional robustness in question answering. |
| Outcome: | The proposed methods are more robust to distribution shifts than fully fine-tuned models, and few-shot prompt models exhibit better robustness than few- shot prompt models. |
Compositional Questions Do Not Necessitate Multi-hop Reasoning (P19-1)
Copied to clipboard
| Challenge: | a single-hop reasoning model can solve much more of the dataset than previously thought. |
| Approach: | They propose a single-hop BERT-based RC model that achieves 67 F1 . they propose an evaluation setting where humans are not shown all paragraphs . |
| Outcome: | The proposed model achieves 67 F1—comparable to state-of-the-art multi-hop models. |
Prompt Waywardness: The Curious Case of Discretized Interpretation of Continuous Prompts (2022.naacl-main)
Copied to clipboard
Daniel Khashabi, Xinxi Lyu, Sewon Min, Lianhui Qin, Kyle Richardson, Sean Welleck, Hannaneh Hajishirzi, Tushar Khot, Ashish Sabharwal, Sameer Singh, Yejin Choi
| Challenge: | Recent work shows the surprising power of continuous prompts to language models for controlled generation and solving a wide range of tasks. |
| Approach: | They propose to extract a discrete (textual) interpretation of continuous prompts faithful to the problem they solve. |
| Outcome: | The proposed model can find prompts that solve a task while being projected to an arbitrary text with a smaller drop in accuracy. |
Retrieval-based Language Models and Applications (2023.acl-tutorials)
Copied to clipboard
| Challenge: | In this tutorial, we will provide a comprehensive overview of retrieval-based language models. |
| Approach: | This tutorial will provide a comprehensive overview of recent advances in retrieval-based language models. |
| Outcome: | This tutorial will provide a comprehensive overview of recent advances in retrieval-based language models. |
Beyond Paragraphs: NLP for Long Sequences (2021.naacl-tutorials)
Copied to clipboard
| Challenge: | In this tutorial, we will introduce document-level representation learning techniques . document-based learning is challenging due to the limited sequence length of many models . |
| Approach: | They will provide an overview of established long sequence NLP techniques and discuss memory-saving methods that are key to processing long sequences. |
| Outcome: | The tutorial will introduce the latest and ongoing techniques for document-level representation learning. |
A Discrete Hard EM Approach for Weakly Supervised Question Answering (D19-1)
Copied to clipboard
| Challenge: | Existing work on question answering tasks only provide weak supervision for how the answer should be computed . weak supervision is attractive because it is relatively easy to gather, allowing for large datasets . but weak supervision complicates learning because there are many different spurious ways to derive the correct answer. |
| Approach: | They propose a method to convert question answering tasks into discrete latent variable learning problems with a precomputed set of possible solutions that contains one correct option. |
| Outcome: | The proposed approach outperforms previous methods on six QA tasks and achieves state-of-the-art on five of them. |
REPLUG: Retrieval-Augmented Black-Box Language Models (2024.naacl-long)
Copied to clipboard
Weijia Shi, Sewon Min, Michihiro Yasunaga, Minjoon Seo, Richard James, Mike Lewis, Luke Zettlemoyer, Wen-tau Yih
| Challenge: | Existing retrieval-augmented language models require access to internal representations to enhance performance. |
| Approach: | They introduce a retrieval-augmented language modeling framework that treats the language model as a black box and augments it with a tuneable retrieval model. |
| Outcome: | The proposed framework improves performance on language modeling tasks by 6.3% and 5.1%. |
FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation (2023.emnlp-main)
Copied to clipboard
Sewon Min, Kalpesh Krishna, Xinxi Lyu, Mike Lewis, Wen-tau Yih, Pang Koh, Mohit Iyyer, Luke Zettlemoyer, Hannaneh Hajishirzi
| Challenge: | Evaluating the factuality of long-form text generated by large language models (LMs) is non-trivial because (1) generations often contain a mixture of supported and unsupported pieces of information, making binary judgments of quality inadequate and (2) human evaluation is time-consuming and costly. |
| Approach: | They introduce a new evaluation that breaks a generation into a series of atomic facts and computes the percentage of atom facts supported by a reliable knowledge source. |
| Outcome: | The proposed model breaks a generation into atomic facts and computes the percentage of atomic fact supported by a reliable knowledge source. |
CopyBench: Measuring Literal and Non-Literal Reproduction of Copyright-Protected Text in Language Model Generation (2024.emnlp-main)
Copied to clipboard
Tong Chen, Akari Asai, Niloofar Mireshghallah, Sewon Min, James Grimmelmann, Yejin Choi, Hannaneh Hajishirzi, Luke Zettlemoyer, Pang Wei Koh
| Challenge: | Existing studies focus on literal copying, but current methods reduce literal copy but not non-literal copying. |
| Approach: | They propose a benchmark to measure literal and non-literal copying in LMs . they use copyrighted fiction books as text sources to assess literal copying . |
| Outcome: | The proposed model measures literal and non-literal copying in copyrighted texts . large models show significantly more copying, with literal copying rates increasing . |
Dense Passage Retrieval for Open-Domain Question Answering (2020.emnlp-main)
Copied to clipboard
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, Wen-tau Yih
| Challenge: | Open-domain question answering relies on efficient passage retrieval to select candidate contexts. |
| Approach: | They propose a dual-encoder framework that can be implemented to retrieve passages from a small number of questions and passages. |
| Outcome: | The proposed system outperforms a strong Lucene-BM25 system in top-20 passage retrieval accuracy on multiple open-domain QA benchmarks. |
On Making Reading Comprehension More Comprehensive (D19-58)
Copied to clipboard
| Challenge: | Getting machines to "understand" text is a vast and long-standing problem, made more challenging by the fact that it is not even clear what it means to understand text. |
| Approach: | They propose a question-based approach to machine reading comprehension that uses a natural language question to test a system's comprehension of a passage of text. |
| Outcome: | The proposed questions have surface cues or other biases that allow a model to shortcut the intended reasoning process. |
Re-Examining Calibration: The Case of Question Answering (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Existing calibration methods do not provide significant gains in accuracy. |
| Approach: | They propose a new calibration metric that better captures whether the model assigns low confidence to wrong predictions and high confidence to correct predictions. |
| Outcome: | The proposed calibration method better captures whether the model assigns low confidence to wrong predictions and high confidence to correct predictions. |
Zero- and Few-Shot NLP with Pretrained Language Models (2022.acl-tutorials)
Copied to clipboard
| Challenge: | a tutorial aims to introduce NLP researchers to the latest techniques for learning from little-to-no data . aims at bringing interested researchers up to speed about the latest and ongoing techniques . |
| Approach: | They aim to introduce techniques for learning from little-to-no data using pretrained language models. |
| Outcome: | This tutorial aims to bring interested NLP researchers up to speed about recent techniques . it will cover methods from manual engineering, better inference algorithms to better tuning methods . |
Noisy Channel Language Model Prompting for Few-Shot Text Classification (2022.acl-long)
Copied to clipboard
| Challenge: | Prior work has suggested methods for finding better prompt or scoring of the output from the model. |
| Approach: | They propose a noisy channel approach for language model prompting in few-shot text classification by in-context demonstration or prompt tuning. |
| Outcome: | The proposed model outperforms direct models in both demonstration and prompt tuning. |
AmbigQA: Answering Ambiguous Open-domain Questions (2020.emnlp-main)
Copied to clipboard
| Challenge: | Existing open-domain question answering systems assume questions have a single welldefined answer. |
| Approach: | They propose an open-domain question answering task which involves finding every plausible answer and rewriting the question for each one to resolve the ambiguity. |
| Outcome: | The proposed task is based on a dataset covering 14,042 open-domain questions . it shows that strong models benefit from weakly supervised learning . |
Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? (2022.emnlp-main)
Copied to clipboard
Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe, Mike Lewis, Hannaneh Hajishirzi, Luke Zettlemoyer
| Challenge: | Large language models can in-context learn by conditioning on a few input-label pairs and making predictions for new inputs. |
| Approach: | They propose to use ground truth demonstrations to replace labels in demonstrations . they also show that other aspects of the demonstrations are key drivers of endtask performance . |
| Outcome: | The proposed model outperforms zeroshot inference on a wide range of tasks using ground truth demonstrations. |
Joint Passage Ranking for Diverse Multi-Answer Retrieval (2021.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to multi-answer retrieval cannot reason about the set of passages jointly. |
| Approach: | They propose a joint passage retrieval model focusing on reranking to solve multi-answer retrieval problem. |
| Outcome: | The proposed model outperforms baseline models on three multi-answer datasets. |