Papers by Sewon Min

31 papers
Multi-hop Reading Comprehension through Question Decomposition and Rescoring (P19-1)

Copied to clipboard

Challenge: Existing systems for multi-hop reading comprehension decompose compositional questions into simpler sub-questions . authors propose a system that learns to break compositional multi- hop questions into simple singlehop sub-question .
Approach: They propose a system that decomposes a compositional question into simpler sub-questions . they propose recast subquestion generation as a span prediction problem .
Outcome: The proposed system generates as effective as human-authored sub-questions using 400 examples . it also provides explainable evidence for its decision making in the form of sub-questions .
FaVIQ: FAct Verification from Information-seeking Questions (2022.acl-long)

Copied to clipboard

Challenge: Existing fact verification datasets with crowdsourced claims introduce subtle biases that are difficult to control for.
Approach: They construct a large-scale fact verification dataset with ambiguous questions . they use a corpus of 188k claims to construct false and true claims .
Outcome: The proposed dataset outperforms models trained on the dataset FEVER or in-domain data by up to 17% absolute.
Efficient and Robust Question Answering from Minimal Context over Documents (P18-1)

Copied to clipboard

Challenge: Recent work shows that neural QA models are sensitive to adversarial inputs.
Approach: They propose a sentence selector to select the minimal set of sentences to feed into a QA model.
Outcome: The proposed system reduces training time and inference time by up to 13 times . it is comparable to or better than the state-of-the-art on SQuAD, NewsQA, TriviaQA and SQu AD-Open .
Measuring and Narrowing the Compositionality Gap in Language Models (2023.findings-emnlp)

Copied to clipboard

Challenge: a language model can correctly answer all sub-problems but not generate the overall solution.
Approach: They propose a method that asks itself and then answers follow-up questions to narrow the compositionality gap by reasoning explicitly instead of implicitly.
Outcome: The proposed method improves on chain of thought by asking itself and answering follow-up questions.
InSCIt: Information-Seeking Conversations with Mixed-Initiative Interactions (2023.tacl-1)

Copied to clipboard

Challenge: In information-seeking conversations, a user may ask questions that are under-specified or unanswerable.
Approach: They present a dataset for information-seeking conversations with mixed-initiative interactions . they use Wikipedia to search for answers and provide relevant information .
Outcome: The proposed system significantly underperforms humans in two of the most recent studies.
Efficient One-Pass End-to-End Entity Linking for Questions (2020.emnlp-main)

Copied to clipboard

Challenge: Existing models for entity linking are limited to entity disambiguation and require mention boundaries to be given in the input.
Approach: They propose a fast end-to-end entity linking model that uses a biencoder to jointly detect mentions and link in one pass.
Outcome: The proposed model outperforms the current state of the art on WebQSP and GraphQuestions with extended annotations that cover multiple entities per question.
UNIFIEDQA: Crossing Format Boundaries with a Single QA System (2020.findings-emnlp)

Copied to clipboard

Challenge: Question answering (QA) tasks have been posed using a variety of formats . a new study aims to develop specialized QA models that can be used to train QA systems .
Approach: They build a pre-trained question answering model that performs well across 19 QA datasets . they argue that format-specialized models can limit the ability to teach reasoning .
Outcome: a new model that trains on QA datasets performs on par with 8 models trained on individual datasets . a single model that trained on UNIFIEDQA performs well on 19 QA data .
CREPE: Open-Domain Question Answering with False Presuppositions (2023.acl-long)

Copied to clipboard

Challenge: Existing question answering datasets assume all questions have well defined answers.
Approach: They propose a QA dataset containing a distribution of false presuppositions . they find that 25% of questions contain false presumptions .
Outcome: The proposed model finds that 25% of questions contain false presuppositions . the model can find presuffpositions moderately well, but struggle when predicting correctness .
OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens (2025.acl-demo)

Copied to clipboard

Challenge: tracing language models' outputs back to training data is a problem because they are trained on text corpora with trillions of tokens . existing methods for tracers have not been scaled to work within this multi-trillion-token setting .
Approach: They propose a system that traces language models' outputs verbatim back to training data . OLMOTRACE retrieves documents from the model's training data that contain exact matches .
Outcome: The proposed system can find verbatim matches between LM output and training data . it can be used to explore fact checking, hallucination, and creativity of language models .
Nonparametric Masked Language Modeling (2023.findings-acl)

Copied to clipboard

Challenge: Existing language models (LMs) predict tokens with a softmax over a finite vocabulary, which can make it difficult to predict rare tokens or phrases.
Approach: They introduce a nonparametric masked language model that replaces a softmax with a distribution over every phrase in a reference corpus and uses an in-batch approximation to train it.
Outcome: The proposed model outperforms larger parametric models on 16 tasks including classification, fact probing and question answering.
Towards Understanding Chain-of-Thought Prompting: An Empirical Study of What Matters (2023.acl-long)

Copied to clipboard

Challenge: Chain-of-Thought (CoT) prompting can dramatically improve the multi-step reasoning abilities of large language models (LLMs).
Approach: They propose to use Chain-of-Thought (CoT) prompting to encourage the LLM to generate intermediate rationales for solving a problem by providing a series of reasoning steps in the demonstrations.
Outcome: The proposed model can generate coherent lines of reasoning even with invalid demonstrations while still generating coherent lines during inference.
MetaICL: Learning to Learn In Context (2022.naacl-main)

Copied to clipboard

Challenge: Large language models can do in-context learning by conditioning on a few training examples with no parameter updates or task-specific templates.
Approach: They propose a meta-training framework where a pretrained language model is tuned to do in-context learning on a large set of training tasks.
Outcome: The proposed framework outperforms baseline models on 142 NLP datasets and a range of target tasks with domain shifts.
Z-ICL: Zero-Shot In-Context Learning with Pseudo-Demonstrations (2023.acl-long)

Copied to clipboard

Challenge: Existing methods for zero-shot learning are based on in-context training, but performance drops when no demonstrations are available.
Approach: They propose a new method that constructs pseudo-demonstrations for a given test input using a raw text corpus and applies techniques to reduce copying.
Outcome: The proposed method outperforms previous zero-shot methods on nine classification datasets and is on par with in-context learning with labeled training data in the few-shot setting.
RECONSIDER: Improved Re-Ranking using Span-Focused Cross-Attention for Open Domain Question Answering (2021.naacl-main)

Copied to clipboard

Challenge: State-of-the-art Machine Reading Comprehension (MRC) models for Open-domain Question Answering (QA) achieve high recall amongst top few predictions, but low overall accuracy, motivating the need for answer re-ranking.
Approach: They propose a method to make answer re-ranking successful for span-extraction tasks even beyond large pre-training.
Outcome: The proposed approach achieves 45.5% Exact Match accuracy on Natural Questions and 61.7% on TriviaQA.
Exploring The Landscape of Distributional Robustness for Question Answering Models (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for predicting distributional robustness fail to generalize reliably in a variety of test conditions.
Approach: They conduct a large empirical evaluation to investigate the landscape of distributional robustness in question answering.
Outcome: The proposed methods are more robust to distribution shifts than fully fine-tuned models, and few-shot prompt models exhibit better robustness than few- shot prompt models.
Compositional Questions Do Not Necessitate Multi-hop Reasoning (P19-1)

Copied to clipboard

Challenge: a single-hop reasoning model can solve much more of the dataset than previously thought.
Approach: They propose a single-hop BERT-based RC model that achieves 67 F1 . they propose an evaluation setting where humans are not shown all paragraphs .
Outcome: The proposed model achieves 67 F1—comparable to state-of-the-art multi-hop models.
Prompt Waywardness: The Curious Case of Discretized Interpretation of Continuous Prompts (2022.naacl-main)

Copied to clipboard

Challenge: Recent work shows the surprising power of continuous prompts to language models for controlled generation and solving a wide range of tasks.
Approach: They propose to extract a discrete (textual) interpretation of continuous prompts faithful to the problem they solve.
Outcome: The proposed model can find prompts that solve a task while being projected to an arbitrary text with a smaller drop in accuracy.
Retrieval-based Language Models and Applications (2023.acl-tutorials)

Copied to clipboard

Challenge: In this tutorial, we will provide a comprehensive overview of retrieval-based language models.
Approach: This tutorial will provide a comprehensive overview of recent advances in retrieval-based language models.
Outcome: This tutorial will provide a comprehensive overview of recent advances in retrieval-based language models.
Beyond Paragraphs: NLP for Long Sequences (2021.naacl-tutorials)

Copied to clipboard

Challenge: In this tutorial, we will introduce document-level representation learning techniques . document-based learning is challenging due to the limited sequence length of many models .
Approach: They will provide an overview of established long sequence NLP techniques and discuss memory-saving methods that are key to processing long sequences.
Outcome: The tutorial will introduce the latest and ongoing techniques for document-level representation learning.
A Discrete Hard EM Approach for Weakly Supervised Question Answering (D19-1)

Copied to clipboard

Challenge: Existing work on question answering tasks only provide weak supervision for how the answer should be computed . weak supervision is attractive because it is relatively easy to gather, allowing for large datasets . but weak supervision complicates learning because there are many different spurious ways to derive the correct answer.
Approach: They propose a method to convert question answering tasks into discrete latent variable learning problems with a precomputed set of possible solutions that contains one correct option.
Outcome: The proposed approach outperforms previous methods on six QA tasks and achieves state-of-the-art on five of them.
REPLUG: Retrieval-Augmented Black-Box Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Existing retrieval-augmented language models require access to internal representations to enhance performance.
Approach: They introduce a retrieval-augmented language modeling framework that treats the language model as a black box and augments it with a tuneable retrieval model.
Outcome: The proposed framework improves performance on language modeling tasks by 6.3% and 5.1%.
FActScore: Fine-grained Atomic Evaluation of Factual Precision in Long Form Text Generation (2023.emnlp-main)

Copied to clipboard

Challenge: Evaluating the factuality of long-form text generated by large language models (LMs) is non-trivial because (1) generations often contain a mixture of supported and unsupported pieces of information, making binary judgments of quality inadequate and (2) human evaluation is time-consuming and costly.
Approach: They introduce a new evaluation that breaks a generation into a series of atomic facts and computes the percentage of atom facts supported by a reliable knowledge source.
Outcome: The proposed model breaks a generation into atomic facts and computes the percentage of atomic fact supported by a reliable knowledge source.
CopyBench: Measuring Literal and Non-Literal Reproduction of Copyright-Protected Text in Language Model Generation (2024.emnlp-main)

Copied to clipboard

Challenge: Existing studies focus on literal copying, but current methods reduce literal copy but not non-literal copying.
Approach: They propose a benchmark to measure literal and non-literal copying in LMs . they use copyrighted fiction books as text sources to assess literal copying .
Outcome: The proposed model measures literal and non-literal copying in copyrighted texts . large models show significantly more copying, with literal copying rates increasing .
Dense Passage Retrieval for Open-Domain Question Answering (2020.emnlp-main)

Copied to clipboard

Challenge: Open-domain question answering relies on efficient passage retrieval to select candidate contexts.
Approach: They propose a dual-encoder framework that can be implemented to retrieve passages from a small number of questions and passages.
Outcome: The proposed system outperforms a strong Lucene-BM25 system in top-20 passage retrieval accuracy on multiple open-domain QA benchmarks.
On Making Reading Comprehension More Comprehensive (D19-58)

Copied to clipboard

Challenge: Getting machines to "understand" text is a vast and long-standing problem, made more challenging by the fact that it is not even clear what it means to understand text.
Approach: They propose a question-based approach to machine reading comprehension that uses a natural language question to test a system's comprehension of a passage of text.
Outcome: The proposed questions have surface cues or other biases that allow a model to shortcut the intended reasoning process.
Re-Examining Calibration: The Case of Question Answering (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing calibration methods do not provide significant gains in accuracy.
Approach: They propose a new calibration metric that better captures whether the model assigns low confidence to wrong predictions and high confidence to correct predictions.
Outcome: The proposed calibration method better captures whether the model assigns low confidence to wrong predictions and high confidence to correct predictions.
Zero- and Few-Shot NLP with Pretrained Language Models (2022.acl-tutorials)

Copied to clipboard

Challenge: a tutorial aims to introduce NLP researchers to the latest techniques for learning from little-to-no data . aims at bringing interested researchers up to speed about the latest and ongoing techniques .
Approach: They aim to introduce techniques for learning from little-to-no data using pretrained language models.
Outcome: This tutorial aims to bring interested NLP researchers up to speed about recent techniques . it will cover methods from manual engineering, better inference algorithms to better tuning methods .
Noisy Channel Language Model Prompting for Few-Shot Text Classification (2022.acl-long)

Copied to clipboard

Challenge: Prior work has suggested methods for finding better prompt or scoring of the output from the model.
Approach: They propose a noisy channel approach for language model prompting in few-shot text classification by in-context demonstration or prompt tuning.
Outcome: The proposed model outperforms direct models in both demonstration and prompt tuning.
AmbigQA: Answering Ambiguous Open-domain Questions (2020.emnlp-main)

Copied to clipboard

Challenge: Existing open-domain question answering systems assume questions have a single welldefined answer.
Approach: They propose an open-domain question answering task which involves finding every plausible answer and rewriting the question for each one to resolve the ambiguity.
Outcome: The proposed task is based on a dataset covering 14,042 open-domain questions . it shows that strong models benefit from weakly supervised learning .
Rethinking the Role of Demonstrations: What Makes In-Context Learning Work? (2022.emnlp-main)

Copied to clipboard

Challenge: Large language models can in-context learn by conditioning on a few input-label pairs and making predictions for new inputs.
Approach: They propose to use ground truth demonstrations to replace labels in demonstrations . they also show that other aspects of the demonstrations are key drivers of endtask performance .
Outcome: The proposed model outperforms zeroshot inference on a wide range of tasks using ground truth demonstrations.
Joint Passage Ranking for Diverse Multi-Answer Retrieval (2021.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to multi-answer retrieval cannot reason about the set of passages jointly.
Approach: They propose a joint passage retrieval model focusing on reranking to solve multi-answer retrieval problem.
Outcome: The proposed model outperforms baseline models on three multi-answer datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations