Challenge: Existing question-answering systems focus on answering individual questions, assuming they are devoid of context.
Approach: They propose to ask multiple related questions in a dataset that includes human-authored questions.
Outcome: The proposed system can answer human-authored questions better than existing systems.

Similar Papers

Can You Unpack That? Learning to Rewrite Questions-in-Context (D19-1)

Copied to clipboard

Challenge: Existing QA datasets lack key NLP problems like coreference and ellipsis resolution.
Approach: They propose a task of question-in-context rewriting to rewrite a context-dependent question into a self-contained question with the same answer.
Outcome: The proposed task is based on a dataset of 40,527 questions based in QuAC . it requires models to link questions together to resolve conversational dependencies .
Open-Domain Question Answering (2020.acl-tutorials)

Copied to clipboard

Challenge: tutorial provides a comprehensive overview of cutting-edge research in open-domain question answering (QA)
Approach: tutorial provides a comprehensive overview of cutting-edge research in open-domain question answering . focus will shift to cutting- edge models proposed for open- domain QA .
Outcome: The tutorial will cover cutting-edge research in open-domain question answering (QA) it will cover two-stage retriever-reader approaches, dense retriever and end-to-end training, and retriever free methods .
A guide to the dataset explosion in QA, NLI, and commonsense reasoning (2020.coling-tutorials)

Copied to clipboard

Challenge: a tutorial aims to provide an up-to-date guide to the recent datasets . the target audience is the NLP practitioners who are lost in dozens of the recent data sets.
Approach: This tutorial provides an up-to-date guide to the recent datasets . it surveys old and new methodological issues with dataset construction .
Outcome: This tutorial aims to provide an up-to-date guide to the recent datasets . it surveys the old and new methodological issues with dataset construction .
C-MORE: Pretraining to Answer Open-Domain Questions by Consulting Millions of References (2022.acl-short)

Copied to clipboard

Challenge: Existing approaches to pretrain open-domain question answering systems lack task-specific annotations.
Approach: They propose to pretrain a two-stage open-domain question answering system with strong transfer capabilities by using a dictionary and a large-scale corpus.
Outcome: The proposed approach leads to 2%-10% gains in top-20 accuracy and improves with reader.
QuAC: Question Answering in Context (D18-1)

Copied to clipboard

Challenge: a dataset for Question Answering in Context contains 14K information-seeking QA dialogs . questions are often more open-ended, unanswerable, or only meaningful within the dialog context .
Approach: They propose a dataset for Question Answering in Context that contains 14K dialogs . they use a student to ask questions about a Wikipedia section and a teacher to answer them .
Outcome: The proposed dataset underperforms humans in a number of reference models . the dataset contains 14K information-seeking dialogs over sections from Wikipedia .
IfQA: A Dataset for Open-domain Question Answering under Counterfactual Presuppositions (2023.emnlp-main)

Copied to clipboard

Challenge: Existing open-domain QA tasks focus on questions whose answer can be deduced directly from global factual knowledge.
Approach: They propose a dataset where each question is based on a counterfactual presupposition via an "if" clause.
Outcome: The IfQA dataset contains 3,800 questions that were annotated by crowdworkers on relevant Wikipedia passages.
DBQR-QA: A Question Answering Dataset on a Hybrid of Database Querying and Reasoning (2024.findings-acl)

Copied to clipboard

Challenge: Question answering (QA) is a fundamental task in the field of Natural Language Processing (NLP).
Approach: They propose a database querying and reasoning dataset for question answering that is designed to accommodate sequential questions and multi-hop queries.
Outcome: The proposed dataset better mirrors the dynamics of real-world information retrieval and analysis with a particular focus on the financial reports of US companies.
Question Answering in the Biomedical Domain (P19-2)

Copied to clipboard

Challenge: False positive questions require specific knowledge, common sense or a procedure due to ambiguity or the scope of the question.
Approach: False q is a question answering technique that uses natural language to find an answer . Falsity is based on a lexical gap and quality of answer spans .
Outcome: Using the proposed system, patients can self-diagnose without sacrificing quality of answer spans.
Question Answering over Tabular Data with DataBench: A Large-Scale Empirical Evaluation of LLMs (2024.lrec-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) are showing emerging abilities, but they are not large enough to assess their capabilities.
Approach: They propose a benchmark that compares large language models with open and closed source models.
Outcome: The proposed benchmark compares open and closed-source models with open-source and closed source models.
Question and Answer Test-Train Overlap in Open-Domain Question Answering Datasets (2021.eacl-main)

Copied to clipboard

Challenge: a recent study examines the ability of Open-Domain Question Answering models to produce answers to factoid questions . a large number of models have been used to study the performance of open-domain QA datasets .
Approach: They evaluate open-domain question answering models to see what they can generalize . they find that all models perform substantially worse on questions that cannot be memorized from train sets .
Outcome: The proposed model outperforms a closed-book QA model on the open-domain datasets.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations