Challenge: Unlike previous datasets, the general knowledge is textual and not tied to a fixed set of relationships.
Approach: They introduce the first open-domain dataset, called QuaRTz, for reasoning about textual qualitative relationships.
Outcome: The proposed dataset is the first open-domain dataset for reasoning about qualitative relationships.

Similar Papers

QUARTZ: QA-based Unsupervised Abstractive Refinement for Task-oriented Dialogue Summarization (2025.findings-emnlp)

Copied to clipboard

Challenge: a framework for task-oriented utility-based dialogue summarization is proposed . QUARTZ is a tool for task summarizing dialogues, but its outputs lack task-specific focus.
Approach: They propose a framework for task-oriented utility-based dialogue summarization . QUARTZ generates summaries and question-answer pairs from a dialogue in a zero-shot manner .
Outcome: The proposed framework achieves competitive results in zero-shot settings, rivaling fully-supervised State-of-the-Art methods.
A guide to the dataset explosion in QA, NLI, and commonsense reasoning (2020.coling-tutorials)

Copied to clipboard

Challenge: a tutorial aims to provide an up-to-date guide to the recent datasets . the target audience is the NLP practitioners who are lost in dozens of the recent data sets.
Approach: This tutorial provides an up-to-date guide to the recent datasets . it surveys old and new methodological issues with dataset construction .
Outcome: This tutorial aims to provide an up-to-date guide to the recent datasets . it surveys the old and new methodological issues with dataset construction .
emrQA: A Large Corpus for Question Answering on Electronic Medical Records (D18-1)

Copied to clipboard

Challenge: Existing annotations for other NLP tasks are used to generate domain-specific large-scale question answering (QA) datasets.
Approach: They propose to re-purpose existing annotations for other NLP tasks by generating a large-scale question answering corpus using 1 million questions-logical form and 400,000+ question-answer evidence pairs.
Outcome: The proposed model can be trained to learn domain-specific large-scale question answering (QA) datasets.
PragmatiCQA: A Dataset for Pragmatic Question Answering in Conversations (2023.findings-acl)

Copied to clipboard

Challenge: Mars? - PragmatiCQA
Approach: Mars? - The Paper .
Outcome: The proposed dataset features 6873 QA pairs that explores pragmatic reasoning in conversations over a diverse set of topics.
IfQA: A Dataset for Open-domain Question Answering under Counterfactual Presuppositions (2023.emnlp-main)

Copied to clipboard

Challenge: Existing open-domain QA tasks focus on questions whose answer can be deduced directly from global factual knowledge.
Approach: They propose a dataset where each question is based on a counterfactual presupposition via an "if" clause.
Outcome: The IfQA dataset contains 3,800 questions that were annotated by crowdworkers on relevant Wikipedia passages.
Open-WikiTable : Dataset for Open Domain Question Answering with Complex Reasoning over Table (2023.findings-acl)

Copied to clipboard

Challenge: Open-WikiTable is the first open domain question answering dataset that requires complex reasoning over tables.
Approach: They propose to use open-domain question answering over tables to extract questions from tables.
Outcome: The dataset is publicly available. it is built upon WikiSQL and WikiTableQuestions.
TIGQA: An Expert-Annotated Question-Answering Dataset in Tigrinya (2024.lrec-main)

Copied to clipboard

Challenge: Existing annotated datasets for NLP tasks in languages with limited resources are limited.
Approach: They propose to use machine translation to convert existing Tigrinya dataset into a Tigrina dataset in SQuAD format.
Outcome: The proposed dataset is an expert-annotated Tigrinya dataset with 2,685 question-answer pairs covering 122 diverse topics.
SParC: Cross-Domain Semantic Parsing in Context (P19-1)

Copied to clipboard

Challenge: Xu et al., 2017): a dataset for cross-domain semantic parsing in context with 4,298 question sequences.
Approach: They present a dataset for cross-domainSemanticParsing inContext that consists of 4,298 coherent question sequences.
Outcome: The proposed dataset demonstrates that it has greater semantic diversity and can be generalized to unseen domains due to its cross-domain nature and the unseened databases at test time.
More Data, More Relations, More Context and More Openness: A Review and Outlook for Relation Extraction (2020.aacl-main)

Copied to clipboard

Challenge: Existing methods for extracting relational facts from text have been successful . but with explosion of Web text, human knowledge is increasing drastically .
Approach: They propose to improve relation extraction methods to extract relational facts from text . they analyze existing methods and show promising directions towards more powerful RE .
Outcome: The proposed methods can extract relational facts from text, but they are still lacking in the current field.
QLEVR: A Diagnostic Dataset for Quantificational Language and Elementary Visual Reasoning (2022.findings-naacl)

Copied to clipboard

Challenge: Synthetic datasets have been used to test visual question-answering datasets for reasoning abilities.
Approach: They propose a visual question-answering dataset that is minimally biased and diagnostic . they propose to use the dataset to test visual reasoning abilities .
Outcome: The proposed dataset is compared with existing models and shows it is far superior to existing models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations