Challenge: SQA is an emerging application of NLP in the medical, geography, and legal domains.
Approach: They propose a dataset of 1,981 scenarios and 4,110 multiple-choice questions in geography domain at high school level.
Outcome: The proposed dataset consists of 1,981 scenarios and 4,110 multiple-choice questions in the geography domain at high school level.

Similar Papers

GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning (2021.findings-acl)

Copied to clipboard

Challenge: Existing methods to solve geometric problems are dependent on handcraft rules and limited on small-scale datasets.
Approach: They propose a Geometric Question Answering dataset with 5,010 geometric problems with corresponding annotated programs to illustrate the solving process.
Outcome: The proposed method is significantly lower than human performance on the proposed dataset than on a publicly available dataset.
SPARTQA: A Textual Question Answering Benchmark for Spatial Reasoning (2021.naacl-main)

Copied to clipboard

Challenge: Existing studies have focused on the spatial reasoning capabilities of modern language models (LMs) however, there has been limited research into the spatial thinking capabilities of LMs.
Approach: They propose a question-answering (QA) benchmark for spatial reasoning on natural language text which contains more realistic spatial phenomena not covered by prior work.
Outcome: The proposed method significantly improves LMs' ability on spatial understanding, which in turn helps solve two external datasets, bAbI, and boolQ.
When Retriever-Reader Meets Scenario-Based Multiple-Choice Questions (2021.findings-emnlp)

Copied to clipboard

Challenge: Scenario-based question answering (SQA) requires retrieving and reading paragraphs from a large corpus to answer a question contextualized by a long scenario description.
Approach: They propose a model where the retriever is implicitly supervised only using QA labels via a novel word weighting mechanism.
Outcome: The proposed model outperforms strong baselines on multiple-choice questions in three datasets.
TIGQA: An Expert-Annotated Question-Answering Dataset in Tigrinya (2024.lrec-main)

Copied to clipboard

Challenge: Existing annotated datasets for NLP tasks in languages with limited resources are limited.
Approach: They propose to use machine translation to convert existing Tigrinya dataset into a Tigrina dataset in SQuAD format.
Outcome: The proposed dataset is an expert-annotated Tigrinya dataset with 2,685 question-answer pairs covering 122 diverse topics.
MultiSpanQA: A Dataset for Multi-Span Question Answering (2022.naacl-main)

Copied to clipboard

Challenge: Existing reading comprehension datasets focus on single-span answers, but multi-spread questions are less studied.
Approach: They propose a new reading comprehension dataset that focuses on multi-span questions . they introduce new metrics for the purposes of multi--spontaneous question answering evaluation .
Outcome: The proposed model beats baselines and achieves state-of-the-art on the existing dataset.
LifeQA: A Real-life Dataset for Video Question Answering (2020.lrec-1)

Copied to clipboard

Challenge: Existing video question answering datasets consist of movies and TV shows, but they are not representative of our day-to-day lives.
Approach: They propose a benchmark dataset for video question answering that focuses on day-to-day situations.
Outcome: The proposed dataset analyzes the challenging but realistic aspects of LifeQA . it consists of video clips and over 2.3k multiple-choice questions .
KQA Pro: A Dataset with Explicit Compositional Programs for Complex Question Answering over Knowledge Base (2022.acl-long)

Copied to clipboard

Challenge: Existing benchmarks for Complex KBQA lack compositional reasoning capabilities . Existing methods for Complex questions are poor in diversity or scale .
Approach: They propose a compositional programming language to represent the reasoning process of complex questions.
Outcome: The proposed dataset includes around 120K diverse natural language questions . it provides a compositional and interpretable programming language to represent the reasoning process of complex questions based on the proposed model .
ComQA: A Community-sourced Dataset for Complex Factoid Question Answering with Paraphrase Clusters (N19-1)

Copied to clipboard

Challenge: ComQA dataset captures question phenomena and the diverse ways in which they are formulated.
Approach: They propose a large dataset of real user questions that captures question phenomena and the diverse ways in which they are formulated.
Outcome: The proposed dataset can be a driver of future research on factoid question answering (QA).
ConditionalQA: A Complex Reading Comprehension Dataset with Conditional Answers (2022.acl-long)

Copied to clipboard

Challenge: Existing datasets for reading comprehension have deterministic answers, but questions in the real world do not always have definite answers.
Approach: They propose a Question Answering (QA) dataset that contains complex questions with conditional answers.
Outcome: The proposed dataset will motivate further research in answering complex questions over long documents.
Chart Question Answering from Real-World Analytical Narratives (2025.acl-srw)

Copied to clipboard

Challenge: a dataset for chart question answering is constructed from visualization notebooks . data visualizations are an essential modality for communicating complex information about data.
Approach: They propose a dataset for chart question answering constructed from visualization notebooks . they use real-world, multi-view charts paired with natural language questions .
Outcome: The proposed dataset is constructed from student-authored visualization notebooks . it features real-world, multi-view charts paired with natural language questions . initial evaluations highlight significant performance gaps .

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations