Challenge: a large body of work has highlighted the brittleness of reading comprehension systems . a crowdsourced, adversarially-created, 55k-question benchmark requires a more comprehensive understanding of paragraphs .
Approach: They propose a reading comprehension benchmark that requires Discrete Reasoning over the content of paragraphs.
Outcome: The proposed benchmarks show that the best systems only achieve 38.4% F1 on the generalized accuracy metric, while human performance is 96%.

Similar Papers

Numerical reasoning in machine reading comprehension tasks: are we there yet? (2021.emnlp-main)

Copied to clipboard

Challenge: Numerical reasoning based machine reading comprehension models have achieved near-human performance on a variety of benchmarks, but are they capable of learning to reason?
Approach: They propose to use a DROP benchmark to measure machine reading comprehension and investigate models that have achieved near-human performance over standard metrics.
Outcome: The DROP benchmark has inspired the design of specialized BERT and embedding the results into a specialized model.
How Much Reading Does Reading Comprehension Require? A Critical Investigation of Popular Benchmarks (D18-1)

Copied to clipboard

Challenge: Recent research addresses reading comprehension, where examples consist of (question, passage, answer) tuples.
Approach: They establish sensible baselines for bAbI, SQuAD, CBT, CNN and Who-did-What datasets and compare them to their previous work.
Outcome: The proposed models perform on 14 out of 20 bAbI, SQuAD, CBT, CNN and Who-did-What datasets.
Giving BERT a Calculator: Finding Operations and Arguments with Reading Comprehension (D19-1)

Copied to clipboard

Challenge: End-to-end reading comprehension models have been successful at extracting text answers, but there are still problems with generalizing them to abstractive numerical reasoning.
Approach: They propose to augment a BERT-based reading comprehension model with a set of executable ‘programs’ which encompass simple arithmetic as well as extraction.
Outcome: The proposed model can perform 33% absolute improvement on the DROP dataset, with very few training examples.
Quoref: A Reading Comprehension Dataset with Questions Requiring Coreferential Reasoning (D19-1)

Copied to clipboard

Challenge: Existing reading comprehension benchmarks do not contain complex coreferential phenomena . obtaining questions focused on such phenomena is difficult because of lexical cues .
Approach: They propose to use a crowdsourced dataset to examine the ability of models to resolve coreference among entities in Wikipedia paragraphs.
Outcome: The proposed model performs significantly worse than humans on the reading comprehension benchmark . paragraphs and other longer texts typically make multiple references to the same entities .
NumNet: Machine Reading Comprehension with Numerical Reasoning (D19-1)

Copied to clipboard

Challenge: Existing numerical MRC models are weak in numerical reasoning, such as addition, subtraction, sorting and counting.
Approach: They propose a numerical MRC model that integrates numerical reasoning into existing MRC models and achieves an EM-score of 64.56% on the DROP dataset.
Outcome: The proposed model outperforms all existing machine reading comprehension models by considering the numerical relations among numbers on the DROP dataset.
Looking Beyond the Surface: A Challenge Set for Reading Comprehension over Multiple Sentences (N18-1)

Copied to clipboard

Challenge: Using a dataset of 6,500+ questions, we found that human solvers achieved an F1-score of 88.1%.
Approach: They propose a reading comprehension challenge in which questions can only be answered by taking into account information from multiple sentences.
Outcome: The proposed reading comprehension challenge is based on a reading comprehension dataset with 6,500+ questions and 1000+ paragraphs across 7 domains.
On Making Reading Comprehension More Comprehensive (D19-58)

Copied to clipboard

Challenge: Getting machines to "understand" text is a vast and long-standing problem, made more challenging by the fact that it is not even clear what it means to understand text.
Approach: They propose a question-based approach to machine reading comprehension that uses a natural language question to test a system's comprehension of a passage of text.
Outcome: The proposed questions have surface cues or other biases that allow a model to shortcut the intended reasoning process.
Benchmarking Machine Reading Comprehension: A Psychological Perspective (2021.eacl-main)

Copied to clipboard

Challenge: MRC is a task that tests the ability of a machine to read and understand unstructured text.
Approach: They propose a theoretical basis for the design of MRC datasets based on psychology and psychometrics and propose shortcut-proof questions and explanations as a part of the task design.
Outcome: The proposed datasets should evaluate the model's ability to understand context-dependent situations and ensure substantive validity by shortcut-proof questions and explanation as a part of the task design.
LILA: A Unified Benchmark for Mathematical Reasoning (2022.emnlp-main)

Copied to clipboard

Challenge: Towards evaluating and improving AI systems in this domain, we propose a mathematical reasoning benchmark based on 23 diversetasks .
Approach: They propose a mathematical reasoning benchmark that includes 23 diverse tasks . they extend the benchmark by collecting task instructions and solutions in the form of Python programs .
Outcome: The proposed model improves on multi-tasking while the best performing model only achieves 60.40%.
Probing Neural Network Comprehension of Natural Language Arguments (P19-1)

Copied to clipboard

Challenge: Argument Reasoning Comprehension Task (ARCT) focuses on inferences, not just discovering warrants.
Approach: They propose to build an adversarial dataset on which all models achieve random accuracy.
Outcome: The proposed dataset provides a more robust assessment of argument comprehension and should be adopted as the standard in future work.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations