Challenge: Existing models have outperformed humans on question answering datasets, but they have yet to outperform humans on the task of question answering itself.
Approach: They evaluate BERT-based question answering models on their generalizability to out-of-domain examples, responses to missing or incorrect data, and ability to handle question variations.
Outcome: The proposed models outperform human baselines on the widely-used SQuAD 1.1 and SQu AD 2.0 datasets.

Similar Papers

Generalizing Question Answering System with Pre-trained Language Model Fine-tuning (D19-58)

Copied to clipboard

Challenge: Existing methods focus on improving in-domain performance, leaving open the question of how they can generalize to out-of-domain and unseen RC tasks.
Approach: They propose a multi-task learning framework that learns the shared representation across different tasks and builds on a large pre-trained language model and fine-tuned on multiple RC datasets.
Outcome: The proposed framework improves the BERT-Large baseline by 8.39 and 7.22 respectively.
Towards Interpreting BERT for Reading Comprehension Based QA (2020.emnlp-main)

Copied to clipboard

Challenge: Pretrained language models such as ELMO and XLNet have achieved state-of-the-art performance on various NLP tasks.
Approach: They propose to define a layer’s role or functionality using Integrated Gradients and perform preliminary analysis across all layers.
Outcome: The proposed model performs better than existing models on RCQA and ELMO, but it lacks the human-level performance needed to perform the task.
Single-dataset Experts for Multi-dataset Question Answering (2021.emnlp-main)

Copied to clipboard

Challenge: Prior work has focused on training one network on multiple datasets to build a model that performs well on all of the training datasets and generalizes and transfers better to new datasets.
Approach: They combine multiple reading comprehension datasets to build a multi-dataset question answering model with an ensemble of single-data set experts.
Outcome: The proposed model outperforms baseline models in in-distribution accuracy and generalization and transfer performance.
A Recurrent BERT-based Model for Question Generation (D19-58)

Copied to clipboard

Challenge: Existing QG models rely on recurrent neural networks (RNNs) but the inherent sequential nature of the RNN models suffers from the problem of handling long sequences.
Approach: They propose to employ a pre-trained BERT language model to tackle question generation tasks.
Outcome: The proposed model outperforms the existing models on the question-answering dataset SQuAD and advances the BLEU 4 score from 16.85 to 22.17.
Bend but Don’t Break? Multi-Challenge Stress Test for QA Models (D19-58)

Copied to clipboard

Challenge: a gap remains in reasoning ability compared to a human, and performance tends to degrade when models are exposed to less-constrained tasks.
Approach: They conduct extensive qualitative and quantitative analyses on the results of four models across four datasets . they relate common errors to model capabilities and discuss a way forward .
Outcome: The proposed model performance is based on the results of four models across four datasets.
Learning to Explain: Datasets and Models for Identifying Valid Reasoning Chains in Multihop Question-Answering (2020.emnlp-main)

Copied to clipboard

Challenge: despite rapid progress in multihop question-answering, models still have trouble explaining why an answer is correct.
Approach: They propose three explanation datasets in which explanations from corpus facts are annotated . they first annotate multiple candidate explanations for each answer, then use crowd-sourcing perturbations to test generalization .
Outcome: The proposed datasets improve explanation quality but still behind the upper bound . the proposed dataset can be used to improve explanations using a BERT-based classifier .
Look at the First Sentence: Position Bias in Question Answering (2020.emnlp-main)

Copied to clipboard

Challenge: Extractive question answering models are trained to predict start and end positions of answers . recent QA models outperform humans in some datasets due to their simplicity and effectiveness.
Approach: They propose to use prior distribution of answer positions as a bias model to reduce position bias.
Outcome: The proposed model outperforms BERT from 37.48% to 81.64% when trained on a biased SQUAD dataset.
Knowing More About Questions Can Help: Improving Calibration in Question Answering (2021.findings-acl)

Copied to clipboard

Challenge: Existing work on calibration focuses on model confidence, such as the max probability of the predicted class.
Approach: They propose a calibration method which estimates whether model correctly predicts answer for each question.
Outcome: The proposed calibration method achieves 5-10% gains on reading comprehension benchmarks.
What do we expect from Multiple-choice QA Systems? (2020.findings-emnlp)

Copied to clipboard

Challenge: Recent work has shown that good performance on a dataset might not correlate well with human’s expectations from models that “understand” language.
Approach: They propose to train a top performing multiple choice question answering model against expectations from models that "understand" language.
Outcome: The proposed training paradigm leads to a model that performs on par with the original model while better satisfying our expectations.
Handling Anomalies of Synthetic Questions in Unsupervised Question Answering (2020.coling-main)

Copied to clipboard

Challenge: Existing approaches to improve unsupervised Question Answering (UQA) are expensive and require additional datasets.
Approach: They propose an unsupervised QA approach that generates QA training data automatically.
Outcome: The proposed method improves unsupervised QA significantly across a number of QA tasks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations