Challenge: Despite advancement of question answering systems, generalizability of QA models is a topic of concern.
Approach: They propose to use a neural paraphrasing model to generate multiple paraphrased questions for a given source question and a set of paraphrase suggestions to re-train the models.
Outcome: The proposed approach requires no human intervention to re-train the models for improved robustness to question paraphrasing.

Similar Papers

Improving the Robustness of QA Models to Challenge Sets with Variational Question-Answer Pair Generation (2021.acl-srw)

Copied to clipboard

Challenge: Existing data augmentation methods for reading comprehension lack robustness to challenge sets whose distribution is different from that of training sets.
Approach: They propose a question-answer pair generation method that generates multiple diverse QA pairs from a paragraph to mitigate this problem.
Outcome: The proposed model improves the accuracy of 12 challenge sets and the in-distribution accuracy.
Bend but Don’t Break? Multi-Challenge Stress Test for QA Models (D19-58)

Copied to clipboard

Challenge: a gap remains in reasoning ability compared to a human, and performance tends to degrade when models are exposed to less-constrained tasks.
Approach: They conduct extensive qualitative and quantitative analyses on the results of four models across four datasets . they relate common errors to model capabilities and discuss a way forward .
Outcome: The proposed model performance is based on the results of four models across four datasets.
Handling Anomalies of Synthetic Questions in Unsupervised Question Answering (2020.coling-main)

Copied to clipboard

Challenge: Existing approaches to improve unsupervised Question Answering (UQA) are expensive and require additional datasets.
Approach: They propose an unsupervised QA approach that generates QA training data automatically.
Outcome: The proposed method improves unsupervised QA significantly across a number of QA tasks.
Exploring The Landscape of Distributional Robustness for Question Answering Models (2022.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for predicting distributional robustness fail to generalize reliably in a variety of test conditions.
Approach: They conduct a large empirical evaluation to investigate the landscape of distributional robustness in question answering.
Outcome: The proposed methods are more robust to distribution shifts than fully fine-tuned models, and few-shot prompt models exhibit better robustness than few- shot prompt models.
How to Ask Good Questions? Try to Leverage Paraphrases (2020.acl-main)

Copied to clipboard

Challenge: Existing methods to generate human-like questions rely on paraphrases to generate good questions.
Approach: They propose to integrate paraphrase knowledge into question generation to generate human-like questions by combining paraphrases with a back-translation method.
Outcome: The proposed model achieves obvious performance gain over several strong baselines and human evaluation validates that it can ask questions of high quality by leveraging paraphrase knowledge.
Can Pre-training help VQA with Lexical Variations? (2020.findings-emnlp)

Copied to clipboard

Challenge: Visual Question Answering (VQA) models are closing the gap between oracle performance and model robustness.
Approach: They propose to use language & cross-modal pre-training to investigate the robustness of VQA models towards lexical variations.
Outcome: The proposed model architectures and training techniques improve the performance of the VQA-Rephrasings dataset on rephrased questions.
Beyond Static Synthetic Noise: Assessing the Robustness of Large Language Models to Natural Context Variation in the Real World (2026.findings-acl)

Copied to clipboard

Challenge: Current robustness evaluation methods rely on static synthetic perturbations to stress-test models.
Approach: They propose a framework for automatically evaluating QA models under naturally occurring textual perturbations by replacing context passages with revised Wikipedia edit histories.
Outcome: The proposed framework replaces context passages with revised Wikipedia edit histories to improve model performance.
Robust Question Answering Through Sub-part Alignment (2021.naacl-main)

Copied to clipboard

Challenge: Current textual question answering models fail to generalize to out-of-domain settings.
Approach: They propose to decompose question and context into smaller units and align them to find the answer.
Outcome: The proposed model is more robust than the standard BERT QA model on adversarial and out-of-domain datasets.
Towards Robust Extractive Question Answering Models: Rethinking the Training Methodology (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing models lack robustness against distribution shifts and adversarial attacks when training on unanswerable questions in EQA datasets.
Approach: They propose a novel loss function for the EQA problem to improve the robustness of extractive question answering models by adding adversarial questions to a crowdsourcing process.
Outcome: The proposed method maintains in-domain performance while improving on out-of-domain datasets.
Addressing Semantic Drift in Question Generation for Semi-Supervised Question Answering (D19-1)

Copied to clipboard

Challenge: Existing QG models suffer from a “semantic drift” problem, i.e., the semantics of the model-generated question drifts away from the given context and answer.
Approach: They propose two semantics-enhanced rewards obtained from downstream question paraphrasing and question answering tasks to regularize the QG model to generate semantically valid questions.
Outcome: The proposed method achieves state-of-the-art performance w.r.t. traditional evaluation metrics and performs best on QA-based evaluation metrics.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations