Challenge: Open-domain Question Answering (OpenQA) aims at answering factual questions using an external large-scale knowledge corpus.
Approach: They propose a retrieval-augmented approach to QA that focuses on retrieving relevant knowledge from an external corpus.
Outcome: The proposed model can generalize to completely different knowledge domains while adapting to updated versions of the same knowledge corpus and switching to completely new knowledge domain.

Similar Papers

To Adapt or to Annotate: Challenges and Interventions for Domain Adaptation in Open-Domain Question Answering (2023.acl-long)

Copied to clipboard

Challenge: Recent advances in open-domain question answering have demonstrated impressive accuracy on general-purpose domains like Wikipedia.
Approach: They propose a more realistic end-to-end domain shift evaluation setting covering five diverse domains to assess model adaption.
Outcome: The proposed model improves by 24 points when adapted to unsupervised datasets.
RobustQA: Benchmarking the Robustness of Domain Adaptation for Open-Domain Question Answering (2023.findings-acl)

Copied to clipboard

Challenge: Existing ODQA datasets consist mainly of Wikipedia corpus, and are insufficient to study models’ generalizability across diverse domains.
Approach: They propose a benchmark to evaluate ODQA's domain robustness using Wikipedia corpus . they annotate QA pairs in retrieval datasets with rigorous quality control .
Outcome: The proposed benchmark improves model performance on annotated QA pairs in retrieval datasets with rigorous quality control.
Exploring Language Model Generalization in Low-Resource Extractive QA (2025.coling-main)

Copied to clipboard

Challenge: Existing LLMs struggle with dataset demands of closed domains such as medicine and law . current LLM performance in closed domain is lacking, even on traditional tasks such as Natural Language Inference .
Approach: They investigate Extractive Question Answering (EQA) with Large Language Models (LLMs) under domain drift . they find that LLMs struggle with dataset demands of closed domains .
Outcome: The proposed model performs poorly in extractive question answering tasks under domain drift . the proposed model can generalize to domains that require specific knowledge without training .
Learning to Generalize for Cross-domain QA (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods for QA are hampered by increased training costs . current methods suffer significant performance degradation when applied to out-of-domain examples.
Approach: They propose a method that combines prompting methods and linear probing with fine-tuning strategy, which does not entail additional cost.
Outcome: The proposed method outperforms state-of-the-art baselines with an average increase in F1 score of 4.5%-7.9%.
Challenges in Generalization in Open Domain Question Answering (2022.findings-naacl)

Copied to clipboard

Challenge: Recent work on Open Domain Question Answering has shown that there is a large discrepancy in model performance between novel test questions and those that largely overlap with training questions.
Approach: They introduce and annotate questions according to three categories that measure training set overlap, compositional generalization, and novel-entity generalization.
Outcome: The proposed models perform better on established datasets and lower on comp-gen/novel-entity questions than on the full test set.
A Survey for Efficient Open Domain Question Answering (2023.acl-long)

Copied to clipboard

Challenge: Open domain question answering (ODQA) is a longstanding task that can answer factoid questions without explicit evidence in natural language processing (NLP).
Approach: They propose to use open domain question answering to answer factual questions from a large knowledge corpus without explicit evidence.
Outcome: The proposed models can answer factoid questions from a large knowledge corpus without explicit evidence.
Relevance-guided Supervision for OpenQA with ColBERT (2021.tacl-1)

Copied to clipboard

Challenge: Recent work has focused on learning to retrieve passages for open-domain question answering . if notions of relevance are not tailored to questions, the MRC model will not reliably see the best passages .
Approach: They propose a retrieval model that uses coarse-grained vector representations of questions and passages to adapt it to OpenQA.
Outcome: The proposed system improves OpenQA retrieval on Natural Questions, SQuAD, and TriviaQA.
Generalizing Question Answering System with Pre-trained Language Model Fine-tuning (D19-58)

Copied to clipboard

Challenge: Existing methods focus on improving in-domain performance, leaving open the question of how they can generalize to out-of-domain and unseen RC tasks.
Approach: They propose a multi-task learning framework that learns the shared representation across different tasks and builds on a large pre-trained language model and fine-tuned on multiple RC datasets.
Outcome: The proposed framework improves the BERT-Large baseline by 8.39 and 7.22 respectively.
C-MORE: Pretraining to Answer Open-Domain Questions by Consulting Millions of References (2022.acl-short)

Copied to clipboard

Challenge: Existing approaches to pretrain open-domain question answering systems lack task-specific annotations.
Approach: They propose to pretrain a two-stage open-domain question answering system with strong transfer capabilities by using a dictionary and a large-scale corpus.
Outcome: The proposed approach leads to 2%-10% gains in top-20 accuracy and improves with reader.
Improving Retrieval Augmented Open-Domain Question-Answering with Vectorized Contexts (2024.findings-acl)

Copied to clipboard

Challenge: Retrieval Augmented Generation can be used to process long contexts in Open-Domain Question-Answering tasks.
Approach: They propose a method to cover longer contexts in Open-Domain Question-Answering tasks by using a small encoder language model and cross-attention with origin inputs.
Outcome: The proposed method can cover longer contexts while keeping the computing requirements close to the baseline.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations