Challenge: Pre-trained retrieval models often face challenges in zero-shot retrieval for knowledge-based question answering .
Approach: SQUARE is a method for corpus-specific unsupervised retrieval customization . it generates synthetic question-answer pairs from the corpus and fine-tunes it .
Outcome: SQUARE is a new method for corpus-specific unsupervised retrieval customization.

Similar Papers

Unsupervised Adaptation of Question Answering Systems via Generative Self-training (2020.emnlp-main)

Copied to clipboard

Challenge: Supervised self-training methods have transformed applied machine learning . however, adapting to target data has received little attention .
Approach: They propose a method to generate synthetic QA pairs for unsupervised self adaptation . they use massive amounts of data to simulate self-supervised tasks .
Outcome: The proposed method improves QA systems significantly by using less data and training computation than existing augmentation approaches.
Adaptive Document Retrieval for Deep Question Answering (D18-1)

Copied to clipboard

Challenge: Existing methods for deep question answering do not understand the exact interplay between document retrieval and machine comprehension.
Approach: They propose an adaptive document retrieval model that learns the optimal document number, conditional on the size of the corpus and the query.
Outcome: The proposed model outperforms state-of-the-art methods on multiple benchmark datasets and in the context of corpora with variable sizes.
Search-Adaptor: Embedding Customization for Information Retrieval (2024.acl-long)

Copied to clipboard

Challenge: Existing methods to embed text in large language models are limited to zero-shot setups and can be integrated with any LLM.
Approach: They propose a method for customizing LLMs for information retrieval by modifying the embeddings generated by pre-trained LLM models and can be integrated with any LLM.
Outcome: The proposed method improves performance on English, multilingual, and multimodal retrieval datasets by 5% over 14 BEIR datasets.
On Synthetic Data Strategies for Domain-Specific Generative Retrieval (2025.acl-long)

Copied to clipboard

Challenge: Generative retrieval models can be used to generate ranked lists of potentially relevant document identifiers for a user query.
Approach: They propose a synthetic data generation strategy for a two-stage training framework that focuses on learning to decode document identifiers from queries and a strategy for mining hard negatives based on initial model's predictions.
Outcome: The proposed model can generate ranked lists of potentially relevant document identifiers for a user query and then refine ranking through preference learning.
Synthetic QA Corpora Generation with Roundtrip Consistency (P19-1)

Copied to clipboard

Challenge: Existing methods for generating synthetic question answering corpora are not suitable for QA, but can be constructed from widely available natural text.
Approach: They propose a method for generating synthetic question answering corpora by combining question generation and answer extraction models and filtering the results to ensure roundtrip consistency.
Outcome: The proposed model achieves exact match and F1 at less than 0.1% and 0.4% from human performance on SQuAD2 and NQ.
DUQGen: Effective Unsupervised Domain Adaptation of Neural Rankers by Diversifying Synthetic Query Generation (2024.naacl-long)

Copied to clipboard

Challenge: State-of-the-art rankers pre-trained on large task-specific training data such as MS-MARCO exhibit strong performance on various ranking tasks without domain adaptation, also called zero-shot.
Approach: They propose a method to generate unsupervised domain adaptation for ranking using large-scale task-specific training data such as MS-MARCO and Wikipedia retrieval.
Outcome: The proposed method outperforms all zero-shot baselines and significantly outperfies the SOTA baselines on 16 out of 18 datasets, for an average of 4% relative improvement across all datasets.
Zero-shot Neural Passage Retrieval via Domain-targeted Synthetic Question Generation (2021.eacl-main)

Copied to clipboard

Challenge: Recent advances in neural retrieval have led to advancements on document, passage and knowledge-base benchmarks.
Approach: They propose an approach to zero-shot learning for passage retrieval that uses synthetic question generation to close this gap.
Outcome: The proposed approach can exceed term-based techniques on document retrieval benchmarks by using domain-targeted synthetic question generation.
UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods for information retrieval tasks require large labeled datasets for fine-tuning, but they can experience significant drops in accuracy due to distribution shifts from the training to the target domain.
Approach: They propose a method for using large language models to generate large numbers of synthetic queries cheaply using an expensive LLM.
Outcome: The proposed method boosts zero-shot accuracy in long-tail domains and achieves substantially lower latency than standard reranking methods.
Disentangling Questions from Query Generation for Task-Adaptive Retrieval (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing work generates synthetic queries from domain-specific documents to jointly train the retriever.
Approach: They propose a query generator that better adapts to wide search intents expressed in the BeIR benchmark.
Outcome: The proposed query generator outperforms baselines and existing models on tasks with underexplored intents while using a query generator 47 times smaller than the previous state-of-the-art.
Back-Training excels Self-Training at Unsupervised Domain Adaptation of Question Generation and Passage Retrieval (2021.emnlp-main)

Copied to clipboard

Challenge: Using self-training to train unsupervised domains can be expensive, resulting in poor generalization due to distributional shift.
Approach: They propose to use unaligned data to train unsupervised domain adaptation models using cheap synthetically generated labeled data.
Outcome: The proposed method significantly outperforms self-training on question generation and passage retrieval domains and on MLQuestions and PubMedQA.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations