SQUARE: Unsupervised Retrieval Adaptation via Synthetic Data (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Pre-trained retrieval models often face challenges in zero-shot retrieval for knowledge-based question answering . |
| Approach: | SQUARE is a method for corpus-specific unsupervised retrieval customization . it generates synthetic question-answer pairs from the corpus and fine-tunes it . |
| Outcome: | SQUARE is a new method for corpus-specific unsupervised retrieval customization. |
Similar Papers
Unsupervised Adaptation of Question Answering Systems via Generative Self-training (2020.emnlp-main)
Copied to clipboard
| Challenge: | Supervised self-training methods have transformed applied machine learning . however, adapting to target data has received little attention . |
| Approach: | They propose a method to generate synthetic QA pairs for unsupervised self adaptation . they use massive amounts of data to simulate self-supervised tasks . |
| Outcome: | The proposed method improves QA systems significantly by using less data and training computation than existing augmentation approaches. |
Adaptive Document Retrieval for Deep Question Answering (D18-1)
Copied to clipboard
| Challenge: | Existing methods for deep question answering do not understand the exact interplay between document retrieval and machine comprehension. |
| Approach: | They propose an adaptive document retrieval model that learns the optimal document number, conditional on the size of the corpus and the query. |
| Outcome: | The proposed model outperforms state-of-the-art methods on multiple benchmark datasets and in the context of corpora with variable sizes. |
Search-Adaptor: Embedding Customization for Information Retrieval (2024.acl-long)
Copied to clipboard
| Challenge: | Existing methods to embed text in large language models are limited to zero-shot setups and can be integrated with any LLM. |
| Approach: | They propose a method for customizing LLMs for information retrieval by modifying the embeddings generated by pre-trained LLM models and can be integrated with any LLM. |
| Outcome: | The proposed method improves performance on English, multilingual, and multimodal retrieval datasets by 5% over 14 BEIR datasets. |
On Synthetic Data Strategies for Domain-Specific Generative Retrieval (2025.acl-long)
Copied to clipboard
| Challenge: | Generative retrieval models can be used to generate ranked lists of potentially relevant document identifiers for a user query. |
| Approach: | They propose a synthetic data generation strategy for a two-stage training framework that focuses on learning to decode document identifiers from queries and a strategy for mining hard negatives based on initial model's predictions. |
| Outcome: | The proposed model can generate ranked lists of potentially relevant document identifiers for a user query and then refine ranking through preference learning. |
Synthetic QA Corpora Generation with Roundtrip Consistency (P19-1)
Copied to clipboard
| Challenge: | Existing methods for generating synthetic question answering corpora are not suitable for QA, but can be constructed from widely available natural text. |
| Approach: | They propose a method for generating synthetic question answering corpora by combining question generation and answer extraction models and filtering the results to ensure roundtrip consistency. |
| Outcome: | The proposed model achieves exact match and F1 at less than 0.1% and 0.4% from human performance on SQuAD2 and NQ. |
DUQGen: Effective Unsupervised Domain Adaptation of Neural Rankers by Diversifying Synthetic Query Generation (2024.naacl-long)
Copied to clipboard
| Challenge: | State-of-the-art rankers pre-trained on large task-specific training data such as MS-MARCO exhibit strong performance on various ranking tasks without domain adaptation, also called zero-shot. |
| Approach: | They propose a method to generate unsupervised domain adaptation for ranking using large-scale task-specific training data such as MS-MARCO and Wikipedia retrieval. |
| Outcome: | The proposed method outperforms all zero-shot baselines and significantly outperfies the SOTA baselines on 16 out of 18 datasets, for an average of 4% relative improvement across all datasets. |
Zero-shot Neural Passage Retrieval via Domain-targeted Synthetic Question Generation (2021.eacl-main)
Copied to clipboard
| Challenge: | Recent advances in neural retrieval have led to advancements on document, passage and knowledge-base benchmarks. |
| Approach: | They propose an approach to zero-shot learning for passage retrieval that uses synthetic question generation to close this gap. |
| Outcome: | The proposed approach can exceed term-based techniques on document retrieval benchmarks by using domain-targeted synthetic question generation. |
UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers (2023.emnlp-main)
Copied to clipboard
Jon Saad-Falcon, Omar Khattab, Keshav Santhanam, Radu Florian, Martin Franz, Salim Roukos, Avirup Sil, Md Sultan, Christopher Potts
| Challenge: | Existing methods for information retrieval tasks require large labeled datasets for fine-tuning, but they can experience significant drops in accuracy due to distribution shifts from the training to the target domain. |
| Approach: | They propose a method for using large language models to generate large numbers of synthetic queries cheaply using an expensive LLM. |
| Outcome: | The proposed method boosts zero-shot accuracy in long-tail domains and achieves substantially lower latency than standard reranking methods. |
Disentangling Questions from Query Generation for Task-Adaptive Retrieval (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing work generates synthetic queries from domain-specific documents to jointly train the retriever. |
| Approach: | They propose a query generator that better adapts to wide search intents expressed in the BeIR benchmark. |
| Outcome: | The proposed query generator outperforms baselines and existing models on tasks with underexplored intents while using a query generator 47 times smaller than the previous state-of-the-art. |
Back-Training excels Self-Training at Unsupervised Domain Adaptation of Question Generation and Passage Retrieval (2021.emnlp-main)
Copied to clipboard
| Challenge: | Using self-training to train unsupervised domains can be expensive, resulting in poor generalization due to distributional shift. |
| Approach: | They propose to use unaligned data to train unsupervised domain adaptation models using cheap synthetically generated labeled data. |
| Outcome: | The proposed method significantly outperforms self-training on question generation and passage retrieval domains and on MLQuestions and PubMedQA. |