Challenge: a new approach to training with binary relevance labels uses synthetic data . contrastive learning with binary correlations leaves out subtle nuances useful for ranking .
Approach: They propose to use waterstein distance as a loss function for training transformer-based retrievers with graduated relevance labels instead of real documents.
Outcome: The proposed method outperforms conventional training with InfoNCE by a large margin on MARCO and BEIR benchmarks without using real documents.

Similar Papers

Unsupervised Dense Retrieval with Relevance-Aware Contrastive Pre-Training (2023.findings-acl)

Copied to clipboard

Challenge: Dense retrievers have impressive performance, but their demand for abundant training data limits their application scenarios.
Approach: They propose a method which uses unlabeled data to construct pseudo-positive examples from unlabelled data and then contrastively weighs the contrastive loss of different pairs according to the estimated relevance.
Outcome: The proposed method beats the SOTA unsupervised Contriever model on BEIR and open-domain QA retrieval benchmarks and is a good few-shot learner.
Contrastive Multi-document Question Generation (2021.eacl-main)

Copied to clipboard

Challenge: Multi-document question generation focuses on generating a question that covers the common aspect of multiple documents, but a naive model trained only using the targeted document set may generate too generic questions that cover a larger scope than delineated by the document set.
Approach: They propose a contrastive learning strategy where given ‘positive’ and ‘negative’ sets of documents, generate a question that is closely related to the ‘positive' set but far away from the ‘negative' set.
Outcome: The proposed model significantly outperforms several strong baselines, as measured by automatic metrics and human evaluation.
On Synthetic Data Strategies for Domain-Specific Generative Retrieval (2025.acl-long)

Copied to clipboard

Challenge: Generative retrieval models can be used to generate ranked lists of potentially relevant document identifiers for a user query.
Approach: They propose a synthetic data generation strategy for a two-stage training framework that focuses on learning to decode document identifiers from queries and a strategy for mining hard negatives based on initial model's predictions.
Outcome: The proposed model can generate ranked lists of potentially relevant document identifiers for a user query and then refine ranking through preference learning.
Multi-stage Training with Improved Negative Contrast for Neural Passage Retrieval (2021.emnlp-main)

Copied to clipboard

Challenge: Existing neural firststage retrieval models overcome lexical gap issue by projecting query and document to a shared dense space.
Approach: They propose a multi-stage framework for neural passage retrieval using synthetic data, negative sampling, and fusion techniques.
Outcome: The proposed framework improves retrieval accuracy and enhances the negative contrast in both stages.
Exploring efficient zero-shot synthetic dataset generation for Information Retrieval (2024.findings-eacl)

Copied to clipboard

Challenge: Recent advances in large language models offer a new avenue of generating synthetic training data to train neural retrieval models for unlabelled data collections.
Approach: They propose a method to generate high-quality synthetic datasets using a small language model and a filtering mechanism to ensure the quality of generated questions.
Outcome: The proposed method outperforms unsupervised retrieval methods such as BM25 and pretrained monoT5.
Beyond Contrastive Learning: A Variational Generative Model for Multilingual Retrieval (2023.acl-long)

Copied to clipboard

Challenge: Contrastive learning is the dominant paradigm for learning text representations from parallel text, but finding negative examples can be expensive in terms of compute or manual effort.
Approach: They propose a generative model for learning multilingual text embeddings which encourages source separation in multilingual contexts by an approximation.
Outcome: The proposed model outperforms both a strong contrastive and generative baseline on a suite of tasks including semantic similarity, bitext mining, and cross-lingual question retrieval.
Beyond [CLS] through Ranking by Generation (2020.emnlp-main)

Copied to clipboard

Challenge: Recent work on generative ranking models for Information Retrieval has focused on discriminative methods that learn a similarity function to compare questions and candidates answers.
Approach: They propose to use a language model to train a ranking function that model the semantic similarity of documents and queries instead of discriminative ranking functions.
Outcome: The proposed approaches are as effective as state-of-the-art discriminative models for the answer selection task and show unlikelihood losses are reduced for IR.
Towards Robust Ranker for Text Retrieval (2023.findings-acl)

Copied to clipboard

Challenge: Existing methods for text retrieval are based on a 'retrieval & rerank' pipeline, which uses a fast retriever to fetch a set of top document candidates, while a robust ranker is based upon a weak negative mining during contrastive learning.
Approach: They propose a multi-adversarial training strategy that leverages multiple retrievers as generators to challenge a ranker.
Outcome: The proposed model outperforms the existing de facto ranker training paradigms on the passage retrieval benchmarks using BM25-reranking, full-ranking and retriever distillation.
Exploiting Asymmetry for Synthetic Training Data Generation: SynthIE and the Case of Information Extraction (2023.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) have great potential for synthetic data generation.
Approach: They show that large language models can generate useful data even for complex tasks . they use a symmetric task difficulty asymmetry to prompt an LLM to generate plausible input text for a target output structure.
Outcome: The proposed approach outperforms existing models by a substantial margin on closed information extraction tasks with 1.8M data points and 770M parameters.
Expand, Highlight, Generate: RL-driven Document Generation for Passage Reranking (2023.emnlp-main)

Copied to clipboard

Challenge: Existing studies use large language models to generate training data for ranking models.
Approach: They propose a pipeline that generates synthetic documents from queries using large language models . they propose RL-based reinforcement learning to optimize the pipeline .
Outcome: The proposed pipeline outperforms existing state-of-the-art methods in generating synthetic documents more effectively.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations