Challenge: Contextual document embedding reranking is an efficient and efficient retrieval framework.
Approach: They propose a highly efficient retrieval framework that uses contextual document embedding reranking to incorporate ranking context into training.
Outcome: The proposed framework reduces the computational overhead of a first-stage method and can be used as stand-alone retrieval models.

Similar Papers

Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document Embeddings (2025.emnlp-main)

Copied to clipboard

Challenge: Modern document retrieval embedding methods typically encode passages (chunks) from documents independently, often overlooking contextual information from the rest of the document.
Approach: They propose a benchmark to evaluate retrieval models' ability to leverage document-wide context.
Outcome: The proposed method significantly improves retrieval quality on ConTEB without sacrificing base model performance.
Improving Embedding-based Large-scale Retrieval via Label Enhancement (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing methods for large-scale retrieval are trained with 0-1 hard labels that indicate whether a query is relevant to a document, ignoring rich information of the relevance degree.
Approach: They propose to introduce label enhancement for the first time to characterize query-document relevance degree by embedding label distribution into contextual embeddables.
Outcome: The proposed method significantly outperforms existing retrieval models and its counterparts equipped with two alternative methods on English and Chinese large-scale retrieval tasks.
Query-as-context Pre-training for Dense Passage Retrieval (2023.emnlp-main)

Copied to clipboard

Challenge: Existing methods to improve passage retrieval performance by using context-supervised pre-training are weakly correlated.
Approach: They propose to use query-as-context pre-training to train passage-query pairs . they evaluate the pre-trained models on large-scale passage retrieval benchmarks .
Outcome: The proposed technique improves performance on large-scale passage retrieval benchmarks and out-of-domain zero-shot benchmarks.
Leveraging Cognitive Complexity of Texts for Contextualization in Dense Retrieval (2025.emnlp-main)

Copied to clipboard

Challenge: Existing approaches to estimate semantic similarity of queries and documents rely on token-level information derived from query/document interactions.
Approach: They propose a new DRM that leverages query/document interactions based on full embedding representations generated by a Transformer-based model.
Outcome: The proposed model outperforms fine-tuning techniques on lightweight bi-encoders and traditional late-interaction models.
Dimension Reduction for Efficient Dense Retrieval via Conditional Autoencoder (2022.emnlp-main)

Copied to clipboard

Challenge: Existing work reserves the principle dimensions of query and document embeddings for building more efficient retrieval systems.
Approach: They propose to use Conditional Autoencoder to compress high-dimensional embeddings to maintain the same embeddable distribution and better recover ranking features.
Outcome: The proposed algorithm achieves comparable ranking performance with its teacher model and makes the retrieval system more efficient.
Dwell in the Beginning: How Language Models Embed Long Documents for Dense Retrieval (2024.acl-short)

Copied to clipboard

Challenge: Existing studies have shown that Transformer-based language models lose information in the middle of input sequences, especially in the context of web document retrieval.
Approach: They examine position biases at multiple stages of the training pipeline for an encoder-decoder neural retrieval model, namely language model pre-training, contrastive pre- training, and contrastive fine-tuning.
Outcome: The proposed model generates embeddings that better capture the beginning of the input content, with fine-tuning further aggravating this effect.
Efficient Document Retrieval by End-to-End Refining and Quantizing BERT Embedding with Contrastive Product Quantization (2022.emnlp-main)

Copied to clipboard

Challenge: Existing semantic hashing methods only learn a binary code for each document and use Hamming distance to evaluate document distances.
Approach: They propose to leverage BERT embeddings to perform efficient retrieval based on product quantization technique . they transform original BERT embedded codewords and feed it into a probabilistic product quantizer module .
Outcome: The proposed method outperforms current state-of-the-art methods on three benchmarks.
SACL: Understanding and Combating Textual Bias in Code Retrieval with Semantic-Augmented Reranking and Localization (2025.findings-emnlp)

Copied to clipboard

Challenge: a recent study shows that code retrievers exhibit a strong bias towards well-documented code .
Approach: They propose a framework that augments textual information with semantic information to mask specific features while preserving code functionality.
Outcome: The proposed framework enhances textual information and reduces bias by augmenting code or structural knowledge with semantic information.
SDR: Efficient Neural Re-ranking using Succinct Document Representation (2022.acl-long)

Copied to clipboard

Challenge: BERT based ranking models have been successful on various information retrieval tasks, but they are prone to storage and network fetching latency.
Approach: They propose a late-interaction architecture that allows pre-computation of intermediate document representations, thus reducing latency.
Outcome: The proposed model achieves 4x–11.6x higher compression rates on the MSMARCO passage re-reranking task compared to existing methods.
ED2LM: Encoder-Decoder to Language Model for Faster Document Re-ranking Inference (2022.findings-acl)

Copied to clipboard

Challenge: State-of-the-art neural models typically encode document-query pairs using cross-attention for re-ranking.
Approach: They propose to fine tune a pretrained encoder-decoder model using document to query generation.
Outcome: The proposed model achieves comparable results to more expensive approaches while being 6.8X faster.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations