| Challenge: | Existing lexical search models suffer from lexica gap problems and are not fast enough to solve these problems. |
| Approach: | They build a dense retrieval based semantic search engine on scientific articles from Elsevier that generates high-quality pseudo training labels. |
| Outcome: | The proposed model significantly outperforms the currently deployed lexical search engine on the two test sets. |
Similar Papers
GPL: Generative Pseudo Labeling for Unsupervised Domain Adaptation of Dense Retrieval (2022.naacl-main)
Copied to clipboard
| Challenge: | Dense retrieval approaches suffer from the lexical gap and require large amounts of training data. |
| Approach: | They propose an unsupervised method for domain adaptation that uses query generator and pseudo labeling from a cross-encoder to improve retrieval performance. |
| Outcome: | The proposed method outperforms state-of-the-art retrieval methods on domain-specialized datasets by 9.3 points nDCG@10 on six tasks. |
Unsupervised Dense Retrieval with Relevance-Aware Contrastive Pre-Training (2023.findings-acl)
Copied to clipboard
| Challenge: | Dense retrievers have impressive performance, but their demand for abundant training data limits their application scenarios. |
| Approach: | They propose a method which uses unlabeled data to construct pseudo-positive examples from unlabelled data and then contrastively weighs the contrastive loss of different pairs according to the estimated relevance. |
| Outcome: | The proposed method beats the SOTA unsupervised Contriever model on BEIR and open-domain QA retrieval benchmarks and is a good few-shot learner. |
Unsupervised Multilingual Dense Retrieval via Generative Pseudo Labeling (2024.findings-eacl)
Copied to clipboard
| Challenge: | Existing sparse retrieval methods often yield inferior performance in multilingual retrieval, requiring a large amount of paired data, which is costly. |
| Approach: | They propose an Unsupervised Multilingual dense Retriever trained without paired data which iteratively improves performance of multilingual retrievers. |
| Outcome: | The proposed framework outperforms supervised baselines on two benchmark datasets and shows that iterative training improves the performance. |
PseudoSeer: a Search Engine for Pseudocode (2026.findings-acl)
Copied to clipboard
| Challenge: | PseudoSeer is a search engine for academic pseudocode that indexes over 320,000 implementations extracted from 2.2 million arXiv papers. |
| Approach: | They propose to use caption-reference pairs to match short queries with a median length of five words against long documents composed primarily of natural language with limited LaTeX notation. |
| Outcome: | The proposed algorithm outperforms the best pretrained model by 8.7 points and achieves 66.5% R@10 . |
How to Train Your Dragon: Diverse Augmentation Towards Generalizable Dense Retrieval (2023.findings-emnlp)
Copied to clipboard
Sheng-Chieh Lin, Akari Asai, Minghan Li, Barlas Oguz, Jimmy Lin, Yashar Mehdad, Wen-tau Yih, Xilun Chen
| Challenge: | Existing techniques to improve dense retrieval suffer from effectiveness tradeoffs between supervised and zero-shot retrieval, some argue due to the limited model capacity. |
| Approach: | They propose to use diverse queries and sources of supervision to train a generalizable DR to achieve high accuracy in both supervised and zero-shot retrieval. |
| Outcome: | The proposed DR can achieve state-of-the-art in supervised and zero-shot evaluations without increasing model size. |
MixGR: Enhancing Retriever Generalization for Scientific Domain through Complementary Granularity (2024.emnlp-main)
Copied to clipboard
| Challenge: | Recent studies show the importance of document retrieval in the scientific domain. |
| Approach: | They propose a zero-shot approach to measure query-document similarity using atomic components in queries and documents to combine them into a united score. |
| Outcome: | The proposed approach outperforms previous document retrieval methods by 24.7%, 9.8%, and 6.9% on nDCG@5 with unsupervised, supervised, and LLM-based retrievers. |
Breaking Boundaries in Retrieval Systems: Unsupervised Domain Adaptation with Denoise-Finetuning (2023.findings-emnlp)
Copied to clipboard
| Challenge: | Existing domain adaptation methods for dense retrieval models use unadapted rerank models, leading to imprecise labels. |
| Approach: | They propose to adapt a rerank model to the target domain before using it for label generation. |
| Outcome: | The proposed model achieves better results across three retrieval datasets. |
Domain Adaptation for Dense Retrieval and Conversational Dense Retrieval through Self-Supervision by Meticulous Pseudo-Relevance Labeling (2024.lrec-main)
Copied to clipboard
| Challenge: | Recent studies have shown that dense retrieval models generalize less well than interaction-based models on out-of-distribution data sets. |
| Approach: | They propose to combine query-generation approach with self-supervision approach in which pseudo-relevance labels are automatically generated on the target domain. |
| Outcome: | The proposed approach is based on a T5-3B model for pseudo-positive labeling and hard negatives on conversational dense retrieval models. |
Scientific Paper Retrieval with LLM-Guided Semantic-Based Ranking (2025.findings-emnlp)
Copied to clipboard
| Challenge: | Recent studies also use large language models (LLMs) for query understanding, but these methods lack grounding in corpus-specific knowledge and may generate unreliable or unfaithful content. |
| Approach: | They propose a paper retrieval framework that combines large language models (LLMs) with a concept-based semantic index to capture scientific concepts. |
| Outcome: | The proposed framework improves the performance of various base retrievers, surpasses strong existing LLM-based baselines, and remains highly efficient. |
Lexicon-Enhanced Self-Supervised Training for Multilingual Dense Retrieval (2022.findings-emnlp)
Copied to clipboard
| Challenge: | Recent multilingual pre-trained models perform poorly on multilingual retrieval tasks due to lack of multilingual training data. |
| Approach: | They propose to mine and generate self-supervised training data based on large-scale unlabeled corpus and introduce query generator to generate more queries in target languages for unlabed passages. |
| Outcome: | The proposed method performs better than baselines on a Mr. TYDI dataset and an industrial dataset from a commercial search engine. |