Optimized Text Embedding Models and Benchmarks for Amharic Passage Retrieval (2025.findings-acl)
Copied to clipboard
| Challenge: | Existing methods for multilingual and cross-lingual retrieval are lacking in low-resource, morphologically rich languages such as Amharic. |
| Approach: | They propose to train Amharic-specific dense retrieval models based on pre-trained Amharican BERT and RoBERTa backbones. |
| Outcome: | The proposed model achieves 17.6% improvement in MRR@10 and 9.86% gain in Recall@10 over the strongest multilingual baseline, Arctic Embed 2.0. |
Similar Papers
Large Language Models as Foundations for Next-Gen Dense Retrieval: A Comprehensive Empirical Assessment (2024.emnlp-main)
Copied to clipboard
| Challenge: | Pre-trained language models have limited generalization capabilities and performance challenges. |
| Approach: | They evaluate 15 different backbone LLMs and non-LLMs to evaluate their performance . larger models and extensive pre-training consistently enhance in-domain accuracy and data efficiency . |
| Outcome: | The results show that larger models and extensive pre-training enhance in-domain accuracy and data efficiency. |
Llama2Vec: Unsupervised Adaptation of Large Language Models for Dense Retrieval (2024.acl-long)
Copied to clipboard
| Challenge: | Dense retrieval requires discriminative embeddings to represent the semantic relationship between query and document. |
| Approach: | They propose an unsupervised approach that performs unsupervised adaptation of large language models for dense retrieval. |
| Outcome: | The proposed model improves on a variety of dense retrieval benchmarks and is available on github. |
Shall We Pretrain Autoregressive Language Models with Retrieval? A Comprehensive Study (2023.emnlp-main)
Copied to clipboard
Boxin Wang, Wei Ping, Peng Xu, Lawrence McAfee, Zihan Liu, Mohammad Shoeybi, Yi Dong, Oleksii Kuchaiev, Bo Li, Chaowei Xiao, Anima Anandkumar, Bryan Catanzaro
| Challenge: | a recent study shows that retrieval-augmented LMs can improve text generation quality and accuracy. |
| Approach: | They propose a model that reproduces RETRO parameters while retrieving a text corpus . they find RETRO outperforms GPT on text generation with less repetition . |
| Outcome: | The proposed model outperforms standard retrieval-augmented GPT and retrieval augmented GTP on text generation and accuracy tasks. |
SkMTEB: Slovak Massive Text Embedding Benchmark and Model Adaptation (2026.acl-long)
Copied to clipboard
| Challenge: | Slovak embeddings are core infrastructure for semantic search, retrieval-augmented generation (RAG), clustering, and classification. |
| Approach: | They propose a MTEB-style text embedding benchmark for Slovak, a low-resource West Slavic language . they use 31 datasets across 7 task types to evaluate the performance of the models . |
| Outcome: | The proposed model achieves competitive performance with proprietary APIs while remaining locally deployable for RAG . the model is based on 31 datasets across 7 task types and is 4 the depth of existing benchmark for Slovak . |
Unsupervised Corpus Aware Language Model Pre-training for Dense Passage Retrieval (2022.acl-long)
Copied to clipboard
| Challenge: | Recent research shows that fine-tuning dense retrievers to realize their capacity requires carefully designed fine-cuning techniques. |
| Approach: | They propose a pre-training architecture that learns to condense information into the dense vector through LM pre-training and a coCondenser architecture which adds an unsupervised corpus-level contrastive loss to warm up the passage embedding space. |
| Outcome: | The proposed architecture reduces the need for heavy data engineering and large batch training. |
UNKs Everywhere: Adapting Multilingual Language Models to New Scripts (2021.emnlp-main)
Copied to clipboard
| Challenge: | Massively multilingual language models offer state-of-the-art cross-lingual transfer performance on a range of NLP tasks, but there is a profound performance gap between resource-rich and resource-poor target languages. |
| Approach: | They propose a series of data-efficient methods that enable quick and effective adaptation of pretrained multilingual models to low-resource languages and unseen scripts. |
| Outcome: | The proposed methods improve learning of the new dedicated embedding matrix in the target language and for low-resource languages written in unseen scripts. |
Mini-Model Adaptation: Efficiently Extending Pretrained Models to New Languages via Aligned Shallow Training (2023.findings-acl)
Copied to clipboard
| Challenge: | Existing approaches to pretrain Masked Language Models (MLMs) are expensive and require a full forward and backward pass over the entire model. |
| Approach: | They propose to learn a shallow mini-model from a fraction of a large model's parameters and plug it into a larger model for rapid cross-lingual transfer. |
| Outcome: | Experiments on XNLI, MLQA and PAWS-X show that mini-model adaptation matches the standard approach using up to 2.3x less compute on average. |
Leveraging Cognitive Complexity of Texts for Contextualization in Dense Retrieval (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing approaches to estimate semantic similarity of queries and documents rely on token-level information derived from query/document interactions. |
| Approach: | They propose a new DRM that leverages query/document interactions based on full embedding representations generated by a Transformer-based model. |
| Outcome: | The proposed model outperforms fine-tuning techniques on lightweight bi-encoders and traditional late-interaction models. |
Parameter-Efficient Neural Reranking for Cross-Lingual and Multilingual Retrieval (2022.coling-1)
Copied to clipboard
| Challenge: | State-of-the-art neural rankers are notoriously data-hungry and rarely used in multilingual and cross-lingual retrieval settings. |
| Approach: | They propose to use Sparse Fine-Tuning Masks and Adapters to transfer rankers trained on English data to other languages and cross-lingual setups by means of multilingual encoders. |
| Outcome: | The proposed methods outperform standard zero-shot transfer with full MMT fine-tuning while being more modular and reducing training times. |
Dwell in the Beginning: How Language Models Embed Long Documents for Dense Retrieval (2024.acl-short)
Copied to clipboard
| Challenge: | Existing studies have shown that Transformer-based language models lose information in the middle of input sequences, especially in the context of web document retrieval. |
| Approach: | They examine position biases at multiple stages of the training pipeline for an encoder-decoder neural retrieval model, namely language model pre-training, contrastive pre- training, and contrastive fine-tuning. |
| Outcome: | The proposed model generates embeddings that better capture the beginning of the input content, with fine-tuning further aggravating this effect. |