Landmark Embedding: A Chunking-Free Embedding Method For Retrieval Augmented Long-Context Large Language Models (2024.acl-long)
Copied to clipboard
| Challenge: | Existing methods for retrieval augmentation work with chunked contexts, which leads to poor quality of semantic representation and incomplete retrieval of useful information. |
| Approach: | They propose a method for retrieval augmentation of long-context language modeling using landmark embedding. |
| Outcome: | The proposed method outperforms existing retrieval methods with a notable advantage. |
Similar Papers
Situated Embedding Models for Context-Aware Dense Retrieval (2026.acl-short)
Copied to clipboard
| Challenge: | Existing embedding models are not well-equipped to encode situated context effectively, i.e., situating a chunk’s meaning within its context. |
| Approach: | They propose to represent short chunks in a way that is conditioned on a broader context window to enhance retrieval performance. |
| Outcome: | The proposed model outperforms state-of-the-art embedding models on a book-plot retrieval dataset. |
LongEmbed: Extending Embedding Models for Long Context Retrieval (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing embedding models support only 512 input tokens, hindering their application in scenarios requiring long inputs. |
| Approach: | They evaluate the performance of existing embedding models by using a new benchmark and a training-free context window extension strategy. |
| Outcome: | The proposed model extends the input window of existing models by several folds. |
Tagging-Augmented Generation: Assisting Language Models in Finding Intricate Knowledge In Long Contexts (2025.emnlp-industry)
Copied to clipboard
Anwesan Pal, Karen Hovsepian, Tinghao Guo, Mengnan Zhao, Somendra Tripathi, Nikos Kanakaris, George Mihaila, Sumit Nigam
| Challenge: | Recent studies into effective context lengths of flagship large language models (LLMs) have revealed major limitations in effective question answering (QA) and reasoning over long and complex contexts for even the largest and most impressive cadre of models. |
| Approach: | They propose a lightweight data augmentation strategy that boosts LLM performance in long-context scenarios without degrading and altering the integrity and composition of retrieved documents. |
| Outcome: | The proposed strategy boosts performance in long-context scenarios without degrading and altering the integrity and composition of retrieved documents. |
Context is Gold to find the Gold Passage: Evaluating and Training Contextual Document Embeddings (2025.emnlp-main)
Copied to clipboard
| Challenge: | Modern document retrieval embedding methods typically encode passages (chunks) from documents independently, often overlooking contextual information from the rest of the document. |
| Approach: | They propose a benchmark to evaluate retrieval models' ability to leverage document-wide context. |
| Outcome: | The proposed method significantly improves retrieval quality on ConTEB without sacrificing base model performance. |
Extending LLM Context Window with Adaptive Grouped Positional Encoding: A Training-Free Method (2025.acl-long)
Copied to clipboard
| Challenge: | Existing long-context training data is scarce and requires substantial GPU resources for training. |
| Approach: | They propose a training-free plug-and-play method to enhance long-context understanding in existing large language models. |
| Outcome: | The proposed method outperforms existing LLMs on various tasks and surpasses baseline methods. |
Untie the Knots: An Efficient Data Augmentation Strategy for Long-Context Pre-Training in Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Large language models (LLMs) have prioritized expanding the context window from which they can incorporate more information. |
| Approach: | They propose a data augmentation strategy to enable large language models to gain long-context capabilities without the need to modify existing data mixture. |
| Outcome: | The proposed model outperforms existing models on 20 billion tokens and achieves 75% and 84.5% accuracy on RULER at 128K context length. |
Document Segmentation Matters for Retrieval-Augmented Generation (2025.findings-acl)
Copied to clipboard
Zhitong Wang, Cheng Gao, Chaojun Xiao, Yufei Huang, Shuzheng Si, Kangyang Luo, Yuzhuo Bai, Wenhao Li, Tangjian Duan, Chuancheng Lv, Guoshan Lu, Gang Chen, Fanchao Qi, Maosong Sun
| Challenge: | Existing rule-based chunking methods lead to suboptimal splits, where overly large chunks introduce irrelevant information and small chunks lack semantic coherence. |
| Approach: | They propose a method that leverages document summaries as pseudo-instructions to guide chunking by computing semantic similarity between sentences and the summary. |
| Outcome: | Experiments on multiple open-domain question-answering benchmarks show that PIC significantly improves retrieval accuracy (Hits@k) and end-to-end QA performance (Exact Match) without any additional training. |
MoC: Mixtures of Text Chunking Learners for Retrieval-Augmented Generation System (2025.acl-long)
Copied to clipboard
| Challenge: | Existing methods for text chunking are limited by text chunks and lack of domain-specific knowledge. |
| Approach: | They propose a dual-metric evaluation method to quantify text chunking quality . they aim to generate a structured list of chunking regular expressions . |
| Outcome: | The proposed method enables direct quantification of chunking quality . it substantiates the need to integrate LLMs into chunking process . |
Pretraining Context Compressor for Large Language Models with Embedding-Based Memory (2025.acl-long)
Copied to clipboard
| Challenge: | Efficient processing of long contexts in large language models is essential for real-world applications such as retrieval-augmented generation and in-context learning. |
| Approach: | They propose a decoupled compressor-LLM framework that preserves contextual information within condensed embedding representations. |
| Outcome: | The proposed framework outperforms baseline models in three domains and across eight datasets while adapting to different downstream LLMs. |
Stronger Baselines for Retrieval-Augmented Generation with Long-Context Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing long-context language models (LMs) can handle tens of thousands of tokens in a single context window. |
| Approach: | They compare two recent multi-stage pipelines, ReadAgent and RAPTOR, against three baselines. |
| Outcome: | The proposed pipelines outperform more complex methods on multiple long-context QA benchmarks. |