| Challenge: | Transformer-based language models have a finite context window and expensive computational cost of processing long text documents. |
| Approach: | They propose to adapt pre-trained LMs into AutoCompressors to compress text into summary vectors . authors propose to use summary vector to speed up inference over long contexts based on a finite context window . |
| Outcome: | The proposed model can compress long contexts into summary vectors, which are accessible as soft prompts. |
Similar Papers
Pretraining Context Compressor for Large Language Models with Embedding-Based Memory (2025.acl-long)
Copied to clipboard
| Challenge: | Efficient processing of long contexts in large language models is essential for real-world applications such as retrieval-augmented generation and in-context learning. |
| Approach: | They propose a decoupled compressor-LLM framework that preserves contextual information within condensed embedding representations. |
| Outcome: | The proposed framework outperforms baseline models in three domains and across eight datasets while adapting to different downstream LLMs. |
Extending Context Window of Large Language Models via Semantic Compression (2024.findings-acl)
Copied to clipboard
| Challenge: | Existing models rely on a quadratic computation to generate long texts . current models impose limitations on the length of text inputs . |
| Approach: | They propose a semantic compression method that extends the context window of large language models . the method reduces the semantic redundancy of long inputs before passing them to the LLMs . |
| Outcome: | The proposed method extends the context window of large language models across tasks . it exhibits consistent fluency in text generation while reducing associated computational overhead. |
In-Context Former: Lightning-fast Compressing Context for Large Language Model (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods to reduce inference costs of transformer-based large language models entail quadratic complexity . et al., 2017): transformer-derived large language model performance is a major challenge. |
| Approach: | They propose a method that compresses long contexts into short soft prompts . they use the self-attention mechanism of the large model to extract and condense information . |
| Outcome: | The proposed method reduces compression costs by 68 to 112 times while achieving 90% of baseline performance. |
Context Compression for Auto-regressive Transformers with Sentinel Tokens (2023.emnlp-main)
Copied to clipboard
| Challenge: | Existing Transformer-based LLMs have limited performance due to complexity of attention module . key-value cache is the major memory footprint and inference latency problem . |
| Approach: | They propose a plug-and-play approach that incrementally compresses token activation into compact ones . they also profile the benefit of context compression on improving the system throughout . |
| Outcome: | The proposed approach reduces memory footprint and inference latency by compressing tokens into compact ones. |
KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches (2024.findings-emnlp)
Copied to clipboard
Jiayi Yuan, Hongyi Liu, Shaochen Zhong, Yu-Neng Chuang, Songchen Li, Guanchu Wang, Duy Le, Hongye Jin, Vipin Chaudhary, Zhaozhuo Xu, Zirui Liu, Xia Hu
| Challenge: | Long context capability is a crucial competency for large language models as it mitigates the human struggle to digest long-form texts. |
| Approach: | They propose to evaluate 10+ state-of-the-art approaches for long context-capable LLMs. |
| Outcome: | The proposed methods are compared against 10+ state-of-the-art approaches across seven categories of long context tasks. |
Dodo: Dynamic Contextual Compression for Decoder-only LMs (2024.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to NLP are sparsifying attention patterns or approximating the attention computation with kernel methods. |
| Approach: | They propose a method for dynamic contextual compression for decoder-only LMs. |
| Outcome: | The proposed method reduces the cost of self-attention to a fraction of typical time and space. |
Can Large Language Models Understand Context? (2024.findings-eacl)
Copied to clipboard
Yilun Zhu, Joel Moniz, Shruti Bhargava, Jiarui Lu, Dhivya Piraviperumal, Site Li, Yuan Zhang, Hong Yu, Bo-Hsiang Tseng
| Challenge: | Existing evaluation methodologies for Large Language Models (LLMs) have been inadequate to evaluate their ability to understand contextual features. |
| Approach: | They propose a benchmark to assess large language models' ability to understand context by adapting existing datasets to suit their evaluation. |
| Outcome: | The proposed model performs better under the in-context learning pretraining scenario than state-of-the-art models. |
500xCompressor: Generalized Prompt Compression for Large Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Prompt compression is important for large language models to increase inference speed, reduce computation cost, and improve user experience. |
| Approach: | They propose a method that compresses natural language contexts into a special token . they propose to reduce computations and memory costs by reducing the complexity . |
| Outcome: | The proposed method reduces computations and memory costs by 27-90% . it retains 70-74% and 77-84% of the LLM capabilities at high compression ratios . |
Perception Compressor: A Training-Free Prompt Compression Framework in Long Context Scenarios (2025.findings-naacl)
Copied to clipboard
| Challenge: | Long prompts contain redundant information and are sensitive to the position of key information in long context scenarios. |
| Approach: | They propose a training-free prompt compression framework that retains key information at token level while removing distracting tokens. |
| Outcome: | The proposed framework outperforms existing methods on long context benchmarks. |
LLoCO: Learning Long Contexts Offline (2024.emnlp-main)
Copied to clipboard
Sijun Tan, Xiuyu Li, Shishir G Patil, Ziyang Wu, Tianjun Zhang, Kurt Keutzer, Joseph Gonzalez, Raluca Popa
| Challenge: | Large language models are still unable to handle long contexts due to the quadratic computational and memory overhead of the self-attention mechanism and the substantial KV cache sizes during generation. |
| Approach: | They propose a method to learn contexts offline through context compression and in-domain parameter-efficient finetuning with LoRA. |
| Outcome: | The proposed model outperforms in-context learning while using 30 fewer tokens during inference and significantly reduces the cost of long document question answering. |