Challenge: InfiniPot is a KV cache control framework that can handle long input contexts without additional training.
Approach: They propose a KV cache control framework that can handle long input contexts efficiently without additional training.
Outcome: The proposed framework outperforms models trained for long contexts in various NLP tasks and is highly efficient and versatile.

Similar Papers

KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches (2024.findings-emnlp)

Copied to clipboard

Challenge: Long context capability is a crucial competency for large language models as it mitigates the human struggle to digest long-form texts.
Approach: They propose to evaluate 10+ state-of-the-art approaches for long context-capable LLMs.
Outcome: The proposed methods are compared against 10+ state-of-the-art approaches across seven categories of long context tasks.
EdgeInfinite: A Memory-Efficient Infinite-Context Transformer for Edge Devices (2025.acl-industry)

Copied to clipboard

Challenge: Existing KV cache optimizations struggle with irreversible token eviction in long-output tasks . alternative sequence modeling architectures prove costly to adopt within established Transformer infrastructures.
Approach: They propose a memory-efficient solution for infinite contexts that integrates compressed memory into Transformer-based LLMs through a trainable memory-gating module.
Outcome: The proposed solution achieves comparable performance to baseline Transformer-based LLMs while optimizing memory consumption and time to first token.
SirLLM: Streaming Infinite Retentive LLM (2024.acl-long)

Copied to clipboard

Challenge: Large Language Models (LLMs) are becoming increasingly prevalent in various domains, requiring a one-off input of overly long texts to maintain a degree of memory.
Approach: They propose a Streaming Infinite Retentive LLM which allows LLMs to maintain longer memory during infinite-length dialogues without fine-tuning.
Outcome: The proposed model can achieve stable and significant improvements across different LLMs and tasks, compellingly proving its effectiveness.
LM-Infinite: Zero-Shot Extreme Length Generalization for Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Currently, large language models (LLMs) train on short text segments due to the computational overhead quadratic in the input lengths of their Transformer architectures.
Approach: They propose a method that allows LLMs pre-trained with 2K or 4K-long segments to generalize to up to 200M length inputs while retaining perplexity.
Outcome: The proposed method achieves 2.7 decoding speed up and 7.5 memory saving over the original model.
Memorize Step by Step: Efficient Long-Context Prefilling with Incremental Memory and Decremental Chunk (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods to optimize LLM for long sequences for long documents are slow and consume memory.
Approach: They propose a method that starts with a small memory size and gradually increases it . they propose Decremental Chunk based on Incremental Memory (IMDC) which reduces chunk size while increasing memory size .
Outcome: The proposed method is faster (1.45x) and reduces GPU memory consumption by 23.3% compared to fixed-size memory.
Efficient Solutions For An Intriguing Failure of LLMs: Long Context Window Does Not Mean LLMs Can Analyze Long Sequences Flawlessly (2025.coling-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable capabilities in comprehending and analyzing lengthy sequential inputs.
Approach: They propose to implement ad-hoc solutions that enhance LLMs’ performance on long input sequences by up to 50% while reducing API cost and latency by up . to address this limitation, they propose to use three datasets and two tasks to analyze news categorization and sentence analysis to evaluate their models.
Outcome: The proposed solutions significantly improve LLMs’ performance on long input sequences by up to 50% while reducing API cost and latency by up . to 93% and 50%, respectively.
Discovering the Gems in Early Layers: Accelerating Long-Context LLMs with 1000x Input Token Reduction (2026.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable capabilities in handling long context inputs, but this comes at the cost of increased computational resources and latency.
Approach: They propose an algorithm that uses early LLM layers as filters to select and compress input tokens, reducing the context length for subsequent processing.
Outcome: The proposed method outperforms existing techniques on the Needle in a Haystack task while demonstrating comparable performance on the LongBench challenge.
InfiniteICL: Breaking the Limit of Context Window Size via Long Short-term Memory Transformation (2025.findings-acl)

Copied to clipboard

Challenge: InfiniteICL is a framework that parallels context and parameters in large language models with short- and long-term memory in human cognitive systems.
Approach: They propose a framework that parallels context and parameters in large language models with short- and long-term memory in human cognitive systems and enables infinite context integration.
Outcome: The proposed framework reduces context length by 90% while achieving 103% average performance of full-context prompting across fact recall, grounded reasoning, and skill acquisition tasks.
KV Pareto: Systems-Level Optimization of KV Cache and Model Compression for Long Context Inference (2026.eacl-industry)

Copied to clipboard

Challenge: Long-context Large Language Models (LLMs) face significant memory bottlenecks due to the linear growth of key-value (KV) cache with sequence length.
Approach: They propose a framework that maps the trade-off frontier between total memory consumption and task accuracy across three complementary optimization techniques.
Outcome: The proposed model-specific configurations achieve 68-78% total memory reduction with minimal (1-3%) accuracy degradation on long-context tasks.
LongReD: Mitigating Short-Text Degradation of Long-Context Large Language Models via Restoration Distillation (2025.acl-long)

Copied to clipboard

Challenge: Large language models (LLMs) have extended context windows through scaling positional encodings and lightweight continual pre-training, but performance degradation is still not fully explored.
Approach: They propose a novel approach to reduce short-text performance degradation by minimizing distribution drift in hidden states and attention scores.
Outcome: The proposed approach minimizes the distribution discrepancy between the extended and original models while maintaining or even enhancing the model's long-context abilities.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations