Challenge: Adaptive Retention is a probabilistic, layer-wise token selection mechanism that learns which representations to keep under a strict global budget M.
Approach: They propose a probabilistic token selection mechanism that learns which representations to keep under a strict global budget M.
Outcome: The proposed method reduces memory usage by 35–45% while improving throughput by 1.8.

Similar Papers

LAMB: A Training-Free Method to Enhance the Long-Context Understanding of SSMs via Attention-Guided Token Filtering (2025.acl-short)

Copied to clipboard

Challenge: Recent work attributes performance degradation to an exponential decay in hidden-state memory.
Approach: They propose a token filtering strategy that is training-free and attention-guided . they propose 'LAMB' to preserve critical tokens during inference .
Outcome: The proposed token filtering improves long-context performance by 30.35% over state-of-the-art methods on benchmarks.
Efficient Sparse Attention needs Adaptive Token Release (2024.findings-acl)

Copied to clipboard

Challenge: Large Language Models (LLMs) have demonstrated remarkable capabilities across a wide array of text-centric tasks, however, their ‘large’ scale introduces significant computational and storage challenges, particularly in managing the key-value states of the transformer, which limits their wider applicability.
Approach: They propose to release resources from caches and rebuild key-value states by a lightweight controller module to approximate an ideal top-K sparse attention.
Outcome: The proposed method achieves a significant throughput improvement of 221.8% over full attention and a model with 7 billion tokens.
Long-Range Language Modeling with Selective Cache (2023.findings-emnlp)

Copied to clipboard

Challenge: Existing models that use transformers to model language cost quadratically increase with sequence length.
Approach: They propose a selective cache which stores key-value pairs from previous contexts.
Outcome: The proposed selective cache outperforms XL cache and compressive cache by considerable margins.
Flashback: Memory Mechanism for Enhancing Memory Efficiency and Speed in Deep Sequential Models (2025.coling-main)

Copied to clipboard

Challenge: Existing deep sequential processing models have problems with memory degradation and inaccurate gradient backpropagation.
Approach: They propose a Flashback property that preserves memory as an identity mapping until it is overwritten by a hidden state at a different time step.
Outcome: The proposed model can be implemented in Transformers and Mamba, and it performs well.
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection (2025.emnlp-main)

Copied to clipboard

Challenge: Rapid advances in Large Language Models have spurred demand for processing extended context sequences . however, performance degradation due to sequence lengths out-of-distribution and excessively long inference times are limiting LLMs in long-context scenarios.
Approach: They propose a training-free method for efficient and accurate long-context inference . they selectively involves a few critical KV cache tokens in attention calculation .
Outcome: The proposed method speeds up attention computation and accelerates inference time while reducing selection overhead.
RECALL: REpresentation-aligned Catastrophic-forgetting ALLeviation via Hierarchical Model Merging (2025.emnlp-main)

Copied to clipboard

Challenge: Existing models that require task labels or performance trade-offs are susceptible to catastrophic forgetting.
Approach: They propose a representation-aware model merging framework for continual learning without access to historical data.
Outcome: The proposed framework outperforms baselines in knowledge retention and generalization across five NLP tasks and multiple continual learning scenarios.
ABC: Attention with Bounded-memory Control (2022.acl-long)

Copied to clipboard

Challenge: Existing approaches to attention with bounded-memory control (ABC) have a quadratic complexity in sequence lengths, making it prohibitive for long sequences.
Approach: They propose a new abstraction that bounds memory size to improve efficiency . they propose bounded-memory control, which connects several efficient attention variants .
Outcome: The proposed approach outperforms existing approaches on language modeling, machine translation, and masked language model finetuning.
DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning (2026.findings-acl)

Copied to clipboard

Challenge: Large reasoning models generate long chains of intermediate steps, but their inference cost is dominated by decoding, where each new token must attend to the entire growing sequence.
Approach: They propose a training-free sparse attention mechanism that reduces inference cost by evicting entries from the key-value cache.
Outcome: The proposed model matches or surpasses full attention on reasoning benchmarks . it reduces the number of attended tokens by up to 4.25 and delivers 1.54 speedup .
Forgetting Curve: A Reliable Method for Evaluating Memorization Capability for Long-Context Models (2024.emnlp-main)

Copied to clipboard

Challenge: Existing methods, tasks and benchmarks to measure model’s effective memory length are limited.
Approach: They propose a method called forgetting curve to measure the memorization capability of long-context models.
Outcome: The proposed method is robust to the tested corpus and experimental settings, and can be applied to any model size.
Not Every Token Needs Forgetting: Selective Unlearning Balancing Forgetting and Utility in Large Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: Conventional unlearning approaches forget all tokens in a target document, including common tokens that carry general knowledge.
Approach: They propose a method that identifies a critical subset of tokens within the forgetting set that is relevant to the unwanted information and unlearns only those tokens.
Outcome: Experiments on two benchmarks and six baseline unlearning algorithms show that selective unlearning achieves effective unlearning on the targeted forget data.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations