Papers with InfiniteBench

    2 papers
    Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference (2026.findings-acl)

    Copied to clipboard

    Challenge: Large language models (LLMs) have demonstrated remarkable capabilities in processing long contexts.
    Approach: They propose a training-free method that adaptively chooses the selection layer for KV cache reduction . they exploit the variance of token ranks ordered by attention score to optimize decoding .
    Outcome: The proposed method outperforms state-of-the-art token pruning methods on InfiniteBench, RULER, and NIAH benchmarks.
    LAVa: Layer-wise KV Cache Eviction with Dynamic Budget Allocation (2025.findings-emnlp)

    Copied to clipboard

    Challenge: Existing methods for cache compression are heuristic and lack dynamic budget allocation . cnn's john mccartney and johnny mccain present a new approach for cache eviction and dynamic budgets .
    Approach: They propose a unified framework for cache compression that minimizes information loss in transformer residual streams.
    Outcome: The proposed method consistently maintains top performance across task types.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations