Papers by Chengruidong Zhang

2 papers
LeanK: Learnable K Cache Channel Pruning for Efficient Decoding (2025.emnlp-main)

Copied to clipboard

Challenge: Existing efforts to optimize the key-value (KV) cache include: (1) Eviction, which discards cache of less important tokens; (2) Selection, which retains the full KV cache but selectively reads relevant entries.
Approach: They propose a learning-based method that prunes unimportant key (K) cache channels by leveraging static channel sparsity.
Outcome: Experiments show that LeanK reduces GPU memory and accelerates decoding without sacrificing accuracy.
Accelerating Prefilling via Decoding-time Contribution Sparsity (2026.findings-acl)

Copied to clipboard

Challenge: Existing acceleration methods exploit attention score sparsity by estimating blocks with high attention scores and applying dynamic sparse attention.
Approach: They propose a method which replaces dense attention with Triangle attention in a subset of layers to reduce the time needed to decode.
Outcome: Experiments show that TriangleMix achieves near-lossless performance on long-context and long-constrast reasoning benchmarks while significantly improving efficiency.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations