Papers by Chengruidong Zhang
LeanK: Learnable K Cache Channel Pruning for Efficient Decoding (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing efforts to optimize the key-value (KV) cache include: (1) Eviction, which discards cache of less important tokens; (2) Selection, which retains the full KV cache but selectively reads relevant entries. |
| Approach: | They propose a learning-based method that prunes unimportant key (K) cache channels by leveraging static channel sparsity. |
| Outcome: | Experiments show that LeanK reduces GPU memory and accelerates decoding without sacrificing accuracy. |
Accelerating Prefilling via Decoding-time Contribution Sparsity (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing acceleration methods exploit attention score sparsity by estimating blocks with high attention scores and applying dynamic sparse attention. |
| Approach: | They propose a method which replaces dense attention with Triangle attention in a subset of layers to reduce the time needed to decode. |
| Outcome: | Experiments show that TriangleMix achieves near-lossless performance on long-context and long-constrast reasoning benchmarks while significantly improving efficiency. |