Papers by Zhiwen Mo
Deep Kernel Fusion for Transformers (2026.acl-short)
Copied to clipboard
| Challenge: | Agentic LLM inference with long contexts is limited by memory bandwidth rather than compute. |
| Approach: | They propose a deeply fused kernel that cuts HBM traffic and boosts cache reuse. |
| Outcome: | The proposed kernel delivers 13.2% speedup on H100 and 9.7% on A100 over SGLang. |