Papers by Michele Marzollo
SSSD: Simply-Scalable Speculative Decoding (2026.acl-long)
Copied to clipboard
Michele Marzollo, Jiawei Zhuang, Niklas Roemer, Niklas Zwingenberger, Lorenz K Muller, Lukas Cavigelli
| Challenge: | Existing methods for accelerating inference in Large Language Models require additional training and training, resulting in a higher deployment and maintenance cost. |
| Approach: | They propose a training-free method that combines lightweight n-gram matching with hardware-aware speculation. |
| Outcome: | SSSD reduces latency by up to 2.9 and is faster than autoregressive decoding methods. |