Towards Fast Multilingual LLM Inference: Speculative Decoding and Specialized Drafters (2024.emnlp-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) have revolutionized natural language processing and are limited by high inference time in multilingual settings. |
| Approach: | They propose a training recipe for an assistant model in speculative decoding, which are leveraged to draft and-then its future tokens are verified by the target LLM. |
| Outcome: | The proposed model significantly speeds up inference time and out-of-domain speedup across various languages. |
Similar Papers
Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding (2024.findings-acl)
Copied to clipboard
Heming Xia, Zhe Yang, Qingxiu Dong, Peiyi Wang, Yongqi Li, Tao Ge, Tianyu Liu, Wenjie Li, Zhifang Sui
| Challenge: | Large Language Models (LLMs) have a high inference latency stemming from autoregressive decoding. |
| Approach: | They propose a novel decoding paradigm that drafts multiple tokens and verifies them in parallel . they aim to provide a catalyst for further research on Speculative Decoding . |
| Outcome: | The proposed method drafts multiple tokens and verifies them in parallel . it can be used to accelerate inference in large language models. |
Draft
& Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding (2024.acl-long)
Copied to clipboard
| Challenge: | Existing methods for accelerating Large Language Models have been criticized for their inference costs and inefficient decoding. |
| Approach: | They propose a self-speculative decoding approach for accelerating Large Language Models without an auxiliary model. |
| Outcome: | The proposed method achieves a speedup of up to 1.99 with no additional neural network training and no extra memory footprint. |
A Drop-In Solution for On-the-Fly Adaptation of Speculative Decoding in Large Language Models (2025.acl-long)
Copied to clipboard
| Challenge: | Large Language Models (LLMs) are highly memory-intensive when performing real-time inference. |
| Approach: | They propose a technique that allows for speculative decoding to be run on the fly to maximize the efficiency of LLM inferences. |
| Outcome: | The proposed solution can lead to 3.55-16.48% speed improvement over the standard speculative decoding, and 1.2-3.4 over the default LLMs. |
Decoding Speculative Decoding (2025.naacl-long)
Copied to clipboard
| Challenge: | Speculative decoding is a widely used technique to speed up inference for Large Language Models (LLMs) Autoregressive decoding has been known to be hardware inefficient, leading to poor resource utilization and low throughput during inference. |
| Approach: | They propose to use a draft model to generate speculative tokens and then use the target LLM to verify those tokens. |
| Outcome: | The proposed model can provide 111% higher throughput than existing draft models and generalizes further to all LLaMA models and supervised fine-tuned models. |
SPECTRA: Faster Large Language Model Inference with Optimized Internal and External Speculation (2025.acl-long)
Copied to clipboard
| Challenge: | Existing approaches to inference with Large Language Models (LLMs) are expensive and time-consuming. |
| Approach: | They propose a framework for accelerating large language model inference without additional training or modification to the original LLM. |
| Outcome: | The proposed framework outperforms state-of-the-art methods and achieves 4.08x speedups across benchmarks and LLM architectures. |
UniSpec: Training-Free Speculative Decoding for Robust LLM Acceleration Across Languages and Hardware (2026.acl-long)
Copied to clipboard
| Challenge: | Existing methods for speculative decoding ignore device-specific verification costs and lack of mechanisms to assess draft token quality. |
| Approach: | They propose a training-free, lossless speculative decoding framework that enables robust, plug-and-play LLM acceleration across diverse hardware configurations and languages. |
| Outcome: | The proposed framework outperforms existing training-free methods while maintaining identical output quality across different hardware environments. |
How Speculative Can Speculative Decoding Be? (2024.lrec-main)
Copied to clipboard
| Challenge: | Large language models (LLMs) have a largely increased latency due to their ability to autoregressively model . speculative decoding is a technique that trades generation quality for speed . |
| Approach: | They propose to use a draft model to draft tokens autoregressively and then verify them in parallel. |
| Outcome: | The proposed model could draft tokens autoregressively and then verify them in parallel . the proposed model trades quality for speed and could fail in verification stage . |
An Empirical Study of Speculative Decoding for Small Language Models (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing studies focus on 7B-70B parameters models, leaving a knowledge gap for small language models. |
| Approach: | They propose a draft-then-verify paradigm that allows for a single forward pass through a model and transfer of all model parameters to the GPU cache. |
| Outcome: | The proposed method can be used to accelerate small language models with low computational overhead. |
FastDraft: How to Train Your Draft (2025.findings-acl)
Copied to clipboard
| Challenge: | Speculative Decoding relies on the availability of efficient draft models, which are often lacking due to a stringent constraint of vocabulary compatibility. |
| Approach: | They propose a novel approach for pre-training and aligning a draft model to any large language model by incorporating efficient pre-train and fine-tuning over synthetic datasets generated by the target model. |
| Outcome: | The proposed model can be trained on a single server with 8 Intel Gaudi 2 accelerators in under 24 hours and achieves 3x acceptance rate, block efficiency and 2x memory bound speedup. |
Speculative Diffusion Decoding: Accelerating Language Generation through Diffusion (2025.naacl-long)
Copied to clipboard
Jacob K Christopher, Brian R. Bartoldson, Tal Ben-Nun, Michael Cardei, Bhavya Kailkhura, Ferdinando Fioretto
| Challenge: | Existing methods to accelerate large language model inference are limited by the reliance on incremental token generation in existing draft models. |
| Approach: | They propose an adaptation of speculative decoding which uses discrete diffusion models to generate draft sequences and allows parallelization of both the drafting and verification steps. |
| Outcome: | The proposed approach provides 7.2x speedups over standard generation processes and 1.75x speed ups over existing speculative decoding approaches. |