Challenge: Speculative decoding (SD) is a promising technique for LLM inference acceleration.
Approach: They propose a method to generate draft tokens in a retrieval-based manner to reduce drafting overhead and improve inference speed.
Outcome: Extensive tests show that *LogitSpec* can achieve 2.61 speedup and 3.28 mean accepted tokens per decoding step.

Similar Papers

DReSD: Dense Retrieval for Speculative Decoding (2025.findings-acl)

Copied to clipboard

Challenge: Speculative decoding (SD) uses an efficient draft model to propose the next few tokens, which are verified by the LLM in a single forward call, reducing latency while preserving its outputs.
Approach: They propose a draft model that proposes the next few tokens from a non-parametric datastore and uses a framework that uses approximate nearest neighbour search with contextualised token embeddings to retrieve the most semantically relevant sequences for SD.
Outcome: The proposed framework achieves (on average) 87% higher acceptance rates, 65% longer accepted tokens and 19% faster generation speeds compared to sparse retrieval (REST).
Double: Breaking the Acceleration Limit via Double Retrieval Speculative Parallelism (2026.acl-long)

Copied to clipboard

Challenge: Parallel Speculative Decoding (PSD) has limitations due to speedup limits and high computational waste . a novel synchronous mechanism solves the Retrieval Precision-Efficiency Dilemma .
Approach: They propose a framework that combines a draft-verification-based approach with a synchronous mechanism to solve the Retrieval Precision-Efficiency Dilemma.
Outcome: The proposed framework breaks speedup limits for Speculative Decoding by overlapping draft generation with verification.
RACER: Retrieval-Augmented Contextual Rapid Speculative Decoding (2026.findings-acl)

Copied to clipboard

Challenge: Existing methods for decoding large language models generate one token per step, causing high inference latency.
Approach: They propose a method that integrates retrieved exact patterns with logit-driven future cues.
Outcome: Experiments on Spec-Bench, HumanEval, and MGSM-ZH show that RACER outperforms training-free methods and accelerates inference.
Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding (2024.acl-long)

Copied to clipboard

Challenge: Existing methods for accelerating Large Language Models have been criticized for their inference costs and inefficient decoding.
Approach: They propose a self-speculative decoding approach for accelerating Large Language Models without an auxiliary model.
Outcome: The proposed method achieves a speedup of up to 1.99 with no additional neural network training and no extra memory footprint.
Speculative Streaming: Efficient and Scalable Speculative Decoding with Multi-Stream Attention (2025.emnlp-main)

Copied to clipboard

Challenge: Speculative decoding is a prominent technique for accelerating LLM inference by leveraging an auxiliary draft model, but its effectiveness is limited by the autoregressive nature of draft generation.
Approach: They propose a method that integrates speculative draft generation directly within the target model using multi-stream attention.
Outcome: The proposed method improves acceptance but also latency and speculation latency, limiting overall speedup.
UniSpec: Training-Free Speculative Decoding for Robust LLM Acceleration Across Languages and Hardware (2026.acl-long)

Copied to clipboard

Challenge: Existing methods for speculative decoding ignore device-specific verification costs and lack of mechanisms to assess draft token quality.
Approach: They propose a training-free, lossless speculative decoding framework that enables robust, plug-and-play LLM acceleration across diverse hardware configurations and languages.
Outcome: The proposed framework outperforms existing training-free methods while maintaining identical output quality across different hardware environments.
RASD: Retrieval-Augmented Speculative Decoding (2025.findings-acl)

Copied to clipboard

Challenge: Existing methods for generating draft tokens rely on lightweight draft models or additional model structures to generate tokens and retrieve context from databases.
Approach: They propose to use a pruning method to enhance model-based speculative decoding by combining the best-fit model with the best retrieval tree.
Outcome: The proposed method achieves state-of-the-art inference acceleration across tasks such as DocQA, Summary, Code, and In-Domain QA.
SpecHub: Provable Acceleration to Multi-Draft Speculative Decoding (2024.emnlp-main)

Copied to clipboard

Challenge: Large Language Models (LLMs) have limited inference speed due to sequential token generation . Spechub is a novel, efficient sampling-verification method for MDSD that improves acceptance rates with only linear computational overhead.
Approach: They propose a method that uses a smaller draft model to generate multiple token sequences . Spechub generates 0.05-0.27 and 0.02-0.16 more tokens per step than RRS and RRS without replacement .
Outcome: The proposed method improves acceptance rates with only linear computational overhead.
Speculative Decoding via Early-exiting for Faster LLM Inference with Thompson Sampling Control Mechanism (2024.findings-acl)

Copied to clipboard

Challenge: Existing approaches to generate draft tokens in large language models are expensive and resource-intensive.
Approach: They propose an approach to generate draft tokens using a segment of the LLM and a self-distillation method to enhance the quality of draft token.
Outcome: The proposed approach generates draft tokens using a segment of the LLM and a self-distillation method to improve quality and speed up generation.
DiffuSpec: Unlocking Diffusion Language Models for Speculative Decoding (2026.findings-acl)

Copied to clipboard

Challenge: Autoregressive (AR) decoding in large language models is latency-bounded by strictly sequential token generation.
Approach: They propose a diffusion-based drafter that proposes multi-token candidates and then verifies them in parallel by the target model.
Outcome: The proposed drafter generates multi-token proposals in a single forward pass while remaining compatible with standard AR verifiers.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations