Speculative Sampling and Draft Models: Why Faster Tokens Need a Verifier?
Speculative Sampling and Draft Models: Why Faster Tokens Need a Verifier? In the performance optimization of LLM inference, a core contradiction exists: gene…

Speculative Sampling and Draft Models: Why Faster Tokens Need a Verifier?
In the performance optimization of LLM inference, a core contradiction exists: generation speed is constrained by the autoregressive nature of the models. The model must generate tokens one by one, and each token requires a full forward pass through the model. Even with state-of-the-art GPU clusters, noticeable latency persists when generating extremely long texts.
To break this bottleneck, speculative decoding has emerged. Its core logic shifts from "slow but steady" to "fast and bold."
What is Speculative Sampling?
Simply put, speculative sampling introduces a lightweight "draft model." This model is very small (e.g., 100M parameters), extremely fast in inference, but has lower accuracy.
The workflow is as follows:
- Drafting Phase: The draft model quickly predicts the next $K$ tokens (e.g., 5). Because it is small, these $K$ tokens are generated much faster than by the main model.
- Verification Phase: These $K$ tokens are submitted to the main model (Verifier Model) for parallel verification in one go. Leveraging its superior capabilities, the main model computes the probability distributions for these $K$ positions through a single forward pass.
- Acceptance and Correction: If the main model deems the draft model's predictions probabilistically acceptable, it accepts them directly. If a deviation occurs at any position, that token and all subsequent ones are discarded, and the main model provides the correct correction.
This mechanism transforms the originally serial generation process into a "predict-verify" loop. Ideally, if the draft model has a high hit rate, a single iteration can produce multiple tokens, significantly boosting end-to-end throughput.
Why Do We Need a Verifier?
Many ask: Since the draft model runs so fast, why not use it directly? The answer lies in the LLM's "hallucinations" and "logical consistency."
Small models excel at handling simple grammar and common phrases but are prone to failure when dealing with complex logic, specialized knowledge, or long-range dependencies. Without verification by the main model, the generated text would quickly lose its logical foundation. The presence of the verifier ensures that: inference speed is determined by the small model, but output quality is backed by the large model.
Key Challenge: Balancing Hit Rate and Overhead
The efficiency of speculative sampling depends on the $\text{Acceptance Rate}$.
- If the acceptance rate is very high $\rightarrow$ speed increases significantly.
- If the acceptance rate is very low $\rightarrow$ the main model must not only verify incorrect tokens but also regenerate the correct ones, thereby increasing computational overhead.
Therefore, it is crucial to select a draft model whose distribution closely matches that of the main model while being sufficiently small. The current trend is to use distillation techniques to transfer knowledge from large models to small ones, thereby improving the hit rate.
Practical Implications
For developers and architects, speculative sampling means that "quantization" is no longer the only path for deploying LLMs. By introducing auxiliary small models or simple N-gram-based predictors, a 2-3x inference speedup can be achieved without sacrificing accuracy.
Core Conclusion: Future AI inference will no longer be a solo performance by a single model, but a collaborative system of "fast prediction + precise verification."