# Is Speculative Decoding Worth It? Profiling vLLM on NVIDIA Blackwell

Sheilah Kirui, Akamai | AI Engineer World's Fair 2026 | 15:17

Source: https://www.youtube.com/watch?v=XTpyNrEgJQ4
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/is-speculative-decoding-worth-it-profiling-vllm-on-nvidia-blackwell
Published: 2026-10-06
Tags: benchmarks, gpus, inference, latency

## TL;DR
- Speculative decoding uses a smaller model to propose several tokens, then has the larger model verify them in one forward pass.
- It needs extra GPU memory for a second model and its KV cache, so it is most suitable when there is spare VRAM and concurrency is modest.
- Structured output reached 1.6x faster generation in the demo, while creative writing had a low acceptance rate and gained much less.

## Summary
Sheilah Kirui explains how speculative decoding changes the decode phase of LLM inference. A smaller draft model proposes several tokens, and the larger target model accepts or rejects them in one forward pass. This can reduce the number of expensive sequential steps, but it requires memory for both models and both KV caches. Kirui profiles the approach with vLLM on one NVIDIA Blackwell GPU, comparing a baseline with speculative decoding. Structured output produced a high acceptance rate and ran 1.6x faster in the demo. Creative writing produced a low acceptance rate because the next token was less predictable, especially at higher temperature. She also explains why speculative decoding does not speed up prompt processing, so long-context workloads may see limited gains. Her checklist covers model size, tokenizer compatibility, accuracy, cost, available VRAM, batch size, concurrency, and the structure of the workload.

## Key ideas
### Inference has a one-time prompt phase and a sequential generation phase
[00:44](https://www.youtube.com/watch?v=XTpyNrEgJQ4&t=44s)
Kirui separates inference into prefill and decode. During prefill, the model reads the prompt and builds the KV cache, which acts as working memory for the request. During decode, the model writes the answer one token at a time because each token depends on the previous tokens. A large model such as a 70-billion-parameter Llama model may need hundreds of forward passes to generate an answer, which adds latency.

### A draft model can propose several tokens for the target model to check together
[01:25](https://www.youtube.com/watch?v=XTpyNrEgJQ4&t=85s)
Speculative decoding uses a smaller model to guess the next few tokens. Kirui says a typical cycle can generate three to five draft tokens. The larger target model then verifies them in one forward pass. It accepts the tokens that match its predictions and recomputes the first rejected token. The target model still controls the final output, so the draft model is used for faster guesswork rather than as a replacement for the accurate model.

### Speculative decoding spends extra GPU memory on models and KV caches
[02:29](https://www.youtube.com/watch?v=XTpyNrEgJQ4&t=149s)
The speedup has a memory cost. The serving setup must host both the target model and the smaller draft model, and it needs extra KV-cache space for both. Kirui recommends considering the technique when there is spare GPU capacity. High-concurrency workloads may already use most of the GPU for active requests, leaving too little room for a second model and its cache.

### The draft model must be much smaller without losing too much prediction accuracy
[03:41](https://www.youtube.com/watch?v=XTpyNrEgJQ4&t=221s)
For her one-GPU Blackwell demo, Kirui chose models that could fit alongside the baseline and speculative setups. The baseline model used about 16 gigabytes for weights, while the draft model used 2.5 gigabytes, leaving room for KV caching. She recommends a draft model that is typically 10 to 50 times smaller, uses the same tokenizer, and ideally comes from the same model family. Its speed, cost, and prediction accuracy all affect whether the setup pays off.

### Structured workloads are easier for a draft model to predict
[05:47](https://www.youtube.com/watch?v=XTpyNrEgJQ4&t=347s)
Kirui expects speculative decoding to work better when the output is constrained or predictable. She gives coding, JSON generation, and SQL prompts as examples. Creative tasks such as poetry and brainstorming have more possible continuations, so the smaller model is less likely to match the target model. The acceptance rate of draft tokens is therefore a central measure of whether a workload can benefit.

### The structured-output demo reached 1.6x faster generation
[09:22](https://www.youtube.com/watch?v=XTpyNrEgJQ4&t=562s)
In the side-by-side vLLM demo, speculative decoding generated the structured response much faster. Kirui reports a 1.6x speedup and points to the high acceptance rate as the important result. She also shows higher tokens-per-second output with speculative decoding, which improves throughput for this workload. The acceptance rate measures how many draft-model tokens the target model accepts out of the total proposed.

### Creative generation produced a low acceptance rate
[09:50](https://www.youtube.com/watch?v=XTpyNrEgJQ4&t=590s)
The second demo uses a creative task with a higher temperature. Kirui says the acceptance rate is low because there is more variety in the possible next tokens. When the target model rejects more of the draft model's guesses, speculative decoding has less opportunity to skip expensive sequential work. The contrast with structured output shows why a benchmark on the actual application matters more than enabling the feature by default.

### Long prompts limit the benefit because speculative decoding speeds up decode only
[10:56](https://www.youtube.com/watch?v=XTpyNrEgJQ4&t=656s)
Speculative decoding accelerates token generation after the prompt has been processed. It does not accelerate prefill or the construction of the KV cache. Applications such as retrieval-augmented generation and document analysis may spend much of their time reading a large context. If the input is much longer than the generated answer, the faster decode phase may contribute only a small part of the total latency.

### The decision depends on VRAM, structure, context, and concurrency
[11:37](https://www.youtube.com/watch?v=XTpyNrEgJQ4&t=697s)
Kirui's checklist asks whether the application has enough VRAM, whether its outputs are structured, whether it uses creative generation, how much context it sends, and how many requests it runs together. Smaller batch sizes are more likely to leave room for the second model. She recommends measuring acceptance rate and generation speed on the real workload before enabling speculative decoding.

## Notable quotes
- "The output is still the same because the target model, the main model that has the accuracy, still has to verify the tokens and it does this in one forward pass." (01:48)
- "Typically 10 to 50 times smaller than your target model." (04:48)
- "The key thing to note here is how high the acceptance rate is." (09:22)
- "With speculative decoding it's supposed to accelerate the generating part, not the part where the model reads the prompt." (10:56)

## Tools & references mentioned
- Akamai
- vLLM
- NVIDIA Blackwell
- Llama
- speculative decoding
- Medusa
- EAGLE
- retrieval-augmented generation
- GitHub
- Akamai Developers

## Who should watch
- You are serving an LLM with spare GPU memory and want to know whether a second, smaller model could reduce decode latency.
- Your application produces structured outputs such as code, JSON, or SQL and you need a practical way to measure draft-token acceptance.
- You run long-context, creative, or highly concurrent workloads and want to understand why speculative decoding may produce limited gains.

## Related talks

- [Deep dive on LLM Inference at Scale](https://aietalks.com/talks/deep-dive-on-llm-inference-at-scale) (Harshul Jain, Audible & Tanmay Sah, Independent AI Researcher, 1:28:12)
- [How fast are LLM inference engines anyway?](https://aietalks.com/talks/how-fast-are-llm-inference-engines-anyway) (Charles Frye, Modal, 16:07)
- [Mastering LLM Inference Optimization From Theory to Cost Effective Deployment](https://aietalks.com/talks/mastering-llm-inference-optimization-from-theory-to-cost-effective-deployment) (Mark Moyou, NVIDIA, 33:39)
- [Hacking the Inference Pareto Frontier](https://aietalks.com/talks/hacking-the-inference-pareto-frontier) (Kyle Kranen, NVIDIA, 20:25)
- [Running LLMs Locally: Practical LLM Performance on DGX Spark](https://aietalks.com/talks/running-llms-locally-practical-llm-performance-on-dgx-spark) (Mozhgan Kabiri chimeh, NVIDIA, 10:16)
