# What Makes Open Models Fast in Production

Sujee Maniyam & Dylan Bristot, Nebius | AI Engineer | 20:32

Source: https://www.youtube.com/watch?v=TRe1u7dHYiA
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/what-makes-open-models-fast-in-production
Published: 2026-10-03
Tags: inference, open-models, quantization

## TL;DR
- Open models have become competitive with proprietary models, so teams can gain choice and lower token costs without vendor lock-in.
- Production inference depends on matching models to serving engines, routing requests by KV-cache location, and separating prefill from decoding.
- Speculative decoding, KV-cache offloading, and carefully chosen quantization can improve speed while preserving model quality.

## Summary
Dylan Bristot presents Nebius Token Factory as a managed alternative to closed APIs and self-hosting. Closed APIs are easy to start with but limit model control and optimization. Self-hosting gives control but requires a large engineering effort. Token Factory connects inference, production data, post-training, and deployment so teams can improve and redeploy models in one loop. Sujee Maniyam then explains why fast production inference depends on the whole serving stack. Open models are now close to proprietary models on capability and can cost less. Nebius matches models with suitable engines, uses cache-aware routing, and applies speculative decoding with smaller draft models. It also caches KV values, offloads them from GPU memory when needed, separates prefill from decoding, and tests quantization levels to find a quality and performance balance. The talk is a concise tour of the engineering hidden behind an inference API.

## Key ideas
### Open models give teams more choice as their quality approaches proprietary models
[10:10](https://www.youtube.com/watch?v=TRe1u7dHYiA&t=610s)
Sujee Maniyam points to an Artificial Analysis benchmark where open models are competitive with, and sometimes better than, proprietary models. He says teams do not always need to choose a proprietary model for intelligence. Open models can also cost less in many cases. Running them gives customers more choice across providers and reduces vendor lock-in. The talk then separates larger teacher models from smaller student models. A large teacher can train a smaller model, allowing teams to use different model sizes for different production needs.

### A production AI system needs a loop from inference through data and post-training to deployment
[04:06](https://www.youtube.com/watch?v=TRe1u7dHYiA&t=246s)
Dylan Bristot describes Token Factory as a platform that connects four stages. Inference runs models through real-time or batch endpoints, structured outputs, and function calling. Data Lab captures production completions so teams can filter, version, and export datasets. Post-training uses those logs and synthetic data for LoRA, full fine-tuning, distillation, and custom speculative decoding. Deployment pushes updated models into production. Bristot says the loop matters because teams need to observe how a model performs, improve it, and redeploy updates without breaking the product.

### The serving engine must be chosen for each model rather than treated as a generic detail
[11:46](https://www.youtube.com/watch?v=TRe1u7dHYiA&t=706s)
Maniyam says Nebius works with open-source serving engines and maintains an internal fork that it continually tunes. Different models do not run equally well on every engine, so the platform selects an engine for the model instead of exposing that decision to the customer. The same open-source model can have different speed and cost characteristics depending on the engine and the surrounding infrastructure. Customers use an API while Nebius handles the engine, GPU, and runtime choices behind it.

### Cache-aware routing keeps related requests near the KV cache they can reuse
[12:55](https://www.youtube.com/watch?v=TRe1u7dHYiA&t=775s)
Random load balancing can fragment an LLM's cache across GPUs. Maniyam explains that inputs now vary widely, from short prompts to large codebases used by coding agents, so routing requests randomly is not enough. A cache-aware router tracks where cached values are stored and sends a request to a GPU that can reuse them. His slide contrasts fragmented cache placement with a more coherent layout. Reusing the cache improves inference speed because the system does not have to recreate work that was already done.

### Speculative decoding uses a small draft model and a large model that verifies its output
[14:24](https://www.youtube.com/watch?v=TRe1u7dHYiA&t=864s)
Large language models generate tokens one at a time, which can make generation slow. Speculative decoding has a smaller model generate candidate tokens while the larger model verifies them. Maniyam compares this to a senior engineer checking work produced by a junior engineer. When the large model accepts the candidates, generation moves faster. If it rejects them, it regenerates the tokens. He says draft models work better when trained on a customer's own production data, and Nebius is building a feature that captures that data and trains the draft model with little manual setup.

### KV-cache offloading trades GPU memory pressure for reuse
[15:59](https://www.youtube.com/watch?v=TRe1u7dHYiA&t=959s)
The KV cache stores information from tokens that have already been processed, so the system can look up cached work instead of generating it again. Maniyam says Nebius has seen speedups ranging from 5x to 10x in situations where caching works well. The cache can consume substantial memory, especially with large inputs and long context windows. Since GPU memory is limited, Token Factory can move cached data into regular memory and bring it back when needed. The platform manages this movement automatically rather than discarding the cache.

### Separating prefill and decode lets each stage use hardware for its own workload
[17:33](https://www.youtube.com/watch?v=TRe1u7dHYiA&t=1053s)
Maniyam divides LLM inference into prefill and decode. Prefill fills the context and is compute-intensive. Decode generates tokens and is memory-intensive. Running both stages on the same GPU makes them compete for resources. Nebius separates them across GPU groups, with one group handling prefill and another handling decoding. The groups transfer the KV cache between them. This arrangement lets each stage run on hardware suited to its workload.

### Quantization has a quality threshold that must be found through testing
[18:13](https://www.youtube.com/watch?v=TRe1u7dHYiA&t=1093s)
Quantization reduces numerical precision so models can run more efficiently. Maniyam warns that reducing precision too far can degrade model quality. Nebius therefore runs experiments to find a sweet spot where the model uses fewer resources without losing too much performance. Quantization is presented alongside speculative decoding, cache-aware routing, KV-cache handling, and prefill/decode separation as part of the work hidden behind a production inference API.

## Notable quotes
- Sujee Maniyam: "The gap between proprietary and open models is very narrow." (10:50)
- Sujee Maniyam: "You can train them using synthetic data, and even if you train them using generic data you get up to 30% improvement." (15:33)
- Sujee Maniyam: "We have seen speedups anywhere from 5 to 10x, not just 5 to 10 percentage, actually 10x." (16:47)
- Sujee Maniyam: "We do a lot of experiments to figure out what's the kind of the sweet spot of a quantization." (18:17)

## Tools & references mentioned
- Nebius Token Factory
- Nebius
- NVIDIA
- Artificial Analysis
- Microsoft
- Meta
- Cursor
- OpenCode
- GLM
- Kimi
- DeepSeek
- NVFP4
- Blackwell Ultra
- GB300

## Who should watch
- You are deciding between a closed model API and operating open models yourself, and want to understand the engineering cost behind each choice.
- You run an LLM service with long prompts, uneven request sizes, or coding-agent traffic where ordinary random load balancing wastes cache reuse.
- You are tuning inference cost and latency and need practical options beyond simply buying more GPUs.

## Related talks

- [Making Open Models 10x Faster and Better for Modern Application Innovation](https://aietalks.com/talks/making-open-models-10x-faster-and-better-for-modern-application-innovation) (Dmytro (Dima) Dzhulgakov, Fireworks AI, 18:55)
- [Customized, production-ready inference with open source models](https://aietalks.com/talks/customized-production-ready-inference-with-open-source-models) (Dmytro (Dima) Dzhulgakov, Fireworks AI, 18:55)
- [The Rise of Open Models in the Enterprise](https://aietalks.com/talks/the-rise-of-open-models-in-the-enterprise) (Amir Haghighat, Baseten, 16:50)
- [How fast are LLM inference engines anyway?](https://aietalks.com/talks/how-fast-are-llm-inference-engines-anyway) (Charles Frye, Modal, 16:07)
- [Trends Across the AI Frontier](https://aietalks.com/talks/trends-across-the-ai-frontier) (George Cameron, Artificial Analysis, 17:52)
