Operating Distributed Inference Systems at Scale

Nishant Gupta, Meta, Naman Ahuja, Meta19:51 · Sept 2026 · 3,670 views
Thumbnail for Operating Distributed Inference Systems at Scale Watch on YouTube
TL;DR
  1. 1

    Inference traffic at Meta has outgrown the largest microservices and requires workload-aware scheduling and admission control.

  2. 2

    Inference layers are tightly coupled, so routing, caching, batching, GPU utilization, autoscaling, and reliability must be managed together.

  3. 3

    The right optimization target is cost per successful task, with an inference control plane coordinating routing, batching, caching, scheduling, and failure handling.

Summary

Nishant Gupta and Naman Ahuja describe inference as a distributed systems problem whose constraints differ from ordinary microservices. Requests can range from 50 to 100,000 tokens, carry expensive KV-cache state, run on costly GPUs, and stream partial results that cannot be safely retried after a failure. Routing decisions affect cache hits, batch composition, GPU utilization, and autoscaling, so teams need an end-to-end view of the stack. Gupta proposes four ways to reduce work: avoid it with caching, share it through batching, move it with routing, or delay it through admission control. The metric should be cost per successful task, including retries, failures, storage, networking, and operations. Ahuja adds observability as the input to an automatic control loop and frames serving choices as tradeoffs among latency, cost, and throughput. They argue that inference needs its own control plane, much as Kubernetes organized virtual machines and cluster operations.

Key ideas
00:12

Inference capacity grows with agent activity, not just user count

Nishant Gupta says inference traffic at Meta already exceeds the largest microservices and is growing faster than any workload Meta has seen. Traditional web capacity often scaled roughly with users and QPS. Agentic serving adds the number of calls per user and the number of tokens. A chatbot may make one model call per turn, while copilots make 10 to 20, research agents make 50, and autonomous workflows can make thousands. Capacity planning therefore needs elasticity, workload-aware scheduling, and admission control rather than a simple fleet-size calculation.

00:50

Inference is following cloud infrastructure toward an orchestration layer

Gupta compares the current AI infrastructure cycle with cloud infrastructure around 2008. Cloud systems began with virtual machines, then moved value and complexity into schedulers, service meshes, autoscalers, and platforms. He says AI is moving through a similar progression in a shorter period. Model serving frameworks came before an emerging orchestration layer for routing, KV-cache management, prefill and decode disaggregation, and multimodel multiplexing. The talk focuses on this control and orchestration layer rather than only on models or kernels.

03:13

Inference differs from microservices in request shape, state, hardware, and failure handling

LLM requests can range from 50 to 100,000 tokens and have different compute profiles during prefill and decode. They need continuous in-flight batching, while traditional microservices may batch at the load-balancer layer or not at all. Inference also carries expensive per-request KV-cache state. GPUs cost much more than ordinary CPU pods, take longer to acquire, and cannot be casually overprovisioned. A crashed microservice can often restart and rebuild state, but a cold model startup takes time, and losing a GPU during decode can drop thousands of in-flight tokens and create a queue buildup.

06:09

Small decisions across the stack create one coupled system

The individual layers of an inference stack are familiar, but their interactions create new behavior. Gupta gives a chain in which routing changes the cache hit rate, the cache hit rate changes batch composition, batch composition changes GPU utilization, and utilization changes the autoscaling decision. A regression therefore cannot be diagnosed inside only one layer. Teams need to inspect routing, caching, admission control, scheduling, serving, and the hardware stack together. The same end-to-end view is needed to identify where a bottleneck actually sits before investing in a fix.

07:01

A streamed inference request behaves like a distributed transaction

A prompt can pass through a gateway, router, cache lookup, scheduler, GPU cluster, serving runtime, and streaming response path. Each arrow is a network hop that can retry, time out, fall back, or fail, and each hop has its own service-level objective. Partial failure is harder when the system has already streamed output. If a GPU is preempted after 200 tokens, the request cannot simply be retried as an ordinary RPC. Gupta argues that reliability must be a property of the control plane because that layer can see the complete workflow.

08:31

An inference scheduler needs workflow and resource context

Gupta says an inference scheduler must consider at least seven dimensions: GPU type and network topology, memory headroom, KV-cache state, whether model weights are warm, tenant priority, workflow context, and latency budget. A request at step three of a five-step workflow has already spent resources on earlier steps, so a failure or retry decision has a different cost from one at the beginning. Workflow-aware scheduling can change admission, priority, placement, and retry behavior. The aim is to place work where it will finish fastest and most cheaply rather than on an arbitrary available GPU.

10:19

Every optimization reduces, shares, moves, or delays work

Gupta groups inference optimizations into four categories. Caching avoids work through prefix, response, or semantic caching. Batching shares work across requests, including continuous batching and chunked prefill. Routing moves work to a smaller model, a cheaper region, or a location closer to the user. Admission control and queueing delay work until a better moment, using priorities and deadlines. He says this framework applies across serving stacks and gives teams a consistent way to compare new optimizations with existing ones.

12:36

Reliability requires breaking feedback loops before they cascade

A cascading failure can begin with a degraded GPU that raises latency, triggers client retries, increases queue depth, and saturates healthy GPUs. Agentic workloads add a KV-cache problem because traffic cannot be casually rerouted to a cold cluster. Gupta recommends deliberate circuit breakers at the routing layer, admission control that rejects load instead of endlessly queuing it, load shedding tied to queue depth rather than only CPU or memory, and retry budgets. A cold pool needs time to warm up, so the hot pool must otherwise absorb the extra traffic.

16:19

Serving decisions trade latency, cost, and throughput

Naman Ahuja describes observability as the input to a control loop rather than only a dashboard. Useful signals include time to first token, utilization ratio, success per dollar, and end-to-end latency. Serving choices move the system around a tradeoff triangle. Larger batches can improve throughput and cost efficiency while hurting tail latency. Speculative decoding may reduce latency while adding compute. A smaller model can lower latency and cost while reducing response quality, and failures or retries can raise cost again. The platform has to choose a setting based on the product's needs.

"What is important is to understand what is the key performance indicator for your product which will add value to the users."Nishant Gupta11:56
Who should watch
  • You are building an agent or copilot whose request volume depends on calls per user, workflow steps, and token counts.
  • Your inference stack has separate teams for routing, caching, GPU scheduling, serving, or reliability, and failures are hard to explain across those boundaries.
  • You need to choose between batching, caching, routing, admission control, or retries while accounting for latency and total task cost.