Dream Machine: Scaling to 1m users in 4 days

Keegan McCallum, Luma AI19:03 · Jul 2025 · 1,794 views
Thumbnail for Dream Machine: Scaling to 1m users in 4 days Watch on YouTube
TL;DR
  1. 1

    Luma AI scaled from about 500 to 5,000 H100 GPUs within six hours after Dream Machine launch traffic overwhelmed its initial capacity.

  2. 2

    Luma replaced a brittle Triton-based setup with a PyTorch serving stack that could run across multiple GPUs, nodes, and chip types.

  3. 3

    Fair scheduling uses service-level objectives to prevent lower-priority jobs from waiting indefinitely when resources are limited.

Summary

Keegan McCallum describes the infrastructure problems behind Dream Machine's launch on June 11, 2024. Luma expected heavy traffic, but its initial allocation of about 500 H100 GPUs was quickly overwhelmed by a queue approaching 100,000 requests. The team manually added capacity, reached about 5,000 H100s within six hours, and then took another 4,000 GPUs from the training cluster. As demand continued, Luma rebuilt its serving system around vanilla PyTorch, decoupled CPU workers from GPU workers, and connected machines across providers through shared queues and storage. McCallum explains the resulting problems, including back pressure, model warm-up time, burst handling, and starvation between user tiers. Luma's scheduler uses product-defined service-level objectives to age jobs fairly. Model versions are stored with their environments and checkpoints, while YAML selects the active version and lets workers roll out changes without restarting.

Key ideas
00:03

Dream Machine's launch overwhelmed Luma's initial GPU allocation

Luma announced Dream Machine at 9:00 a.m. on June 11, 2024, expecting significant traffic. The team had allocated about 500 H100 GPUs, but requests quickly filled a queue that grew to almost 100,000 jobs. Engineers manually ran SSH commands across every provider they could access and reached about 5,000 H100s within six hours. The queue began draining around 2 p.m., until the CEO tweeted that capacity had increased tenfold. Ten minutes later, the queue started rising again. Luma then moved another 4,000 H100 GPUs from its training cluster, but that barely changed the queue.

02:02

Luma is building a general multimodal model lab, not only a video product

McCallum describes Luma as a foundation model lab aiming to build general multimodal intelligence. Its goal is for models to generate, understand, and operate in the physical world like a human can. He demonstrates Modify Video, where users upload iPhone videos and apply text prompts to transform their contents. Luma also offers a public API and SDK. Applications can send raw user prompts and receive generated images and videos without doing complex prompt engineering themselves.

04:28

Triton became a poor fit for Luma's multi-GPU video workloads

At launch, Luma used tightly coupled containers and Triton Inference Server. The arrangement could run on raw machines with few dependencies, but it was brittle. If Triton went down, CPU processes could continue pulling jobs that then failed. Video models often needed multiple GPUs and nodes to meet latency requirements, which Triton did not support well. Luma also needed to run on AMD and other chipsets, while Triton had stronger support for Nvidia hardware. Researchers found the system difficult to develop against because it required different idioms and many special configuration steps.

06:09

A PyTorch serving layer lets Luma spread workers across providers

Luma rebuilt its serving stack on vanilla PyTorch. McCallum says this gave the team a common base because chipset vendors generally make sure PyTorch works on their hardware. Some operations still need optimization, but models can usually be made to run on different chips. The new design decouples CPU workers from GPU workers. CPU workers queue multimedia inputs so GPUs are not blocked waiting for videos or images. GPUs can run on arbitrary providers or virtual machines as long as they connect to Redis and distributed storage such as SeaweedFS. Tailscale connects the pieces.

07:48

Decoupling creates back pressure that needs an explicit dispatch limit

With several clusters pulling from one global queue, too many CPU workers can pull jobs into one cluster and leave them waiting for that cluster's GPUs. Those jobs might have run elsewhere. Luma added a dispatch limitation system. It tracks jobs that have entered a cluster but have not yet been picked up by a GPU, then limits how many jobs may sit in that state. The limit prevents CPU workers from filling a cluster with work that its GPUs cannot process while capacity exists in another cluster.

10:38

Service-level objectives prevent low-priority jobs from waiting for hours

When resources became limited, Luma had to schedule API, enterprise, unlimited, light, and free jobs with different priorities. Strict priority could leave light jobs waiting seven, eight, or nine hours. Luma instead uses service-level objectives defined by the product team. An API job might have a wait target of a couple of minutes, while a light job might have a ten-minute target. Once a job reaches a chosen fraction of its allowed wait, it moves forward. To avoid a second starvation problem, Luma ranks jobs by the percentage of their SLO that has elapsed. An API job waiting one minute can therefore compete with a light job waiting ten minutes.

13:30

Immutable model versions make rollback and fleet-wide rollout practical

Each model has a folder in object storage with immutable versioned subfolders. Versions contain the model code, Python environment, dependencies, and checkpoint. A YAML file at the root identifies the active version. This lets Luma reproduce the exact environment and checkpoint used previously, which makes rollback a practical response when a deployment causes trouble. Luma also built automated rollout on top of the model repository. When the YAML changes, workers switch versions without restarting, allowing updates across thousands of H100 and AMD GPUs.

16:14

PyTorch gives new chipsets a starting point, while Luma's optimization team makes them fast

In the questions, McCallum calls PyTorch support from chip providers a kind of cheat code. A model will often run on a new chipset, but it may initially run slowly. Luma has a team of about 10 people called the Excel team that optimizes low-level PyTorch operations, uses Triton, and works with chip providers to improve speed. Luma works closely with Nvidia and AMD and is exploring other providers, including Amazon and Groq, through partnerships. Cloud providers generally give Luma a working Kubernetes cluster, while Luma does not provision the nodes itself.

"We had the insight that this is a product problem and it really comes down to how long worst case are you okay with different tiers waiting."11:49
Who should watch
  • You are building an inference service for large video, image, or other multimodal models and need to share work across GPU clusters.
  • Your queue has several customer tiers, and strict priority causes lower-priority jobs to wait indefinitely.
  • You deploy models across changing hardware and need rollbacks without rebuilding every worker.