# From 15% to 90% GPU Utilization: Fix the Data Pipeline, Not the Model

Tarun Sunkaraneni, Amazon AGI | AI Engineer World's Fair 2026 | 17:02

Source: https://www.youtube.com/watch?v=tNAV259UhCM
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/from-15-to-90-gpu-utilization-fix-the-data-pipeline-not-the-model
Published: 2026-10-10
Tags: data-pipelines, gpus, inference, multimodal

## TL;DR
- Multimodal training can spend most of its time loading, transforming, and transferring data instead of doing GPU computation.
- Concurrency, prefetching, and Ray's object store reduced the pipeline's wait-time ratio from 85% to 20%.
- At larger scale, spreading workers across nodes and enabling zero-copy retrieval added 50% throughput by avoiding a network card bottleneck.

## Summary
Tarun Sunkaraneni examines a multimodal supervised fine-tuning pipeline for a Qwen3-VL-based model using images stored in S3. The baseline fetched, decoded, resized, normalized, and tokenized each example serially. Data preparation took more than 100 seconds for a batch while training took 15 to 20 seconds, leaving the GPUs at about 15% utilization. He improves the pipeline in stages. Asyncio handles image fetching, Ray actors handle processing, and a background queue prefetches data before the trainer requests it. Ray's object store replaces repeated copies of large image tensors with references. These changes reduce the wait-time ratio from 85% to 20%. Scaling then exposes a network interface card bottleneck caused by colocating the data generator and workers. Spreading workers across nodes and using zero-copy retrieval produces another 50% throughput at scale. The talk's practical lesson is to profile each bottleneck because every fix can expose another one.

## Key ideas
### Multimodal training can be CPU- and IO-bound before the GPU starts work
[00:41](https://www.youtube.com/watch?v=tNAV259UhCM&t=41s)
Sunkaraneni asks whether the training loop really depends on the GPUs, then argues that the hard part is keeping them supplied with data. Image workloads require loading, resizing, and normalizing with libraries such as PIL. Audio and video add their own loading and transformation work. The training cluster may contain more CPUs than GPUs, yet the GPU can still wait. He defines a wait-time ratio as the time spent preparing and transferring data divided by the total preparation, transfer, and training time. In the baseline experiment, data acquisition consumed 85% of the pipeline time, leaving GPU utilization at 15%.

### The baseline serializes every image operation and leaves the trainer waiting
[04:10](https://www.youtube.com/watch?v=tNAV259UhCM&t=250s)
The example uses JSONL records with image paths stored in S3. A data generator loads and transforms the images, then sends partitions to Megatron data-parallel ranks. In the naive path, each example is fetched, tokenized, resized, normalized, and passed to the model one at a time. A batch of 256 therefore requires serial work across all 256 images. Training takes roughly 15 to 20 seconds, while fetching the data takes more than 100 seconds. Profiling shows that S3 loads dominate, with a smaller tail for loading the image through PIL. The result is only 15% GPU utilization.

### Concurrency should match the bottleneck instead of following a familiar pattern
[06:20](https://www.youtube.com/watch?v=tNAV259UhCM&t=380s)
The first change is a worker pool that processes independent samples concurrently. Sunkaraneni separates image fetching from decoding, vision preprocessing, and text tokenization. He describes asyncio as suitable for IO such as network requests, threads as useful when the called operations release Python's GIL, and processes as heavier workers with isolation across cores. The choice depends on what is slow. For this pipeline, asyncio fetches images while Ray actors perform processing. Ray actors provide process-like workers with communication APIs. This reduces the per-batch fetch time to around 40 seconds and gives a 25% speedup, but the trainer still develops gaps while waiting for future batches.

### Prefetching keeps processed examples ready before the trainer asks for them
[09:07](https://www.youtube.com/watch?v=tNAV259UhCM&t=547s)
The second change puts a background producer in front of the trainer. It fills a queue ahead of time, producing examples faster than the trainer consumes them. When the trainer requests a batch, the data is already available. This removes almost all visible data fetching and processing time per batch because the work happened earlier. Sunkaraneni also names the costs. Prefetching makes checkpointing and resumption for the model and data loader more complicated, and the queue has a cold-start period while it fills. Once those details are handled, keeping the producer ahead of the consumer removes both the load and transform bottlenecks.

### Ray's object store avoids repeated copies of large multimodal tensors
[10:33](https://www.youtube.com/watch?v=tNAV259UhCM&t=633s)
After concurrency and prefetching, data transport becomes the remaining problem. Each multimodal sample can contain image arrays measured in megabytes. The default path pickles a sample on the worker, copies it to the coordinator, and pickles it again for the trainer. These copies can cause the data generator to run out of memory and add communication work that the pipeline does not need. The proposed fix places the sample in Ray's object store and passes an object reference instead. The driver keeps the pointer, while Megatron's data-parallel ranks retrieve the tensor directly from the store. This removes an unnecessary round trip through the head node.

### Scaling changes the bottleneck from data preparation to a single machine's network card
[12:55](https://www.youtube.com/watch?v=tNAV259UhCM&t=775s)
The improved pipeline behaves differently when scaled to a production-sized run. Sunkaraneni increases the data-parallel ranks by four, adds 200 data streams, and multiplies the batch size by four. The wait-time ratio rises because the data generator and worker pool are colocated, concentrating traffic on one node's network interface card. Ray's default scheduling preference also places workers on the same nodes. That can saturate one machine before the workload spreads to another. The problem is therefore caused by placement and network throughput, even though the earlier fixes addressed serial processing, prefetching, and tensor copies.

### Zero-copy retrieval and spread scheduling matter when the system is large
[14:35](https://www.youtube.com/watch?v=tNAV259UhCM&t=875s)
The final experiment changes two settings: zero-copy retrieval and spread scheduling. Spread scheduling distributes workers across multiple nodes so image processing and storage do not overload one network card. The workers then provide data to the data-parallel ranks in a lazy, just-in-time pattern. Neither setting produced a useful improvement in the small-scale experiment, but together they had a strong effect at larger scale. Sunkaraneni reports 50% additional throughput from these two flags. His broader point is that a setting can look ineffective in a small test and still matter after the number of workers, streams, and batches increases.

### Every optimization can expose a different bottleneck
[15:35](https://www.youtube.com/watch?v=tNAV259UhCM&t=935s)
Sunkaraneni's final advice is to profile the baseline and ask whether the system is CPU-bound, GPU-bound, IO-bound, or compute-bound. In this pipeline, the bottleneck moved from throughput to wait time, then to memory, and then to retrieval cost. Parallelism and prefetching helped, but the semantics and constraints of the framework affected the result. Scale also changed which settings mattered. Zero-copy retrieval and worker spreading did little at small scale and became important later. The talk treats each improvement as a step in diagnosis rather than a universal configuration. The next bottleneck appears after the current one is removed.

## Notable quotes
- "The hardest part of optimizing multimodal train throughput is not necessarily kernel optimization. It is making sure that the GPUs are constantly fed with data." (00:41)
- "For us, GPUs are not the goal. It's making sure that the data is never the bottleneck is the goal." (02:58)
- "We went from a baseline wait ratio of 85% to 20%." (12:26)
- "Config can do poorly at a small scale but can do really good at a larger scale." (14:46)
- "GPUs are expensive but sometimes the biggest battles happen outside of the GPU." (15:54)

## Tools & references mentioned
- Ray
- Ray actors
- Ray object store
- asyncio
- Megatron
- Qwen3-VL
- S3
- PIL
- Hugging Face Transformers tokenizers
- Python GIL
- Amazon AGI

## Who should watch
- You are training multimodal models and your GPUs show low utilization even though the model code looks efficient.
- Your data loader fetches images from object storage and performs decoding or preprocessing inside the training path.
- You are scaling a Ray-based pipeline and need to understand why a configuration that helps at large scale does little on one machine.

## Related talks

- [The Small Model Infrastructure Nobody Built (So We Did)](https://aietalks.com/talks/the-small-model-infrastructure-nobody-built-so-we-did) (Filip Makraduli, Superlinked, 18:30)
- [Vertical Mobility: Inference from MVP to Trillion-Parameter Workloads](https://aietalks.com/talks/vertical-mobility-inference-from-mvp-to-trillion-parameter-workloads) (Sitanshu Gupta, CoreWeave, 15:22)
- [Weight Folding, CUDA Streams, and the Bug That Made My Model Speak Backwards](https://aietalks.com/talks/weight-folding-cuda-streams-and-the-bug-that-made-my-model-speak-backwards) (Filip Makraduli, Superlinked, 17:18)
- [The Messy Reality of Scale: Synthetic Data and Pre-Training](https://aietalks.com/talks/the-messy-reality-of-scale-synthetic-data-and-pre-training) (Marah Abdin & Robert McHardy, Poolside, 17:31)
- [What Every AI Engineer Needs to Know About GPUs](https://aietalks.com/talks/what-every-ai-engineer-needs-to-know-about-gpus) (Charles Frye, Modal, 19:52)
