DatologyAI seeds synthetic generation with high-quality documents and rephrases them instead of asking a model to create data from scratch.
2
Batching S3 metadata requests reduced initialization time for a trillion-token run from 9–11 days to about two hours.
3
Recoverable retries, cross-cluster resource scheduling, and vLLM parameter sweeps let DatologyAI grow from 30 billion synthetic tokens to about 12 trillion.
Summary
Bogdan Gaza explains how DatologyAI scaled synthetic data generation from limited Slurm runs to about 12 trillion tokens across web, math, and code. The company's BeyondWeb recipe starts with high-quality documents and rephrases them with models, since prompting a model without source material tends to reproduce the modes of its original training distribution. The engineering problems appeared around the data and infrastructure rather than the basic generation loop. Millions of S3 partitions made metadata discovery take 9–11 days until the team batched requests. Large GPU jobs lost too much work when machines failed, so the team combined smaller partitions with checkpointing. Running CPU and GPU work across separate clusters required atomic scheduling. Finally, systematic sweeps of vLLM settings produced about 40% more batch-inference throughput. The talk gives a practical account of the operational work required when synthetic data jobs become large enough to keep substantial GPU capacity waiting on storage, scheduling, or recovery.
Synthetic data addresses the limited amount of usable web text
Gaza says scaling laws require exponentially more compute and data for models to improve linearly, while the web contains a limited amount of usable text. DatologyAI estimates that the web provides about 30 trillion tokens, which is not enough for every large pre-training run. Its approach supplements web data with synthetic data. The company works with customer data, identifies high-quality examples, and then generates much larger datasets from them. The recent run discussed in the talk produced about 12 trillion synthetic tokens across web, math, and code.
BeyondWeb rephrases strong source documents instead of generating from an empty prompt
BeyondWeb starts with high-quality documents from a customer's training data and restructures or reframes them in different ways. Gaza argues that asking a model to produce a synthetic data point without a source document makes it learn the modes of the distribution it was originally trained on. The pipeline prefixes prompts to selected documents, sends them through generation models, and uses GPUs to create new examples. Gaza presents results where BeyondWeb matches or exceeds several comparison datasets at different model sizes and reaches similar performance with fewer tokens. He says the published results are already older than the current recipe.
A unified Ray and Kubernetes workflow replaced a split research and product setup
The first system ran synthetic data research on a Slurm cluster separate from the company's Kubernetes-based product pipeline. The two teams used different codebases, moved data manually, and left GPUs underused. DatologyAI moved toward Kubernetes with Ray, Spark, and vLLM as core components. Its setup uses a control cluster, a compute cluster for Spark and some Ray work, and another Kubernetes cluster for H100 jobs on AWS EKS on HyperPod. The workflow can curate data, generate synthetic data, launch training, and run evaluation as one coordinated sequence.
S3 metadata discovery can stall a trillion-token run before generation starts
At large scale, datasets are stored in formats such as Parquet and may contain millions of partitions. Fetching metadata one object at a time from S3 can hit rate limits and overload Ray metadata handling, leaving GPUs idle. Gaza estimates that the original approach would take 9–11 days to populate metadata for a trillion-token-scale run. The team changed the process to batch S3 requests, using list operations with batches of about 1,000 requests. That reduced API processing time to about two hours. His lesson is to design metadata discovery as part of the synthetic data system before starting expensive GPU work.
Small partitions and checkpointing make GPU failures recoverable
Large synthetic jobs eventually encounter failed machines or bad GPUs. Without recovery, a failure can leave partial outputs, waste compute, and require more cluster time. Gaza gives the example of an eight-hour partition that fails during its final five minutes, losing nearly the whole partition's progress. DatologyAI combines right-sized partitions with periodic checkpoints written back to S3. Failed work can then be retried without repeating a large amount of processing. The system also needs idempotent checkpoint writes and a way to resume execution cleanly.
CPU and GPU capacity must be scheduled as one job across clusters
DatologyAI runs curation in a central location while synthetic jobs may use GPU capacity in different locations. A scheduler can find CPU capacity for a Ray job while failing to find the required GPUs, or find GPUs when the CPU pool is full. Gaza describes this as a bin-packing problem across infrastructure. The solution uses dedicated resource pools and schedules the Ray head and workers together, accounting for both CPU and GPU requirements. This atomic approach avoids separating parts of the same job and lets the pipeline scale across clusters and cloud environments.
Inference tuning produced about 40% more batch throughput
The team found that inference performance depended heavily on the settings used by vLLM and, in testing, SGLang. Synthetic generation uses batch inference, which has different requirements from normal online serving. DatologyAI built a harness to benchmark parameters such as batch size and speculative decoding, then ran grid searches over the settings. Gaza reports about 40% higher throughput from tuning these flags. He says the process needs discipline because the best settings are not obvious from a default configuration, especially when running large workloads.
The infrastructure changes increased output from 30 billion to about 12 trillion tokens
At the end of 2024, DatologyAI could produce around 30 billion synthetic tokens in limited Slurm runs. By the end of the following year, the company had run jobs producing about 7 trillion tokens for web data and another 5 trillion for math and code. Gaza says the team continues to test better synthetic data recipes and plans to examine differences across web, multilingual, math, code, and legal data. The talk's final scale comparison ties the infrastructure work to a large change in the amount of data the company could generate.
"If you just prompt a model and say, hey, give me a synthetic data point, you're going to just learn the modes of the distribution the data was originally trained on."04:44
Who should watch
You are building large synthetic-data jobs and need practical advice on storage metadata, retries, and GPU utilization.
Your pipeline spans CPU and GPU clusters, and separate resource pools are making scheduling unreliable.
You are tuning batch inference for data generation and want an example of why parameter sweeps can matter at production scale.