# Why We Don't Need More Data Centers

Dr. Jasper Zhang, Hyperbolic | AI Engineer World's Fair 2025 | 14:18

Source: https://www.youtube.com/watch?v=M6Vbaig1TsM
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/why-we-dont-need-more-data-centers
Published: 2025-08-01
Tags: cost, gpus, inference

## TL;DR
- Building more data centers will not by itself meet AI's GPU demand because much of the existing capacity sits idle.
- A GPU marketplace can combine supply from different data centers and offer spot, on-demand, and reserved access at lower prices.
- Flexible GPU rentals let startups adjust capacity as their training and inference needs change, while other users can buy their unused capacity.

## Summary
Dr. Jasper Zhang argues that AI infrastructure needs better use of existing GPUs alongside new data centers. He points to long grid-connection waits, high construction costs, energy use, and a projected US data-center supply deficit. At the same time, enterprise GPUs can sit idle for much of the time, while companies struggle to find affordable capacity. Hyperbolic's proposed answer is a GPU marketplace and orchestration layer. Its HyperDOS software connects Kubernetes clusters to a network where users can rent GPUs through spot, on-demand, or long-term reservations. Zhang says this model can lower prices, reduce supplier-vetting work, and let companies release unused machines. His startup example shows a company changing its GPU needs during training and inference instead of holding one large fixed reservation. The talk also explains how HyperDOS provisions machines through a central server and provider-side agents.

## Key ideas
### GPU demand is growing faster than data centers can be built
[01:20](https://www.youtube.com/watch?v=M6Vbaig1TsM&t=80s)
Zhang says AI will become part of every company's work, which is driving demand for GPUs and data centers. He cites a projection that the world will need four times more data centers by 2030, built in one quarter of the current time. He estimates current data-center capacity at 55 gigawatts and says a median demand scenario reaches 219 gigawatts in 2030, with 22% annual growth. The problem is timing as well as scale. A facility in Northern Virginia can wait seven years for a 100-megawatt grid connection, according to Zhang.

### New data centers carry high financial and environmental costs
[02:33](https://www.youtube.com/watch?v=M6Vbaig1TsM&t=153s)
Zhang uses the first Stargate data center as an example of the money required to add capacity, saying it cost more than a billion dollars to build. He also says GPUs and data centers already account for 4% of total US electricity consumption. Even if planned facilities arrive on schedule, he expects a US data-center supply deficit of more than 15 gigawatts by 2030. His argument is that construction alone cannot close the gap quickly enough, and it adds pressure through energy use, land use, and carbon emissions.

### Existing GPU capacity is fragmented and often idle
[03:43](https://www.youtube.com/watch?v=M6Vbaig1TsM&t=223s)
Zhang says enterprise and company GPUs sit idle 80% of the time, citing data attributed to Deote. At the same time, there are more than 100 GPU cloud providers, according to an analysis he mentions. This creates a mismatch. Some users cannot find GPUs or have to pay very high prices, while other data centers and clouds have unused machines. Zhang proposes an aggregation layer that brings these providers together so available capacity can be matched with users who need it.

### Hyperbolic connects provider clusters to a shared GPU marketplace
[04:38](https://www.youtube.com/watch?v=M6Vbaig1TsM&t=278s)
Hyperbolic's orchestration software is called HyperDOS, short for Hyperbolic Distributed Operating System. Zhang compares it with Kubernetes software. A provider installs it on a cluster, and Hyperbolic can add that cluster to its network within five minutes, according to his description. Users can then rent GPUs through spot instances, on-demand access, or long-term reservations. They can also host models on the platform. The marketplace is intended to give users one place to compare supply instead of contacting many separate data centers.

### Aggregating supply can lower GPU prices
[06:13](https://www.youtube.com/watch?v=M6Vbaig1TsM&t=373s)
Zhang says Hyperbolic's modeling indicates potential cost savings of 50% to 75%. He gives a current beta-marketplace example in which an H100 costs 99 cents per hour, compared with about $11 for Google's on-demand GPU and $2 to $3 at Lambda. He attributes the lower price to aggregating more supply and distributing it through a common channel. He also refers to queueing theory, specifically an M/M/c model, as part of the reasoning behind the marketplace economics.

### Flexible reservations let startups match compute to changing work
[08:13](https://www.youtube.com/watch?v=M6Vbaig1TsM&t=493s)
Zhang describes a startup that initially reserves 1,000 GPUs for a year, then needs 1,000 additional GPUs for one month after early experiments, and later needs only 500 GPUs to host its model. In the marketplace model, the company can add short-term capacity in the third month and release its unused GPUs in the sixth month for other customers to rent. Zhang contrasts this with a traditional cloud arrangement that requires large fixed reservations. In his example, the cost falls from $43.8 million to $6.9 million, which he describes as six times lower.

### Lower compute costs can let startups run more of their own training
[10:21](https://www.youtube.com/watch?v=M6Vbaig1TsM&t=621s)
Zhang says the benefit is not limited to reducing a bill. Since model quality improves with more compute under scaling laws, he argues that the same budget can produce more training work. He describes startups that currently rely on closed models such as OpenAI and Anthropic because they cannot afford enough GPUs. Cheaper, more flexible access could let them rent more capacity for their own training. He expects a GPU marketplace to develop into a platform for training, online inference, and offline inference.

### HyperDOS provisions machines through a central server and provider agents
[12:46](https://www.youtube.com/watch?v=M6Vbaig1TsM&t=766s)
In the question period, Zhang explains that HyperDOS is a Kubernetes agent. A data center with Kubernetes can install it, while an individual machine can use MicroK8s to become Kubernetes-ready. Hyperbolic calls its central server the 'monarch' and the providers that own compute 'barons.' When a user requests a GPU, the monarch sends the request to a baron. The baron provisions the machine and sets up an SSH instance for the customer.

## Notable quotes
- "Building data centers alone can solve the problem." (00:01)
- "GPU sit idle 80% of the time for enterprises and companies." (03:43)
- "With the same budget you will increase your productivity by 6x." (10:43)
- "We should better reuse recycle those idle compute by selling it to others." (11:59)

## Tools & references mentioned
- Hyperbolic
- HyperDOS
- Kubernetes
- MicroK8s
- Stargate
- UC Berkeley
- Citadel Securities
- OpenAI
- Anthropic
- Google
- Lambda
- McKinsey

## Who should watch
- You are a startup founder deciding whether to reserve a large GPU fleet for training and later inference.
- Your team is spending time comparing cloud GPU suppliers or paying high prices while capacity elsewhere remains unused.
- You run a data center or GPU cluster and want a way to monetize machines when your own workloads are idle.

## Related talks

- [How to Build Your Own AI Data Center in 2025](https://aietalks.com/talks/how-to-build-your-own-ai-data-center-in-2025) (Paul Gilbert, Arista Networks, 23:00)
- [The Desktop Frontier](https://aietalks.com/talks/the-desktop-frontier) (Ahmad Osman, Osmantic, 18:02)
- [What Every AI Engineer Needs to Know About GPUs](https://aietalks.com/talks/what-every-ai-engineer-needs-to-know-about-gpus) (Charles Frye, Modal, 19:52)
- [The Geopolitics of AI Infrastructure](https://aietalks.com/talks/the-geopolitics-of-ai-infrastructure) (Dylan Patel, SemiAnalysis, 18:29)
- [Stop Renting Your Cognitive Infrastructure](https://aietalks.com/talks/stop-renting-your-cognitive-infrastructure) (Thiyagarajan Maruthavanan, Kalmantic Labs, 07:52)
