# Move Fast and Don't Break Things: Scaling Databases for the AI Era

Ben Dicken, PlanetScale | AI Engineer World's Fair 2026 | 19:07

Source: https://www.youtube.com/watch?v=uKUA1a0Kdfc
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/move-fast-and-dont-break-things-scaling-databases-for-the-ai-era
Published: 2026-10-06
Tags: agents, reliability, tool-use

## TL;DR
- Reliable systems isolate databases from failures in control planes, applications, observability, and other infrastructure.
- Sharding, back pressure, and decoupling let databases handle more data and traffic while limiting the effect of overload.
- Simple interfaces such as VSchema files, resource budgets, database branches, deploy requests, and reverts let AI agents operate infrastructure more safely.

## Summary
Ben Dicken argues that teams can ship quickly without accepting database outages. He starts with two reliability principles: isolate the data plane from components that fail more often, and use redundant database replicas so failures do not become outages. For larger systems, he explains sharding data and queries across many servers, back pressure that rejects some traffic before overload crashes the database, and decoupling that allows backend components to scale independently. The second half applies these ideas to AI agents. A Vitess VSchema is a JSON file that an agent can help design. Traffic Control assigns resource budgets to traffic categories and degrades service before the whole application fails. Database branching, deploy requests, and one-click reverts give agents a Git-like workflow for schema changes. Dicken's argument is practical: developer tools built for safe human operation also give agents safer ways to change infrastructure.

## Key ideas
### Isolation keeps a failing component from taking down the database
[03:31](https://www.youtube.com/watch?v=uKUA1a0Kdfc&t=211s)
Dicken separates the data plane, where critical user data lives, from control planes, observability pipelines, analytics, applications, and MCPs. These surrounding components are deployed more often and are therefore more likely to fail than the database. If a bad control-plane deployment takes the control plane down, users should still be able to access the data plane. The same boundary limits the damage from failures in other parts of the stack.

### Redundant database replicas turn frequent failures into planned failover
[04:32](https://www.youtube.com/watch?v=uKUA1a0Kdfc&t=272s)
Application servers can often scale easily because they are stateless, but databases must preserve state without losing a byte. Dicken describes a primary database node with replica or follower servers in other data centers or availability zones. At small scale, a single server might fail once every couple of years. With thousands of servers, a once-or-twice-a-year failure rate for each server becomes a weekly or monthly event somewhere in the system. Replicas are needed to handle that frequency.

### Sharding spreads data and queries across many independently replicated servers
[07:00](https://www.youtube.com/watch?v=uKUA1a0Kdfc&t=420s)
A single database node may work with 10 gigabytes but becomes unsuitable for 10 terabytes, 100 terabytes, or a petabyte. Sharding distributes data and queries across many servers, with a proxy routing each request to the appropriate shard. Dicken describes Vitess for MySQL and Neki for Postgres. Each shard still needs its own primary and replicas, so a failure within one shard can be repaired without taking down the entire database.

### Back pressure rejects excess work before overload becomes a total outage
[08:26](https://www.youtube.com/watch?v=uKUA1a0Kdfc&t=506s)
Many databases handle overload badly. Too many queries, excessive CPU use, or an exhausted cache can cause a server to crash. Back pressure adds thresholds that make the system push back by failing requests or denying connections. Some traffic is degraded while existing users remain healthy. The rest of the application has to support this behavior, since refusing selected work is less damaging than losing the database for everyone.

### Decoupled backend services can scale without sharing every failure
[09:41](https://www.youtube.com/watch?v=uKUA1a0Kdfc&t=581s)
Combining OLTP, analytics, queues, jobs, and caching in one system may be convenient, but it adds complexity and weakens isolation. A burst of queue jobs should not require the database or unrelated services to scale with it. Separate components can receive their own compute resources and fail independently. Dicken presents decoupling as a way to control both resource contention and the scope of an outage.

### A VSchema gives an agent a simple control surface for complex sharding
[12:56](https://www.youtube.com/watch?v=uKUA1a0Kdfc&t=776s)
Vitess uses VTGate to parse MySQL queries, identify their destination, and route them to the right shard. The sharding configuration exposed to users is much simpler than the machinery underneath. A VSchema is a JSON file that specifies the column to shard on, the sharding method, and where particular tables should go. Dicken says an agent can inspect a database schema and help prototype this file, after which PlanetScale APIs or its command-line interface can create the sharded system.

### Traffic budgets make overload degrade in a controlled way
[14:51](https://www.youtube.com/watch?v=uKUA1a0Kdfc&t=891s)
PlanetScale's Traffic Control lets users categorize and tag database traffic, then assign resource budgets for CPU or backend processes. The system can warn clients when traffic exceeds a budget and tell them to slow down. It can also become stricter and kill queries that are executing. Dicken admits that denying requests is undesirable, but says it is preferable to taking the entire application offline.

### Database branches and reverts give agents a safer schema-change workflow
[15:55](https://www.youtube.com/watch?v=uKUA1a0Kdfc&t=955s)
Dicken applies a Git-like workflow to database schemas. Developers, or agents, can create a branch, make schema changes in isolation, and merge them into production through a deploy request without downtime. If a deployed schema causes a problem, PlanetScale provides a one-click revert to the previous state without data loss. APIs and a CLI allow this workflow to run automatically or with a human reviewing the changes.

### Developer infrastructure primitives also give agents safer operating tools
[17:45](https://www.youtube.com/watch?v=uKUA1a0Kdfc&t=1065s)
Dicken defines developer experience as the tools that help infrastructure owners deploy safely, improve performance, and avoid mistakes. He says the same interfaces built for developers, including APIs, command-line tools, and an MCP server, can let agents work with databases. Agents can automate branching, code changes, pull requests, and reviews while keeping humans in the loop when needed.

## Notable quotes
- "I want to be able to ship fast, build great features, get things out to my customers, but also not break the databases, the infrastructure, the application servers that are powering what I'm doing." (00:47)
- "At a large scale when you have thousands of servers a once or twice a year failure becomes a weekly or monthly failure." (05:29)
- "What back pressure is is it's designing a system to be able to push back and say when I detect that I'm above a certain threshold of resources I can actually say I want to push back and start failing requests denying query uh denying connections." (09:04)
- "It's literally a JSON file where you say here's the column I want to shard on and here's the way I want to do sharding." (12:54)
- "It does kind of turn out that when you really focus on that core of the developer experience, what you actually end up with is also very good primitives for AI being able to work with a database safely and reliably and helping it scale." (18:18)

## Tools & references mentioned
- PlanetScale
- Cursor
- Vitess
- Neki
- MySQL
- Postgres
- VTGate
- Traffic Control
- GitHub
- Cloudflare Workers
- Vercel Functions
- AWS Lambda
- MCP

## Who should watch
- You run an AI product whose traffic can rise faster than your database capacity, and you need concrete ways to limit the damage from overload.
- Your team is considering sharding or database changes managed by AI agents and wants to understand the safety boundaries Dicken recommends.
- You operate infrastructure and want a database workflow with isolated changes, controlled deploys, and reversions without data loss.

## Related talks

- [Scaling AI Agents Without Breaking Reliability](https://aietalks.com/talks/scaling-ai-agents-without-breaking-reliability) (Preeti Somal, Temporal, 15:01)
- [Rethinking how we Scaffold AI Agents](https://aietalks.com/talks/rethinking-how-we-scaffold-ai-agents) (Rahul Sengottuvelu, Ramp, 16:32)
- [Deterministic Infra for Non-Deterministic AI Agents](https://aietalks.com/talks/deterministic-infra-for-non-deterministic-ai-agents) (Nishant Gupta, Meta Superintelligence Labs, 07:14)
- [Designing AI-Intensive Applications](https://aietalks.com/talks/designing-ai-intensive-applications) (swyx, 13:02)
- [Navigating AI's Frontier in 2025](https://aietalks.com/talks/navigating-ais-frontier-in-2025) (Grace Isford, Lux Capital, 17:55)
