# Get Out of the Model's Way

Kevin Hou, Google DeepMind | AI Engineer | 19:01

Source: https://www.youtube.com/watch?v=buHC7bQE1X4
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/get-out-of-the-models-way
Published: 2026-09-27
Tags: agents, design, multi-agent

## TL;DR
- Agent products should improve as their underlying models become faster, cheaper, and more capable.
- Antigravity's agent-teams approach uses dynamically generated subagents, sidecars, and generative UI.
- A team of 93 subagents built an operating system kernel that ran Doom in 12 hours for under $1,000.

## Summary
Kevin Hou argues that agent products should give increasingly capable models room to do more of the work. He describes Antigravity's progression from an agent-first IDE to a separate agent manager, then explains why the company split those applications. The talk traces developer tools from deterministic autocomplete and chat in 2022, through agents in 2024, to agent managers in 2025 and agent teams in 2026. Hou presents three primitives for this latest stage: dynamic subagents that a lead agent creates and coordinates, sidecars that let agents respond to external triggers, and generative UI that creates task-specific views on demand. He gives examples including a photo editor, a messaging app, an operating system kernel that runs Doom, and an evaluation workflow that automated 90% of the analysis. His product lesson is direct: build around the model's changing abilities instead of preserving fixed interfaces and workflows.

## Key ideas
### An agent product should improve as its model improves
[02:58](https://www.youtube.com/watch?v=buHC7bQE1X4&t=178s)
Hou's main design principle is "scaling with intelligence." As the model gets better, the product should expose that improvement in the user's experience. He contrasts this with older developer tools, where much of the behavior was deterministic because models could only handle limited tasks. The product should be built around the frontier of the model being served. That means product teams must expect to remove familiar features when a stronger model enables a better workflow. Hou says teams will often face resistance because users prefer what they already know, and admits that product decisions will not be right every time.

### Developer tools have moved from autocomplete to agent teams
[03:20](https://www.youtube.com/watch?v=buHC7bQE1X4&t=200s)
Hou describes four stages in the tools he has worked on. In 2022, autocomplete and chat sidebars relied on embeddings, rules, files, and syntax-tree parsing. In 2024, agents introduced primitives such as MCPs, custom tools, and permission systems. In 2025, Antigravity's agent manager let users manage many agents in parallel, alongside skills, hooks, and artifacts. For 2026, Hou expects subagents, generative UI, and sidecars to define the next stage. These changes come partly from deliberate product choices and partly from using models that have acquired new abilities.

### Antigravity separates the agent manager from the IDE
[02:05](https://www.youtube.com/watch?v=buHC7bQE1X4&t=125s)
Antigravity 2.0 split the IDE and agent manager into separate applications. Hou compares the IDE to a debugger for the agent manager: useful when a user needs to inspect a deeper layer, but not required for every task. The standalone manager is intended for orchestrating agents and projects, with features including subagents, work trees, scheduled tasks, and voice mode. Hou says the team is betting on agent orchestration, which can also be called agent teams, swarms, or software factories. The split reflects his view that managing work across agents will become a main workflow rather than a feature inside an editor.

### A lead agent can create a specialized team for each task
[08:28](https://www.youtube.com/watch?v=buHC7bQE1X4&t=508s)
In Antigravity's public-preview agent teams mode, a user types /teamwork and gives the system a task. A lead agent manages a team of arbitrary size and can ask the user for more detail when needed. The lead can create roles such as frontend engineer, backend engineer, infrastructure specialist, QA worker, or designer. Each subagent is generated dynamically, can work independently, and can use a different model from the main agent. Hou says the model itself configures and prompts these workers, allowing them to run in parallel and in different secure environments such as sandboxes or remote execution systems.

### Parallel subagents can tackle work that would be too large for one agent
[10:17](https://www.youtube.com/watch?v=buHC7bQE1X4&t=617s)
Hou describes an Antigravity run that built an operating system kernel from scratch and played Doom on it. The run used 93 subagents over 12 hours, made 15,000 requests, processed two billion tokens, and cost under $1,000. He presents the example as evidence that adding more model-driven workers can make a large software task possible at a manageable cost, even though it would not be economical to build an operating system kernel every day. He also mentions a photo editor and a messaging app, each built with hundreds of subagents over almost half a day.

### Research agents can turn evaluation analysis into a parallel investigation
[12:10](https://www.youtube.com/watch?v=buHC7bQE1X4&t=730s)
DeepMind researchers used Antigravity to automate 90% of a side-by-side evaluation workflow. Instead of manually comparing rollout tables and working through notebooks, a researcher can ask about an evaluation in natural language. The agent calculates the difference and then starts a research specialist that proposes 100 hypotheses for why the difference occurred. A separate subagent investigates each hypothesis in parallel before the results are combined into a report. Hou says the workflow replaces manually assembling agent pools, judges, and data pipelines with a dynamic process grounded in skills and an understanding of Google's monorepo.

### Sidecars let agents respond to events outside the application
[15:10](https://www.youtube.com/watch?v=buHC7bQE1X4&t=910s)
Hou introduces sidecars as a plugin protocol for long-lived utility processes. A sidecar listens for events and gives the model triggers from the outside world. Examples include SMS messages, webhooks, cron jobs, and GitHub pull requests. Antigravity already uses the same primitive for scheduled tasks. Hou says Google plans to release the specification so other developers can build on it. The idea expands an agent's work from a single interaction into an ongoing process that can react when a message arrives, a scheduled time occurs, or a change appears in another system.

### Generative UI creates the interface a task needs
[16:05](https://www.youtube.com/watch?v=buHC7bQE1X4&t=965s)
Hou argues that fixed, human-written interfaces are a poor fit for agents whose tasks keep changing. Antigravity can render interfaces inline instead of relying on templates or fixed HTML files. The model can produce a Kanban board, a debugger-style timeline, charts, graphs, tables, or an interactive report when the task calls for one. Hou compares fixed controls to the plastic buttons and keyboards Steve Jobs criticized when introducing the iPhone. In his design, the interface changes with the agent's work. He says this lets users inspect results more richly than a Markdown response or a conversation alone.

## Notable quotes
- "LLMs aren't just role players anymore. They can be your star player if you build the right product around them." (00:58)
- "This means that as the model gets better, so should your product." (03:01)
- "Each subagent is dynamically generated and can operate independently." (09:06)
- "It took 93 subagents over the course of 12 hours, made 15,000 requests, two billion tokens, and it was under $1,000." (10:50)
- "We skipped the heavy infrastructure and mechanical UIs in favor of sidecars and generative UI." (17:16)

## Tools & references mentioned
- Google Antigravity
- Google DeepMind
- Gemini 3.5 Flash
- MCPs
- Windsurf
- Google I/O
- Doom
- Steve Jobs
- GitHub
- Jupyter notebooks

## Who should watch
- You are building an AI coding tool and need product primitives that can take advantage of stronger models.
- Your team is deciding whether to keep a familiar chat or editor workflow as agents become more capable.
- You want concrete examples of parallel agents for software generation or evaluation research.

## Related talks

- [The Evolution of Agentic Surfaces](https://aietalks.com/talks/the-evolution-of-agentic-surfaces) (Gagan Bhat & Isabella Kai He, Anthropic, 31:24)
- [How Google DeepMind Runs Agents at Scale](https://aietalks.com/talks/how-google-deepmind-runs-agents-at-scale) (KP Sawhney & Ian Ballantyne, Google DeepMind, 25:13)
- [The Missing Primitive for Agent Swarms](https://aietalks.com/talks/the-missing-primitive-for-agent-swarms) (Lou Bichard, Ona, 18:37)
- [Defying Gravity](https://aietalks.com/talks/defying-gravity) (Kevin Hou, Google DeepMind, 25:10)
- [Develop at Idea Velocity](https://aietalks.com/talks/develop-at-idea-velocity) (Jeffrey Lee-Chan, Snapchat, 15:28)
