# The 5 Levels of Self-Driving Production

Eric Schwartz, Traversal | AI Engineer World's Fair 2026 | 18:33

Source: https://www.youtube.com/watch?v=y-OVWZD4j6U
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/the-5-levels-of-self-driving-production
Published: 2026-10-02
Tags: agents, debugging, observability, reliability, tracing

## TL;DR
- Coding agents speed up development while leaving teams with more production code, more complexity, and more troubleshooting work.
- Root cause analysis requires tracing causal relationships across services and data, which observability dashboards and general-purpose LLMs do not reliably provide.
- Self-driving production progresses from manual war rooms to systems that diagnose incidents, apply fixes, and verify those fixes without paging engineers.

## Summary
Eric Schwartz argues that coding agents have shortened the development phase of software engineering while increasing the amount of code and complexity that teams must troubleshoot in production. Observability tools can identify broken systems and correlated signals, but they do not usually explain the root cause or recommend a fix. Schwartz describes root cause analysis as a causal problem because an incident may require several hops across many services and large data sets. He presents five levels of self-driving production, from fully manual war rooms through rule-based automation and service-specific agents to enterprise-wide diagnosis and closed-loop fixes. Case studies from PepsiCo and American Express show how Traversal applies alert intelligence and incident analysis. He closes with five questions for evaluating an AI SRE, including whether it can access all production data, model relationships, improve without constant maintenance, and perform fast multi-hop searches.

## Key ideas
### Coding agents move engineering effort toward troubleshooting
[01:02](https://www.youtube.com/watch?v=y-OVWZD4j6U&t=62s)
Schwartz divides software engineering into system design, development, and troubleshooting. Coding agents have made development much faster, allowing people without formal computer science training to contribute production code. In the enterprises Traversal works with, that speed also produces more code and more complex environments. Engineers understand less of the code being pushed into production, so they spend more time fixing issues instead of designing systems. Schwartz says tools such as Claude Code, Codex, and Cursor will increase this pressure because they make it easier to generate still more code.

### Troubleshooting consumes substantial engineering time
[02:32](https://www.youtube.com/watch?v=y-OVWZD4j6U&t=152s)
Schwartz describes troubleshooting as a large operational cost. He cites estimates that enterprises spend upwards of $400 billion a year on it, while 40% of executives say it is a problem for their teams. Engineers lose more than seven hours each week troubleshooting while on call, according to the figures he presents. The result is less time for creative architecture and design. More alerts and more complicated environments make it harder for teams to keep up with the systems they are responsible for.

### Observability tools show symptoms without explaining causes
[03:12](https://www.youtube.com/watch?v=y-OVWZD4j6U&t=192s)
Tools such as Datadog, Elastic, Splunk, and ServiceNow can show that something is broken and identify related failures. Schwartz says they do not usually explain the root cause or tell engineers what to do next. Customers he heard from while working on observability products at ServiceNow described the gap directly: the tools told them what was broken, but not why or how to fix it. Adding more dashboards does not solve the problem when the environment keeps growing in complexity.

### Root cause analysis requires causal reasoning across many hops
[04:01](https://www.youtube.com/watch?v=y-OVWZD4j6U&t=241s)
Schwartz says even LLMs struggle to identify root causes. A failing checkout API may require five to ten hops across dozens of services and petabytes of data before an engineer reaches the actual cause, such as an expired TLS certificate. A single experienced SRE may hold enough tribal knowledge in a small environment, but a Fortune 50 or Fortune 100 company has no one person with the full context. The result is a war room with many engineers, each bringing only one part of the system's history and behavior.

### Self-driving production closes the incident loop
[06:26](https://www.youtube.com/watch?v=y-OVWZD4j6U&t=386s)
Traversal defines self-driving production as a closed-loop system. When an issue appears, the system finds the incident, identifies its root cause, proposes or applies a fix, and verifies that the fix worked. The intended result is that teams can keep building without being paged for every problem. Schwartz presents this as a spectrum of autonomy rather than a state that most companies have already reached.

### The five autonomy levels range from war rooms to verified fixes
[07:06](https://www.youtube.com/watch?v=y-OVWZD4j6U&t=426s)
Level zero is fully manual work in Slack war rooms or Zoom calls. Level one adds rules and structured automations, but these break down when an alert describes a novel situation without a runbook. Levels two and three add more automation, including agents that debug a particular service or alert set. Level four covers an entire environment with hundreds of services, repositories, and log indexes. Level five adds the ability to diagnose issues, apply fixes, and verify them across a large production environment.

### Alert intelligence can reduce the noise reaching on-call engineers
[11:05](https://www.youtube.com/watch?v=y-OVWZD4j6U&t=665s)
At PepsiCo, Traversal is used with applications supporting the supply chain, including the movement of finished products from warehouses to trucks and retailers. Before the engagement, teams received thousands or tens of thousands of alerts each week, and one engineer could have a backlog of 700 alerts. Traversal filters and prioritizes the alerts, identifies which ones deserve investigation, and points to ways to reduce noise. Those actions include changing alert rules, dismissing alerts that do not need immediate attention, and creating tickets for technical debt.

### Incident analysis can limit the number of people pulled into an outage
[12:55](https://www.youtube.com/watch?v=y-OVWZD4j6U&t=775s)
Schwartz describes American Express incidents that previously paged five to ten teams and 20 to 50 engineers. Resolution could take an hour, or sometimes hours or days. Traversal now acts as the first responder when an incident is declared. Within three minutes, it posts a detailed root cause analysis in the incident's Slack channel and can update ServiceNow tickets. Teams may page no one or only one or two teams to verify the findings, rather than bringing every involved group into the incident.

### An AI SRE must search complete data and trace distant causes quickly
[14:44](https://www.youtube.com/watch?v=y-OVWZD4j6U&t=884s)
Schwartz gives five questions for evaluating an AI SRE. It should see all production data, search very large data sets without excessive cost or disruption, map relationships among the entities it sees, improve without engineers maintaining extensive system knowledge, and make multiple hops to find a distant cause. He says searches that take more than five minutes can lose the patience of an on-call team, which may then return to its existing habits.

## Notable quotes
- "Fundamentally the belief that we have at Traversal is that root cause analysis is not an observability problem. It's a causal problem." (05:46)
- "What that means is a closed loop system where as issues break the AI system finds the issue or the incident or the alert, finds the root cause of what caused that issue in the first place, puts up a fix, verifies that it's been fixed and the loop is closed." (06:26)
- "The second the incident is declared Traversal is dispatched within 3 minutes we'll post a very detailed root cause analysis in the Slack channel where that incident is being managed." (13:33)
- "Can it make multiple hops? Can it find the non-trivial, non-obvious root cause that's far away from the initial symptom in a matter of minutes?" (16:14)

## Tools & references mentioned
- Traversal
- Claude Code
- Codex
- Cursor
- Datadog
- Elastic
- Splunk
- ServiceNow
- Google SRE Handbook
- Anthropic
- PepsiCo
- American Express
- DigitalOcean
- Capital One

## Who should watch
- Your team is generating more production code with coding agents, but on-call work and incident debugging are taking more of the week.
- You are evaluating AI SRE products and need questions that test access to data, causal search, system relationships, and speed.
- Your organization has large incident war rooms and wants to understand what automation can do before attempting fully closed-loop production fixes.

## Editor's note

Eric Schwartz says observability tools can show that something is broken without explaining the root cause or what engineers should do next. When an agent run contributes to an incident nobody can explain, Kitaru records every model call and tool result, then replays the run with the same inputs and responses after a change to the model, prompt, or code. That gives engineers a repeatable run to inspect instead of relying on a one-off production failure.

Written by the AIE Talks editors (the Kitaru team), not by the speaker.

## Related talks

- [Production software keeps breaking and it will only get worse](https://aietalks.com/talks/production-software-keeps-breaking-and-it-will-only-get-worse) (Anish Agarwal & Matthew Schoenbauer, Traversal, 18:13)
- [Shipping complex AI applications](https://aietalks.com/talks/shipping-complex-ai-applications) (Giran Moodley, Braintrust & Mayank Soni & Oussama Hafferssas, Trainline, 1:38:34)
- [The 6 Pillars of an Agentic Harness for Production](https://aietalks.com/talks/the-6-pillars-of-an-agentic-harness-for-production) (Varun Krovvidi, Resolve AI, 21:05)
- [Production Evals For Agentic AI Systems](https://aietalks.com/talks/production-evals-for-agentic-ai-systems) (Nishant Gupta, Meta Superintelligence Labs, 08:12)
- [From Blind Spots to Merged PRs: Continuous Agentic Performance Optimization](https://aietalks.com/talks/from-blind-spots-to-merged-prs-continuous-agentic-performance-optimization) (May Walter, Hud, 22:46)
