# From Vibes to Production: Evaluating and Shipping AI Agents That Work 201

Laurie Voss, Arize AI | AI Engineer World's Fair 2026 | 42:17

Source: https://www.youtube.com/watch?v=F0TNSmbo5hE
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/from-vibes-to-production-evaluating-and-shipping-ai-agents-that-work
Published: 2026-10-05
Tags: evals, observability, reliability, tracing

## TL;DR
- Agent traces are the source of truth for behavior because the same input can produce different outputs and different execution paths.
- Evals compress large numbers of traces into scores and explanations, but Signal groups recurring failures and proposes fixes when eval volume exceeds human review capacity.
- The live workshop adds Arize AX tracing to Wonder Toys, finds that its search tool lacks price filtering, and uses a coding agent to implement the fix.

## Summary
Laurie Voss argues that agent development needs an automated improvement loop. Traditional software can often be understood by reading its code, but agents make decisions across multiple turns and may behave differently on repeated runs. Traces therefore show what the application actually did. At production scale, people cannot inspect millions of traces individually, and evals eventually create another review bottleneck. Voss presents Signal as a further layer that groups repeated failures across traces and evals, then suggests actions such as creating a GitHub issue, adding examples to a dataset, creating an evaluator, or opening a pull request. In the workshop, a coding agent instruments the Wonder Toys shopping app, pulls traces from Arize AX, identifies search-quality problems, and adds missing price filtering. Voss is direct about trust and regression concerns. Existing evals should check that a change preserves behavior that already works, and Signal itself is monitored and evaluated.

## Key ideas
### Agent traces show behavior that code alone cannot predict
[02:57](https://www.youtube.com/watch?v=F0TNSmbo5hE&t=177s)
Voss says agents are nondeterministic because they make decisions over multiple turns. The same input can produce a different output, and even the same output may come from a different path. That makes the code an incomplete source of truth for an agent. Traces capture each LLM call, tool call, and agent turn, with nested structure, inputs, outputs, cost, timing, and metadata. In the Wonder Toys example, a trace shows the full workflow rather than only the final shopping response. This lets a developer see wasteful behavior, such as ten turns for a simple request, an unnecessary turn, an expensive operation, or a task that takes twenty minutes instead of one.

### Evals reduce trace volume but eventually create another human bottleneck
[06:06](https://www.youtube.com/watch?v=F0TNSmbo5hE&t=366s)
At production scale, Voss says an application can generate millions of traces, so reading them one at a time cannot give a useful overall view. Evals compress traces into a score and an explanation of why a run passed or failed. Those explanations can be read by a coding agent and fed into an improvement loop. The problem moves when applications and agents run at greater speed. There can be more eval results than a person can review. Voss frames this as a progression from traces, to evals, to a further layer that groups recurring failures across many eval results.

### Signal groups repeated failures and turns them into possible actions
[09:02](https://www.youtube.com/watch?v=F0TNSmbo5hE&t=542s)
Voss introduces Signal as an agent inside Arize AX that continuously reads traces, finds patterns, and suggests fixes. It runs every six hours by default, with a configurable schedule. Instead of presenting thousands of separate failures, it can identify a recurring problem inside that larger set. Its actions include creating a GitHub issue with the failing trace and proposed repair, adding a trace to an evaluation dataset, creating an evaluator, or opening a pull request when the repository is connected. Voss describes this as moving from software that reports what happened toward software that proposes changes based on a definition of good behavior.

### Instrumenting an agent can require configuration rather than new application code
[16:31](https://www.youtube.com/watch?v=F0TNSmbo5hE&t=991s)
The Wonder Toys application initially sends no traces to Arize AX. Voss asks a coding agent to add observability. The application uses the OpenAI Agents SDK, which already has OpenInference instrumentation, so the agent mainly enables it and supplies the Arize API key and endpoint. The workshop repository contains agent skills that explain how to work with Arize AX. After restarting the local app, Voss sees the Wonder Toys project in AX with inputs, outputs, agent requests, and other trace details. The demonstration shows that an agent can add the observability setup without Voss manually writing the instrumentation code.

### Trace inspection finds search-quality failures that ordinary errors miss
[24:41](https://www.youtube.com/watch?v=F0TNSmbo5hE&t=1481s)
The coding agent pulls traces from AX through its command line tools and summarizes span kinds, statuses, and errors. Voss points out that the application has no error spans because requests return HTTP 200 even when the search result is poor. The analysis finds that 42% of searches returned zero results, that searches with all arguments set to null produced meaningless results, and that exact age constraints could narrow searches too far. The agent also identifies the missing price filter after Voss asks it to investigate the request for toys under $6. This is a quality diagnosis, not an exception diagnosis.

### A missing price filter can be fixed and checked through a new trace
[30:28](https://www.youtube.com/watch?v=F0TNSmbo5hE&t=1828s)
The non-pre-run Wonder Toys version has no price-filtering capability. Voss asks the coding agent to add it. The agent finds a previously solved copy of the application beside the working directory and copies the solution, which Voss accepts as a live-demo shortcut. Voss then asks for all toys under $9. The response contains toys under that price, and the trace records both the request and the returned products. The resulting loop has four stages in the demonstration: add observability, inspect what happened, identify a concrete problem, and ask an agent to change the application.

### Regression evals are needed before automated fixes can be trusted
[36:06](https://www.youtube.com/watch?v=F0TNSmbo5hE&t=2166s)
Audience questions focus on deployment, compliance, and the risk that a fix could damage behavior that already works. Voss says a production system should have regression evals for expected good behavior. Those evals can show whether a change preserves existing behavior. She demonstrates creating an evaluator that checks that Wonder Toys never gives instructions for building a bomb. Signal also has its own evaluation suite. Its work is traced, and LLM judges assess whether its proposed pull requests are good. Signal does not automatically create every evaluator yet; Voss says teams should review the suggestion because evaluators take time and money to run.

## Notable quotes
- "Agents are non-deterministic." (03:21)
- "Instead, the traces are your source of truth." (03:41)
- "You become a bottleneck and you can't possibly get an overall sense of what all of your traces are doing by reading them one at a time." (06:46)
- "42% of searches have returned zero results." (28:29)
- "A production system should have a set of regression evals which are things that you are expecting your agent to already be good at." (37:47)

## Tools & references mentioned
- Arize AI
- Arize AX
- Signal
- Wonder Toys
- OpenAI Agents SDK
- OpenInference
- Claude Code
- GitHub
- Vim
- OpenAI

## Who should watch
- You are shipping an agent and need to understand failures that do not appear as exceptions or HTTP errors.
- Your team has more traces or eval results than people can review manually, and you want recurring failures grouped into actionable work.
- You are considering automated code changes and need a practical approach to observability, regression evals, and approval before deployment.

## Editor's note

Laurie Voss says regression evals are needed before automated fixes can be trusted, because a change can damage behavior that already works. Kitaru replays a recorded agent run with the same inputs and tool responses, so a changed model, prompt, or code can be checked against a run that previously worked without touching real systems.

Written by the AIE Talks editors (the Kitaru team), not by the speaker.

## Related talks

- [From Vibes to Production: Evaluating and Shipping AI Agents That Work 101](https://aietalks.com/talks/from-vibes-to-production-evaluating-and-shipping-ai-agents-that-work-101) (Laurie Voss, Arize AI, 1:51:27)
- [How We Built an Agent That Improves Itself](https://aietalks.com/talks/how-we-built-an-agent-that-improves-itself) (Zubin Aysola, Weights & Biases, 17:06)
- [Taming Rogue AI Agents with Observability-Driven Evaluation](https://aietalks.com/talks/taming-rogue-ai-agents-with-observability-driven-evaluation) (Jim Bennett, Galileo, 16:15)
- [The Future of Evals: From LLM as a Judge to Agent as a Judge](https://aietalks.com/talks/the-future-of-evals-from-llm-as-a-judge-to-agent-as-a-judge) (Aparna Dhinakaran, Arize AI, 06:06)
- [From Agent Traces to Agent Simulations](https://aietalks.com/talks/from-agent-traces-to-agent-simulations) (Rustem Feyzkhanov, Snorkel AI, 20:24)
