From Vibes to Production: Evaluating and Shipping AI Agents That Work 201

Laurie Voss, Arize AI42:17 · Oct 2026 · 3,913 views
Thumbnail for From Vibes to Production: Evaluating and Shipping AI Agents That Work 201 Watch on YouTube
TL;DR
  1. 1

    Agent traces are the source of truth for behavior because the same input can produce different outputs and different execution paths.

  2. 2

    Evals compress large numbers of traces into scores and explanations, but Signal groups recurring failures and proposes fixes when eval volume exceeds human review capacity.

  3. 3

    The live workshop adds Arize AX tracing to Wonder Toys, finds that its search tool lacks price filtering, and uses a coding agent to implement the fix.

Summary

Laurie Voss argues that agent development needs an automated improvement loop. Traditional software can often be understood by reading its code, but agents make decisions across multiple turns and may behave differently on repeated runs. Traces therefore show what the application actually did. At production scale, people cannot inspect millions of traces individually, and evals eventually create another review bottleneck. Voss presents Signal as a further layer that groups repeated failures across traces and evals, then suggests actions such as creating a GitHub issue, adding examples to a dataset, creating an evaluator, or opening a pull request. In the workshop, a coding agent instruments the Wonder Toys shopping app, pulls traces from Arize AX, identifies search-quality problems, and adds missing price filtering. Voss is direct about trust and regression concerns. Existing evals should check that a change preserves behavior that already works, and Signal itself is monitored and evaluated.

Key ideas
02:57

Agent traces show behavior that code alone cannot predict

Voss says agents are nondeterministic because they make decisions over multiple turns. The same input can produce a different output, and even the same output may come from a different path. That makes the code an incomplete source of truth for an agent. Traces capture each LLM call, tool call, and agent turn, with nested structure, inputs, outputs, cost, timing, and metadata. In the Wonder Toys example, a trace shows the full workflow rather than only the final shopping response. This lets a developer see wasteful behavior, such as ten turns for a simple request, an unnecessary turn, an expensive operation, or a task that takes twenty minutes instead of one.

06:06

Evals reduce trace volume but eventually create another human bottleneck

At production scale, Voss says an application can generate millions of traces, so reading them one at a time cannot give a useful overall view. Evals compress traces into a score and an explanation of why a run passed or failed. Those explanations can be read by a coding agent and fed into an improvement loop. The problem moves when applications and agents run at greater speed. There can be more eval results than a person can review. Voss frames this as a progression from traces, to evals, to a further layer that groups recurring failures across many eval results.

09:02

Signal groups repeated failures and turns them into possible actions

Voss introduces Signal as an agent inside Arize AX that continuously reads traces, finds patterns, and suggests fixes. It runs every six hours by default, with a configurable schedule. Instead of presenting thousands of separate failures, it can identify a recurring problem inside that larger set. Its actions include creating a GitHub issue with the failing trace and proposed repair, adding a trace to an evaluation dataset, creating an evaluator, or opening a pull request when the repository is connected. Voss describes this as moving from software that reports what happened toward software that proposes changes based on a definition of good behavior.

16:31

Instrumenting an agent can require configuration rather than new application code

The Wonder Toys application initially sends no traces to Arize AX. Voss asks a coding agent to add observability. The application uses the OpenAI Agents SDK, which already has OpenInference instrumentation, so the agent mainly enables it and supplies the Arize API key and endpoint. The workshop repository contains agent skills that explain how to work with Arize AX. After restarting the local app, Voss sees the Wonder Toys project in AX with inputs, outputs, agent requests, and other trace details. The demonstration shows that an agent can add the observability setup without Voss manually writing the instrumentation code.

24:41

Trace inspection finds search-quality failures that ordinary errors miss

The coding agent pulls traces from AX through its command line tools and summarizes span kinds, statuses, and errors. Voss points out that the application has no error spans because requests return HTTP 200 even when the search result is poor. The analysis finds that 42% of searches returned zero results, that searches with all arguments set to null produced meaningless results, and that exact age constraints could narrow searches too far. The agent also identifies the missing price filter after Voss asks it to investigate the request for toys under $6. This is a quality diagnosis, not an exception diagnosis.

30:28

A missing price filter can be fixed and checked through a new trace

The non-pre-run Wonder Toys version has no price-filtering capability. Voss asks the coding agent to add it. The agent finds a previously solved copy of the application beside the working directory and copies the solution, which Voss accepts as a live-demo shortcut. Voss then asks for all toys under $9. The response contains toys under that price, and the trace records both the request and the returned products. The resulting loop has four stages in the demonstration: add observability, inspect what happened, identify a concrete problem, and ask an agent to change the application.

36:06

Regression evals are needed before automated fixes can be trusted

Audience questions focus on deployment, compliance, and the risk that a fix could damage behavior that already works. Voss says a production system should have regression evals for expected good behavior. Those evals can show whether a change preserves existing behavior. She demonstrates creating an evaluator that checks that Wonder Toys never gives instructions for building a bomb. Signal also has its own evaluation suite. Its work is traced, and LLM judges assess whether its proposed pull requests are good. Signal does not automatically create every evaluator yet; Voss says teams should review the suggestion because evaluators take time and money to run.

"You become a bottleneck and you can't possibly get an overall sense of what all of your traces are doing by reading them one at a time."06:46
Who should watch
  • You are shipping an agent and need to understand failures that do not appear as exceptions or HTTP errors.
  • Your team has more traces or eval results than people can review manually, and you want recurring failures grouped into actionable work.
  • You are considering automated code changes and need a practical approach to observability, regression evals, and approval before deployment.