From Vibes to Production: Evaluating and Shipping AI Agents That Work 101

Laurie Voss, Arize AI1:51:27 · Oct 2026 · 4,194 views
Thumbnail for From Vibes to Production: Evaluating and Shipping AI Agents That Work 101 Watch on YouTube
TL;DR
  1. 1

    Agents need traces before evals because traces show every model decision, tool call, input, and output.

  2. 2

    A useful evaluation stack combines deterministic checks, focused LLM judges, and human annotations calibrated against one another.

  3. 3

    Production failures should become labeled datasets that drive controlled experiments and coding-agent fixes.

Summary

Laurie Voss presents a practical workflow for evaluating an AI agent that researches financial data and writes reports. She starts with tracing, then reads actual traces before defining success criteria or writing evaluators. The workshop combines a cheap deterministic ticker check, built-in faithfulness grading, and a custom actionability judge. A correctness judge initially rejects every report because it lacks the live research context used by the agent. Voss shows how human annotations, held-out examples, precision, recall, and judge-bias checks help determine whether an evaluator is trustworthy. Failing traces become datasets for prompt changes and controlled experiments on the same cases. The loop continues in production through sampled online evaluations, monitoring, and recurring failure themes passed to a coding agent. The practical message is direct: define what good means, inspect how the agent actually behaved, and only then automate the grading.

Key ideas
07:42

Traces show the agent's decisions while evals judge the result

Voss gives a simple mental model: traces are like logs for AI, while evals are tests for outputs that can vary from run to run. A trace contains spans for agent turns, LLM calls, and tool calls, along with inputs, outputs, timing, token counts, and metadata. This matters because an agent can search the wrong subject, call the wrong tool, pass bad arguments, or get stuck in a loop before producing a polished answer. A final report alone hides those intermediate decisions. Traces expose where the execution went wrong, while evaluators decide whether the resulting behavior met the requirement.

14:27

Agents need checks for intermediate behavior as well as final answers

Agent evaluation is harder than judging one LLM call because agents make a sequence of decisions. Voss says teams need to check whether the agent used the right tool, passed the right arguments, stopped at the right time, and handled handoffs correctly. She gives a Tesla example where a web search returns information about the 18th-century inventor rather than the car company, producing a polished but irrelevant report. She also cites Anthropic's flight-booking benchmark, where an agent found a valid route around a supposed restriction and was graded as failing. Evaluators must allow creatively correct solutions while still catching wrong outcomes.

17:19

Capability evals become regression evals as the agent improves

Voss separates capability evaluations from regression evaluations. A capability eval asks whether the agent can perform a task at all, so low scores are expected while the team is still improving the behavior. Once the agent can perform that task reliably, the same evaluation becomes a regression test that should pass every time or nearly every time. This gives teams a growing suite of checks for behavior they have already achieved. LLM judges also provide explanations, which can identify what was missing and give a coding agent concrete evidence for the next prompt or code change.

38:57

Read traces and define success before writing evaluators

Voss calls trace review the most important practice in the workshop. Teams should read a dozen or more executions end to end, focus on failures, and ask what specifically broke before automating a grade. For the financial analyst, the requirements include using the correct ticker, including recent actionable financial data, giving recommendations, and separating forward-looking analysis from historical summary. These requirements should involve product managers, QA, support, and other people with domain knowledge. Synthetic queries can provide an initial dataset, but they need varied phrasing and edge cases, and production traffic should replace imagined cases as soon as it exists.

49:16

Cheap deterministic checks provide the first layer of coverage

The first evaluator checks whether the requested stock ticker appears in the report. Voss argues that an LLM judge is unnecessary when the requirement is a specific, deterministic string check. The evaluator runs instantly, costs nothing, and returns a binary label and score. In the workshop, 12 cases pass and one fails because a request about Amazon AWS produces a report focused only on AWS and omits the Amazon ticker. Voss also says deterministic checks can validate JSON, length limits, required fields, forbidden phrases, database values, or API results. They should grade the produced result rather than enforce one exact tool-use path.

58:14

The judging context must match the agent's research context

The correctness evaluator rejects all 13 reports because it judges live financial research against the judge model's own knowledge. The agent used current web searches, including information from 2026, while the judge's knowledge stopped earlier. The judge therefore treats current facts as future information it cannot verify. Voss changes to a faithfulness evaluator and supplies the research context collected by the agent. With the same source material available to the judge, the results split into six faithful reports and seven unfaithful reports. The failure is useful because it measures whether the report stayed grounded in the evidence actually gathered.

01:06:25

Focused rubrics make LLM judge failures easier to diagnose

For actionability, Voss recommends a custom rubric with an explicit judge role, observable pass and fail criteria, tagged inputs and outputs, examples, and simple binary choices. A report should contain specific recommendations and forward-looking analysis rather than vague claims that investors should consider various factors. The rubric comes from observed trace failures, so it tests problems the application actually has. Voss prefers one evaluator per dimension instead of a single evaluator that grades accuracy, tone, completeness, policy, and formatting together. A focused failure tells the team which requirement needs work. She also recommends asking the judge to explain its reasoning before assigning a label.

01:20:00

Human annotations expose judge errors and rubric ambiguity

Voss treats an LLM judge as a classifier and compares its labels with human annotations on the same reports. Human reviewers create a golden dataset by applying concrete criteria in the interface, without writing code. She recommends splitting labeled data into development and held-out test sets, using the development examples to revise the rubric and the held-out examples to check generalization. In the workshop, the human labels are deliberately randomized, producing a 46% agreement rate across 13 examples. The point is to inspect disagreements, read the judge's explanations, and tighten ambiguous criteria. Precision measures how often negative judgments are correct, while recall measures how many genuinely bad reports the judge catches.

01:30:27

Failing traces should drive controlled experiments and production monitoring

Voss turns failed evaluations into datasets and runs the changed agent on those same cases. A coding agent can read the failure explanations and revise prompts based on specific missing requirements, such as absent recommendations or unsupported risks. The comparison must also preserve passing cases so a fix does not create regressions. Experiments hold the inputs and evaluators constant while changing the agent, which makes score differences easier to interpret despite nondeterministic web searches. In production, sampled online evaluators grade new traces, monitors surface recurring failures, and those failures become new regression cases. Voss advises giving coding agents both requirements and failure themes so they fix the underlying behavior instead of merely gaming the eval.

"The explanation is what makes evals into a useful debugging tool and not just a scoreboard."19:37
Who should watch
  • You are shipping an agent after checking a few outputs by hand and need a way to catch regressions before users report them.
  • Your agent uses live retrieval or tools, and a final-answer grader lacks the context needed to judge whether the result is grounded.
  • You are building LLM judges and need human comparison data, held-out tests, and focused rubrics instead of one opaque score.