Tracing gives an agent development team the data needed to find where quality is declining.
2
Topics can expose recurring tasks, sentiment, and failures that existing evaluations do not test for.
3
Production findings become useful when they return to datasets, evaluations, code changes, and reviewable pull requests.
Summary
Doug Guthrie presents Braintrust observability through a hands-on support-agent workshop. He starts with tracing an agent built on the OpenAI Agents SDK, then connects production traces to offline evaluations and online scoring. Code-based scores handle checks such as tool order, while LLM judges assess more subjective qualities. Guthrie stresses that judges need human review and ongoing calibration. Braintrust Topics adds another signal by grouping traces into tasks, sentiment, issues, and custom facets. In the workshop, a custom facet finds order-record lookup failures and other workflow problems in imported support traces. Guthrie then uses SQL, Loop, and a coding-agent skill to inspect representative failures, add examples to evaluation datasets, change the repository, run evaluations, and prepare a pull request. The final section covers deployment models, sampling, project scope, trace access, and remote evaluations, which let collaborators change prompts and models in a playground while execution stays on the evaluation server.
Tracing is the foundation for finding agent quality problems
Guthrie says an agent must expose the steps it takes, including tool inputs and outputs. Without that data, the system is a black box and developers cannot locate where quality declines. The workshop uses the OpenAI Agents SDK with a Braintrust tracing processor, although Guthrie says the platform also supports other frameworks, providers, languages, OpenTelemetry, and custom orchestration. The trace gives later tools and coding agents enough context to inspect failures rather than guessing from the final answer alone.
Production traffic should feed a development flywheel
Guthrie describes a loop in which a team changes an agent, tests the change, deploys it, observes production, and turns the resulting evidence into another development change. Production data must produce usable insight at scale, since a person cannot inspect every interaction. The workshop demonstrates this by scoring incoming traces, finding patterns, selecting representative examples, returning them to evaluation datasets, and checking later changes against those cases. The aim is to make production feedback part of ordinary engineering work.
Scores need different methods and human calibration
Guthrie distinguishes fast code-based scores from LLM judges. Code can check a schema or whether one tool ran before another. An LLM judge can assess a more subjective interaction against stated criteria. He also describes scores that use code-based branching while invoking an LLM underneath. Judges need human review because they are not a set-and-forget system. Guthrie recommends harsh binary scores early on, strong test cases, examples in judge prompts, and sometimes a stronger model for scoring than for generation.
Topics find patterns outside the existing evaluation set
Existing evaluations cover failure modes the team already knows about. Guthrie says Topics helps find the unknown cases by grouping patterns in production data. Built-in facets cover tasks, issues, and sentiment, while custom facets let a team define its own question. Topics can reveal failures, feature requests, or new ways users are trying to use an agent. Guthrie presents this as a way to uncover blind spots and bring new cases into the development process.
The Topics pipeline turns traces into searchable labels
The pipeline preprocesses a trace into a smaller representation of the user, assistant, and tool interaction. A facet then summarizes that representation. Braintrust creates an embedding of the summary and uses it to build a topic map that labels and classifies traces. Those labels can then be queried by a person or a coding agent. Guthrie says the pipeline uses Braintrust-hosted models and is designed to run across large volumes of traces. Custom preprocessors and facets can adapt the pipeline to data outside the default thread representation.
Online automations attach scores to incoming traces
The workshop publishes scores from code and connects them to automations. An automation can run on a whole trace, a particular span, or a grouped conversation. Grouping matters when turns are stored as separate traces, since a conversation ID can connect them. Sampling controls model use and cost. Guthrie also shows filters that restrict scoring to matching traces or spans. This lets a team apply online judges to selected production traffic instead of treating every trace identically.
Representative failures should become regression cases
Loop can query trace data with SQL and gather context for a question such as what should improve in a support agent. Guthrie says the useful workflow is to find an interesting trace, inspect its inputs, outputs, and spans, and add the case to an evaluation dataset. The new case then becomes part of future comparisons. As the agent changes, the dataset and scores should evolve with it, so known failures do not return unnoticed.
Coding-agent skills can automate investigation and code review
The Braintrust CLI lets a coding agent query logs with SQL, inspect spans, create datasets, run evaluations, and work with scores. Guthrie installs skills that teach agents how to use those operations. In the demonstrated workflow, the agent searches 968 support traces, finds scored failures and escalation-tool errors, inspects representative traces, adds regression examples, and proposes code and score changes. A GitHub action can package the investigation, reasoning, evaluation evidence, and links into a pull request for a human reviewer.
Remote evaluations let other collaborators change prompts safely
Guthrie describes remote evaluations that expose an evaluation to a playground. A local or hosted server accepts a request for the evaluation, runs the code and tools on that server, and streams the results back. Parameters can expose system prompts, models, or prompts for different agents. Product managers and subject-matter experts can try changes without requiring an engineer to build a separate application, while the actual execution remains on the evaluation server.
"You should have a really simple and easy to create flywheel that allows you to develop some type of change, fix something within your agent, test whether or not that change regressed the agent in some way or improved it."08:07
Who should watch
You are instrumenting an agent but still inspect production failures manually and need a repeatable path from traces to evaluations.
Your existing eval set covers known failure modes, while production users keep finding cases you did not anticipate.
Your team uses coding agents and wants them to query trace data, add regression examples, and prepare changes for review.