How We Built an Agent That Improves Itself

Zubin Aysola, Weights & Biases17:06 · Sept 2026 · 7,756 views
Thumbnail for How We Built an Agent That Improves Itself Watch on YouTube
TL;DR
  1. 1

    ARIA uses the same byte-for-byte agent in production and offline simulations so evaluation results remain comparable.

  2. 2

    The team turns production traces, including both failures and successes, into YAML evaluation tasks and runs them repeatedly.

  3. 3

    ARIA can inspect a trace, identify a missing `weave.log` SDK call, create a candidate fix, and compare it with the production variant.

Summary

Zubin Aysola explains how Weights & Biases builds and evaluates ARIA, its software agent for work on the Weights & Biases platform. The team keeps production and research code identical, synchronizes them, and logs both environments through Weave. Production traces can then become offline tasks for testing new agent variants. The evaluation harness supports YAML-defined configurations, simulated users, parallel runs, and an unconstrained sandbox. A task can be scored by whether it passes, or by comparing two variants. The team has 886 tasks, including multi-turn user simulations, and adds both production failures and successes to the task set. In the live demo, ARIA converts a production trace into a regression task, finds that the sandbox is missing a `weave.log` call, writes a candidate change, and compares its result with production. Aysola is honest that automation does not remove the need for engineering judgment. The tools reduce manual benchmark writing and leave more time for deciding how the system should improve.

Key ideas
01:31

Benchmarks, evaluations, and agent configurations change together

Aysola says the main evaluation problem is that benchmarks, evaluations, agents, and their configurations are all changing together. For a dynamic system, measurement has to show how the agent performs in both production and offline environments. Weave gives the team one way to observe both sides. The production team logs traces in the same format used offline, so production examples can be moved into the research environment. This lets the team investigate errors and improve the agent without losing the connection between what was tested and what users experienced.

05:38

Production and research run the same agent

The research and production code are byte-for-byte identical. A four-hour synchronization job moves changes between the environments so researchers do not create drift while cutting new agent variants or skills. Aysola describes the two sides as mirrors of the same stack, with a deployment layer and an offline benchmarking layer. The team exposes commands that generate score trajectories, which show how the agent rolled out over time and provide more detail than a final score alone.

06:41

The team generates many traces before deciding how to use them

Aysola's approach is to generate a large number of agent traces and then inspect the behavior in those traces. The team can manually review a task or ask ARIA to review a rollout, identify what went wrong or well, and reinforce the desired behavior through prompting or another change. This treats traces as material for understanding emergent behavior. The goal is to align the agent in specific directions based on observed actions rather than rely only on a single aggregate benchmark result.

07:15

YAML variants make agent experiments repeatable

The harness is designed to work across different models and to keep the surrounding software stack consistent. It handles tasks such as context preparation, compaction, and assembling UI payloads. Agent configurations are defined in YAML, which makes it possible to create many mutations of the same setup and run them in parallel. Aysola connects this to a simple training principle: running more experiments gives the team more evidence about which changes help.

07:57

The sandbox lets ARIA run its own research loop

The team gives the agent an unconstrained sandbox because it wants ARIA to discover useful ways to work, including parallel executions of its own research loop. Aysola says he asked ARIA to create this environment while preparing the presentation. The evaluation pipeline loads a YAML configuration, hydrates it with live data, sets up an environment, rehydrates runtime values that cannot be encoded in YAML, runs the byte-for-byte production agent, and scores the result. Some environments can involve production data, training logs, or simulated GPU executions.

10:58

Tasks simulate both simple and multi-turn users

Evaluation tasks are YAML specifications with a starting condition, user configuration, and desired ending condition. Some tasks are a single text instruction. Others use a language model with a user persona that asks questions in a particular order, creating a simulated multi-turn interaction. The team has 886 tasks, grouped by level, and exposes them to the product team so people can judge whether they reflect the behaviors the benchmarks should measure. The tasks are run repeatedly to produce more trajectories.

10:13

ARIA uses pass-fail and relative scoring

ARIA scores tasks in two ways. Normative scoring asks whether the task passed. Relative scoring compares variants with different behavior, such as one that asks the user questions and one that does not. The team can then see which version behaves better for the chosen comparison. Aysola spends substantial time thinking about the health of these evaluations and the gap between offline results and production behavior. The scoring system is part of the agent-building work, rather than a final reporting step.

12:00

Every production result feeds the evaluation flywheel

The team turns every production miss and every production success into a task for the offline framework. Aysola examines production and offline traces, uses internal tools to understand their behavior, and feeds the resulting tasks back into agent development. In the demo, ARIA takes a production trace, creates a regression task, runs the production and candidate variants, and writes a report. It finds that the sandbox was not calling the `weave.log` SDK call correctly, then injects a prompt or skill change and checks the candidate against production.

14:43

Automation still requires engineering judgment

Aysola says it is easy to put an agent into an automatic mode and let it handle tasks, but that does not remove the need to think about how the system should improve. He says he has been asking Claude to write his code, yet the difficult work remains deciding what changes matter and what guard rails the agent needs. ARIA's self-improvement loop gives him more time to reason about those decisions while the system runs evaluations, traces experiments, and compares variants.

"Using these tools to improve themselves is very valuable because you get to spend more of your time in the gray of actually trying to think how we make this system better."14:59
Who should watch
  • You are building an agent and need production traces to become repeatable offline evaluation tasks.
  • Your research and production versions drift apart, making it hard to tell whether a benchmark result reflects the deployed system.
  • You want an agent to run experiments and propose fixes while you retain responsibility for evaluation design and guard rails.