Tokens Should Have Jobs

Katelyn Lesse, Anthropic, Angela Jiang, Anthropic13:21 · Sept 2026 · 3,455 views
Thumbnail for Tokens Should Have Jobs Watch on YouTube
TL;DR
  1. 1

    Assigning some tokens to advising, grading, or reflection can improve an agent's result without increasing its total token budget.

  2. 2

    On financial analysis tasks, an 80 percent accurate answer is still a failure because the analyst needs a fully correct P&L.

  3. 3

    The best strategy depends on whether a business values lower token use or a higher chance of getting a perfect answer in one run.

Summary

Katelyn Lesse and Angela Jiang argue that agent tokens should have different jobs instead of all being spent on execution. They describe three strategies: an adviser that can help an executor during a task, a grader that checks each attempt against a rubric and triggers another iteration, and a dreamer that reads execution transcripts and writes lessons to memory for the next run. In a financial analysis bench, they compare these strategies under the same budget of about 600,000 tokens. Execution scores 76, while advising reaches 89. They then change the scoring rule so anything below perfect accuracy fails. The execute baseline passes about 42 percent of runs, which makes its expected cost about 1.8 million tokens for one clean answer. Advising and grading use tokens more efficiently, while grading and dreaming are better choices when single-run reliability matters more. The speakers finish by describing how these strategies can be combined from agent primitives.

Key ideas
00:37

Agent budgets assume that every token has the same job

Most agent systems respond to a hard task by increasing the token budget or using more expensive tokens. Lesse and Jiang question the assumption behind that choice: that every token is interchangeable. In a normal setup, the agent receives a task and spends the budget executing it. Their proposal is to assign part of that budget another role. Some tokens can advise the executor, inspect its work, grade it against a rubric, or reflect on past transcripts. The point is to change how a fixed amount of computation is used, rather than only making the budget larger.

02:05

An adviser can intervene while the executor is working

The advising strategy splits the system into an executor and an adviser. The executor works on the task, then calls the adviser when it needs help deciding whether its next step is correct. Lesse and Jiang give a sales agent as an example. An adviser could help flag an overdue follow-up or a deal that is stalling. This gives the executing agent another source of judgment during the task, instead of waiting until the whole attempt is finished.

02:44

A grader turns a rubric into an iteration loop

In the grading strategy, the team defines what a good result looks like in a rubric. A grader checks each executor attempt against that rubric. If the attempt passes, the task can finish. If it does not, the executor tries again. The speakers use customer service refunds as an example. A store may have specific criteria for when a refund should be granted, and the grader can check whether the agent applied those criteria and reached the right outcome.

03:36

A dreamer writes lessons from one run into memory for the next

The dreaming strategy adds a component that inspects the executor's work and transcripts. It extracts findings and writes them to memory, which the executor can use on the next run. Lesse and Jiang connect this to recruiting, where an agent receives repeated feedback about whether a candidate is a good fit. That feedback can sharpen later attempts. Dreaming therefore spends tokens on learning from completed work instead of only producing the current answer.

05:25

The financial analysis benchmark shows an advantage at a fixed budget

The speakers build a benchmark of financial analysis tasks intended to resemble work done by an expert analyst. An initial one-shot comparison is hard to interpret because each strategy chooses how many tokens to spend. The dreaming strategy appears strong while using about 600,000 tokens, so the team fixes that amount as the budget for every strategy. With the same budget, execution improves from 15 to 76, while advising and grading move from the 60s toward the 90s. At this budget, execution scores 76 and advising scores 89, showing that token allocation affects the result.

07:07

A partly correct financial answer has no practical value

The speakers rescore the benchmark from the perspective of someone using an agent for real financial work. If an analyst asks for a profit and loss statement, an answer that is 80 percent accurate still requires the analyst to recompute it or run the task again. The analyst cannot invent a missing income or cost number. Under this scoring rule, only a fully accurate run passes, and anything below 100 percent is a failure. The execute baseline passes about 42 percent of the time, while the more complex strategies reach as high as 75 percent.

08:37

Expected cost depends on the chance of a perfect run

The team converts pass rates into the token cost of obtaining a useful answer. If execution produces a perfect answer about 40 percent of the time, a user should expect to run it roughly three times. At 600,000 tokens per run, that is about 1.8 million tokens for one perfect answer. Advising and grading are more token-efficient under this measure. The right choice still depends on the business goal. Advising fits a focus on token efficiency, while grading or dreaming fits a focus on the percentage of runs that produce a perfect answer.

10:34

Strategies can be composed from agent primitives

The speakers describe a harness for individual agents and a higher-level layer for coordinating multiple agents. That layer can connect an executor with an adviser, send the result to a grader for verification, and then pass the completed work to a dreamer so the next run can improve. They say these combinations can be built from the available primitives, and that teams can create jobs beyond advising, grading, and dreaming. Their longer-term goal is for models and platforms to construct these strategies dynamically for a task.

"If you get really smart about having your tokens do these different jobs and try these different strategies, you're very very likely to be able to get a better outcome for the task at hand within a fixed budget."Angela Jiang10:46
Who should watch
  • You are tuning an agent by increasing its context or token budget and want to test whether splitting that budget across different roles works better.
  • Your application produces outputs where partial accuracy is still unusable, such as financial statements or policy-sensitive customer service decisions.
  • You need to choose between lower expected token cost and a higher probability of a correct answer on each individual run.