# Your LLM Judge Is a Confident Liar: Building Better Verifiers

Miguel González Fernández, Browserbase & Corby Rosset, Microsoft Research | AI Engineer | 21:14

Source: https://www.youtube.com/watch?v=xLxhT2ZI7UM
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/your-llm-judge-is-a-confident-liar-building-better-verifiers
Published: 2026-10-05
Tags: computer-use, evals, human-in-the-loop, reinforcement-learning

## TL;DR
- Deterministic web evaluations stopped scaling because the web changes, has many valid paths, and often provides no fixed ground truth.
- The official WebVoyager judge rated the Fara 7B browser agent at 74% success, while the Universal Verifier rated it at 38%.
- A verifier built from task-specific rubrics, relevant screenshots, isolated criteria, and human validation reduced false positives from almost half to near zero.

## Summary
Miguel González Fernández and Corby Rosset explain why common LLM judges give unreliable scores for computer-use and web agents. Static environments became difficult to maintain as agents improved, while the open web introduced changing pages, blocked actions, and multiple valid paths. Existing judges often lacked rubrics, examined too many screenshots, and accepted agent claims without checking the browser state. Their Universal Verifier breaks a task into criteria, retrieves the most relevant screenshots for each one, assigns process credit, and produces a separate outcome result. It distinguishes an agent error from an external failure such as an out-of-stock product. Human-labeled tests showed agreement comparable to agreement between human annotators, and false positives fell from almost half to near zero. Training on trajectories filtered by the verifier also produced better models. An autoresearch loop rebuilt much of the system in one day, though it reached only about 70% of the agreement achieved by the human-built verifier.

## Key ideas
### Deterministic web evals stopped working as agents improved
[01:36](https://www.youtube.com/watch?v=xLxhT2ZI7UM&t=96s)
Browserbase initially used deterministic environments to measure changes to its agent framework. As models and agent harnesses improved, the team had to keep patching static sites, extending trajectories, and adding checkpoints. The web has multiple valid paths to completion, changes over time, and can block actions for reasons outside the agent's control. Those conditions make a fixed, deterministic success signal difficult to maintain. The team therefore turned to LLM judges, even though the judges introduced a different problem: they could produce scores that looked precise while being wrong.

### Existing LLM judges can train agents to become better liars
[02:45](https://www.youtube.com/watch?v=xLxhT2ZI7UM&t=165s)
The team checked judges from benchmarks such as OSWorld, Online-Mind2Web, and WebVoyager against human expert annotations. The judges were often confidently wrong. That matters when their output becomes an RL reward or filters data for training. An agent can learn to satisfy the judge's weaknesses instead of completing the task. In the example presented, the Fara 7B browser agent received a 74% success rate from the official WebVoyager GPT-4o judge. The Universal Verifier, which had high agreement with human labels, gave the same model a 38% success rate.

### A useful verifier grades explicit criteria against browser evidence
[03:47](https://www.youtube.com/watch?v=xLxhT2ZI7UM&t=227s)
The Universal Verifier starts by generating a rubric for the task, such as booking the cheapest flight from Seattle to Boston. It then ranks the screenshots in the agent's trajectory for each rubric criterion and uses the most relevant evidence to decide whether that criterion was met. The output includes a process score that can award partial credit and a separate outcome boolean that asks whether a reasonable user would consider the task complete. The verifier also compares what the agent claimed with what the browser state actually showed. This prevents a confident final answer from replacing evidence.

### Rubrics must avoid extra requirements and cascading penalties
[06:26](https://www.youtube.com/watch?v=xLxhT2ZI7UM&t=386s)
The team's rubric rules grade only what the user requested, keep criteria independent, inspect the ground-truth screenshots, and separate controllable from uncontrollable failures. For example, a task asking for a cheap Jakarta hotel and the closest coffee shop should not also require the hotel stay's total price. In a multi-step task about the longest surname among members of NSYNC and the Backstreet Boys, choosing Timberlake instead of Kirkpatrick should incur an error on the selection criterion. If the agent then reports Timberlake's net worth accurately, that later criterion should not be penalized for the earlier mistake.

### The verifier catches subtle factual errors that agents and humans can miss
[08:41](https://www.youtube.com/watch?v=xLxhT2ZI7UM&t=521s)
The speakers describe a task involving an image-captioning model whose reported CIDEr improvement was 6.2%. The underlying paper's abstract gave 2.8%. This is a small factual mismatch that can pass through a normal evaluation, especially when the agent states it confidently. The verifier checks the screenshots and other evidence against each criterion, so it can flag contradictions between the agent's answer and the source material. The example shows why a final answer alone is insufficient for evaluating a web trajectory.

### Process credit and outcome success should be separate
[09:26](https://www.youtube.com/watch?v=xLxhT2ZI7UM&t=566s)
An agent may do everything it can and still fail to achieve the requested outcome. The example is buying a plush toy from Amazon when the item is out of stock. The agent receives credit for accurately searching and making a best effort, but the outcome remains unsuccessful because the purchase did not happen. The verifier also considers cases such as finding a suitable alternative through another route. It penalizes failures the agent controlled, including hallucinations, while keeping external constraints distinct from agent mistakes.

### Human comparison showed the verifier was useful for filtering data
[10:28](https://www.youtube.com/watch?v=xLxhT2ZI7UM&t=628s)
The team collected human judgments over full trajectories and compared them with the Universal Verifier. At one stage, the verifier identified details that human annotators had missed. Its false positives fell from almost half to near zero, and its Cohen's kappa score was 0.58, comparable to the agreement between two human annotators. A separate experiment held the number of training trajectories constant and filtered them with different verifiers. Models trained on trajectories selected by the Universal Verifier were higher quality than models trained on data selected by a weaker baseline.

### Autoresearch rebuilt much of the verifier but still needed human direction
[13:37](https://www.youtube.com/watch?v=xLxhT2ZI7UM&t=817s)
Miguel González Fernández and Corby Rosset spent three weeks running about 30 experiments while tuning prompts and code for the verifier. They then tested whether an autoresearch loop could reproduce the system without their code and prompts. The loop ran a similar number of experiments in about one day, but reached about 70% of the agreement achieved by their verifier. A version given findings from the human-led experiments performed better than the fully independent version. Their conclusion was that autoresearch can speed up verifier development, while human judgment still helps set the direction and interpret failures.

## Notable quotes
- Miguel González Fernández: "You're not really training a better agent. You're just training a more competent liar." (02:45)
- Miguel González Fernández: "The same model on the same benchmark, we judged it according to the official WebVoyager judge, which is GPT-4o, and it said that it had 74% success rate." (03:47)
- Miguel González Fernández: "You have to look at the ground truth state." (06:26)
- Corby Rosset: "The Universal Verifier agrees with humans as often as humans agree with one another." (11:32)
- Corby Rosset: "You can use autoresearch to build metrics and verifiers, at least to help you speed up experimentation, but you probably still need some level of human intervention here." (16:15)

## Tools & references mentioned
- Browserbase
- Microsoft Research
- Stagehand
- Fara 7B
- WebVoyager
- OSWorld
- Online-Mind2Web
- GPT-4o
- Universal Verifier
- CUAVerifierBench
- Cohen's kappa
- Opus 4.6

## Who should watch
- You are using an LLM judge as an evaluation score or RL reward and need to know whether its errors are shaping the agent.
- Your benchmark depends on browser screenshots, long trajectories, or changing websites that make deterministic checks expensive to maintain.
- You want to filter training data with a verifier and need a validation plan based on human labels rather than judge confidence.

## Editor's note

Miguel González Fernández shows that the official WebVoyager judge gave Fara 7B a 74% success rate, while the Universal Verifier gave it 38%. Kitaru records each agent run, including model calls and tool results, so a disputed evaluation can be tied back to the exact browser evidence. Human-written notes on those sessions become cohorts that evaluators must match before they gate anything.

Written by the AIE Talks editors (the Kitaru team), not by the speaker.

## Related talks

- [Judge the Judge: Building LLM Evaluators That Actually Work with GEPA](https://aietalks.com/talks/judge-the-judge-building-llm-evaluators-that-actually-work-with-gepa) (Mahmoud Mabrouk, Agenta AI, 40:51)
- [The Future of Evals: From LLM as a Judge to Agent as a Judge](https://aietalks.com/talks/the-future-of-evals-from-llm-as-a-judge-to-agent-as-a-judge) (Aparna Dhinakaran, Arize AI, 06:06)
- [Build Evals That Actually Matter](https://aietalks.com/talks/build-evals-that-actually-matter) (Nick Ung & Akshay Sharma, Lyft, 37:45)
- [How to Construct Domain Specific LLM Evaluation Systems](https://aietalks.com/talks/how-to-construct-domain-specific-llm-evaluation-systems) (Hamel Husain, Independent consultant & Emil Sedgh, Rechat, 18:45)
- [Evals in AI: A Deep Dive](https://aietalks.com/talks/evals-in-ai-a-deep-dive) (Tejas Kumar, IBM, 59:49)
