Deterministic web evaluations stopped scaling because the web changes, has many valid paths, and often provides no fixed ground truth.
2
The official WebVoyager judge rated the Fara 7B browser agent at 74% success, while the Universal Verifier rated it at 38%.
3
A verifier built from task-specific rubrics, relevant screenshots, isolated criteria, and human validation reduced false positives from almost half to near zero.
Summary
Miguel González Fernández and Corby Rosset explain why common LLM judges give unreliable scores for computer-use and web agents. Static environments became difficult to maintain as agents improved, while the open web introduced changing pages, blocked actions, and multiple valid paths. Existing judges often lacked rubrics, examined too many screenshots, and accepted agent claims without checking the browser state. Their Universal Verifier breaks a task into criteria, retrieves the most relevant screenshots for each one, assigns process credit, and produces a separate outcome result. It distinguishes an agent error from an external failure such as an out-of-stock product. Human-labeled tests showed agreement comparable to agreement between human annotators, and false positives fell from almost half to near zero. Training on trajectories filtered by the verifier also produced better models. An autoresearch loop rebuilt much of the system in one day, though it reached only about 70% of the agreement achieved by the human-built verifier.
Deterministic web evals stopped working as agents improved
Browserbase initially used deterministic environments to measure changes to its agent framework. As models and agent harnesses improved, the team had to keep patching static sites, extending trajectories, and adding checkpoints. The web has multiple valid paths to completion, changes over time, and can block actions for reasons outside the agent's control. Those conditions make a fixed, deterministic success signal difficult to maintain. The team therefore turned to LLM judges, even though the judges introduced a different problem: they could produce scores that looked precise while being wrong.
Existing LLM judges can train agents to become better liars
The team checked judges from benchmarks such as OSWorld, Online-Mind2Web, and WebVoyager against human expert annotations. The judges were often confidently wrong. That matters when their output becomes an RL reward or filters data for training. An agent can learn to satisfy the judge's weaknesses instead of completing the task. In the example presented, the Fara 7B browser agent received a 74% success rate from the official WebVoyager GPT-4o judge. The Universal Verifier, which had high agreement with human labels, gave the same model a 38% success rate.
A useful verifier grades explicit criteria against browser evidence
The Universal Verifier starts by generating a rubric for the task, such as booking the cheapest flight from Seattle to Boston. It then ranks the screenshots in the agent's trajectory for each rubric criterion and uses the most relevant evidence to decide whether that criterion was met. The output includes a process score that can award partial credit and a separate outcome boolean that asks whether a reasonable user would consider the task complete. The verifier also compares what the agent claimed with what the browser state actually showed. This prevents a confident final answer from replacing evidence.
Rubrics must avoid extra requirements and cascading penalties
The team's rubric rules grade only what the user requested, keep criteria independent, inspect the ground-truth screenshots, and separate controllable from uncontrollable failures. For example, a task asking for a cheap Jakarta hotel and the closest coffee shop should not also require the hotel stay's total price. In a multi-step task about the longest surname among members of NSYNC and the Backstreet Boys, choosing Timberlake instead of Kirkpatrick should incur an error on the selection criterion. If the agent then reports Timberlake's net worth accurately, that later criterion should not be penalized for the earlier mistake.
The verifier catches subtle factual errors that agents and humans can miss
The speakers describe a task involving an image-captioning model whose reported CIDEr improvement was 6.2%. The underlying paper's abstract gave 2.8%. This is a small factual mismatch that can pass through a normal evaluation, especially when the agent states it confidently. The verifier checks the screenshots and other evidence against each criterion, so it can flag contradictions between the agent's answer and the source material. The example shows why a final answer alone is insufficient for evaluating a web trajectory.
Process credit and outcome success should be separate
An agent may do everything it can and still fail to achieve the requested outcome. The example is buying a plush toy from Amazon when the item is out of stock. The agent receives credit for accurately searching and making a best effort, but the outcome remains unsuccessful because the purchase did not happen. The verifier also considers cases such as finding a suitable alternative through another route. It penalizes failures the agent controlled, including hallucinations, while keeping external constraints distinct from agent mistakes.
Human comparison showed the verifier was useful for filtering data
The team collected human judgments over full trajectories and compared them with the Universal Verifier. At one stage, the verifier identified details that human annotators had missed. Its false positives fell from almost half to near zero, and its Cohen's kappa score was 0.58, comparable to the agreement between two human annotators. A separate experiment held the number of training trajectories constant and filtered them with different verifiers. Models trained on trajectories selected by the Universal Verifier were higher quality than models trained on data selected by a weaker baseline.
Autoresearch rebuilt much of the verifier but still needed human direction
Miguel González Fernández and Corby Rosset spent three weeks running about 30 experiments while tuning prompts and code for the verifier. They then tested whether an autoresearch loop could reproduce the system without their code and prompts. The loop ran a similar number of experiments in about one day, but reached about 70% of the agreement achieved by their verifier. A version given findings from the human-led experiments performed better than the fully independent version. Their conclusion was that autoresearch can speed up verifier development, while human judgment still helps set the direction and interpret failures.
"You can use autoresearch to build metrics and verifiers, at least to help you speed up experimentation, but you probably still need some level of human intervention here."Corby Rosset16:15
Who should watch
You are using an LLM judge as an evaluation score or RL reward and need to know whether its errors are shaping the agent.
Your benchmark depends on browser screenshots, long trajectories, or changing websites that make deterministic checks expensive to maintain.
You want to filter training data with a verifier and need a validation plan based on human labels rather than judge confidence.