Why 99% Accurate Browser Agents Still Fail

Derek Meegan, Browserbase16:49 · Oct 2026 · 4,451 views
Thumbnail for Why 99% Accurate Browser Agents Still Fail Watch on YouTube
TL;DR
  1. 1

    A browser agent with a 99% success rate per step succeeds only about 36% of the time across a 100-step trajectory.

  2. 2

    Production systems should measure whether a transaction completed, allow retries, and verify a concrete artifact such as a receipt or confirmation record.

  3. 3

    The most reliable architecture removes deterministic work from the model, including downloads, authentication, verification, and parts of the critical path.

Summary

Derek Meegan explains why a browser-agent demo can work once while a production system fails across thousands of unattended runs. Every browser step adds cost and failure risk, while the customer receives value only when the full transaction finishes. With a 99% success rate per step, a 100-step workflow succeeds only about 36% of the time. Meegan argues that teams should first optimize performance, then treat cost and maintainability as engineering problems. He recommends measuring success per transaction, allowing retries, and checking concrete outputs such as receipts or records. His example is an insurance portal that downloads an explanation of benefits. The final design uses a deterministic download tool, OCR verification, separate authentication, and an agent skill that describes the critical path. The agent still handles ambiguity, but it no longer controls operations that can be made repeatable and directly testable.

Key ideas
00:00

Production browser agents face unattended variation

A demo usually runs on one site, once, with a person watching it. Production means thousands of unattended runs while the website can change and the model can make an occasional bad decision. Meegan frames the engineering problem around the gap between these conditions. The system has to keep completing customer work when the environment changes or a particular run goes badly. The insurance example later in the talk applies this standard to a real transactional workflow rather than a one-off demonstration.

01:27

Browser agents choose actions from several browser representations

Meegan describes the model as receiving page state, the overall goal, and the actions already taken, then selecting an action from a probability distribution. That action becomes a structured tool call executed against the browser. The browser can be represented through HTML and the accessibility tree, screenshots, or dynamic code using JavaScript or CDP. He says production systems generally use textual representations, screenshots, or a combination. Dynamic execution lets the agent write code against the browser, while tailored tools constrain what it can do.

03:17

Browser trajectories range from fixed tasks to arbitrary requests

Meegan places browser workflows on a spectrum. Purpose-built trajectories complete one known task. Just-in-time automations handle an arbitrary task supplied by the user. Between them are workflows where browser activity supports a larger goal such as deep research or competitor analysis. He focuses on transactional workflows because they complete a unit of work from start to finish. These workflows need the flexibility to handle ambiguity, but their outcome is still a concrete completed transaction.

04:32

Long browser trajectories compound risk because there is no partial credit

Each step in a browser trajectory adds cost and another chance of failure. The customer gets value only after every required step finishes. Meegan calls this the difference between continuous cost accumulation and terminal value realization. In his example, every step succeeds independently 99% of the time, but a 100-step trajectory succeeds only about 36% of the time. A capable model does not remove the problem by itself, because the number of opportunities for failure continues to grow.

08:42

Transaction success should be measured after retries

Meegan distinguishes per-run success from per-transaction success. If an agent succeeds on a run 50% of the time, allowing up to four retries for one transaction raises the transaction success rate to 94% in his example. The customer cares whether the requested work was completed, rather than how many attempts the system used. The outcome should be tied to a concrete artifact, such as a confirmation email for a bill payment, an order ID or receipt, or a new record that can be queried.

07:37

Performance comes before cost and maintainability

Meegan puts performance first because the system must complete the requested task before other improvements matter. Once performance is acceptable, cost and maintainability become engineering problems. Cost includes model usage, infrastructure and compute, and extra integration or tooling. Maintenance includes observability into both model decisions and browser events, developer time when the system breaks, and repeated evaluation as sites, models, and tasks change.

10:46

Retries do not remove the need to account for hostile web conditions

The web can resist automation through antibot systems and other unfriendly behavior. The task itself can change, although Meegan says improved models may become better at understanding intent and correcting course. A site may later provide an API, creating a better method than browser automation, although he says that will not arrive soon for most browser-agent use cases. Models can also stray from the intended path because their behavior is nondeterministic from run to run.

12:06

Deterministic tools reduce the model's responsibility

In the insurance-portal example, the initial agent logs in and downloads an explanation of benefits. Meegan then combines the page interaction, file retrieval, and storage access into one deterministic download tool. An OCR tool extracts entities from the downloaded document and compares them with the system of record. Authentication is moved into a reusable function because that workflow does not change, reducing the number of decisions the model must make for each business operation.

14:36

Skills give the agent a narrower critical path

Meegan describes agent skills as standard operating procedures written for an agent. A skill tells the agent where it is in the workflow and what step comes next, which reduces ambiguity in its decision process. The final architecture uses a serverless authentication function, an agent runtime with the skill, a deterministic download function, and a verification tool. The agent retains responsibility for the ambiguous parts, while repeatable operations are handled outside the model and the result can be checked directly.

"In the end, this system takes steps out of the model's responsibility that simply do not need to be in the responsibility of the agent."16:03
Who should watch
  • You are moving a browser-agent prototype toward unattended transactional work and need a way to reason about full-workflow reliability.
  • Your workflow can produce a receipt, record, confirmation, or other artifact that should be checked after the run.
  • You are deciding which browser operations belong in the model and which should become reusable tools or fixed functions.