Autonomous agents can produce far more code without producing a similar increase in shipped software because human review remains the bottleneck.
2
Passing tests does not show that a change is mergeable, since maintainability, scope, regressions, safety, and outside context remain hard to check.
3
Human review is moving up the stack into the design of review harnesses, benchmarks, rules, and production monitoring.
Summary
Laurie Voss argues that AI agents have made code generation cheap while leaving human review at roughly its old speed. She cites data showing 741% more code but only 30% more shipped software after developers adopted autonomous agents. Reading larger diffs harder does not solve the problem, since reviewer effectiveness drops sharply beyond about 400 lines. Automated review is already common, with GitHub Copilot and Cursor running large-scale systems, but their designs rely on repeated passes, suspicion, tool use, and human acceptance data. Tests alone are a weak proxy for mergeability. Experiments from METR, Cognition, Carlini, Bun, and OpenAI show that agents can pass tests while missing maintainability, hidden dependencies, unsafe code, or prompt injection. Voss's practical advice is to build a review harness that encodes company context and definitions of good, then use production behavior as the final check.
Agent productivity has exposed human review as the bottleneck
Voss cites a study of more than 100,000 GitHub developers. After developers turned on autonomous agents, they wrote 741% more code, but only 30% more software shipped. The route to production still passes through humans, especially code review. Generation has become cheaper while deciding whether to trust a change remains expensive, particularly for sensitive code with a large blast radius.
Reviewing larger diffs makes human reviewers less effective
Voss uses a Cisco study covering 2,500 reviews and 3.2 million lines of code. Reviewers stopped finding defects effectively after reading more than 400 lines in one sitting, and their effectiveness fell off sharply above 450 lines per hour. At that pace, a 10,000-line agent pull request would take three or four working days for a real review. Developers can now run many agents at once, so simply asking people to review harder leads to burnout rather than more throughput.
Passing tests is a poor substitute for mergeability
METR asked maintainers of open-source projects to review pull requests that had already passed SWE-bench. The maintainers considered the changes mergeable only about half the time. The failures involved code quality and changes that quietly broke things outside the test suite, rather than basic test correctness. Cognition's FrontierCode benchmark asks whether a maintainer would merge the result and grades behavioral correctness, regressions, safety, scope discipline, test quality, and maintainability.
Mergeability benchmarks would become training signals
Voss explains why coding models improved quickly on code generation. Compilers and test suites provide cheap verification, so model builders can train against them. A reliable mergeability benchmark would immediately become another training signal. Its criteria would include maintainability, scope discipline, regression safety, and other judgments that humans currently make. The people who define today's review standard would therefore influence the default behavior of future coding models.
Production review systems reduce false positives through repeated passes
GitHub Copilot's reviewer has completed 60 million reviews and accounts for more than one in five reviews on GitHub, according to Voss. Cursor's first reviewer ran eight passes over each diff and shuffled reviewer order because the order changed results. The goal was to filter false positives, since developers ignore a reviewer that frequently flags correct code. Cursor later made the system reason over diffs, call tools, and assume that code may contain a problem instead of accepting it by default.
Cursor's reviewer can spawn a fix agent from its own findings. The system does not only report a bug; it writes a patch and returns a diff for approval. Cursor also wants the reviewer to run code to check whether its own report is real. Voss describes this as a narrowing boundary between reviewing and rewriting. CodeRabbit, Greptile, and Graphite use related approaches, while their success is measured by whether humans accept their suggestions.
Removing humans from the loop still requires humans to build the system
Nicholas Carlini's agents built a C compiler from scratch in Rust and compiled the Linux kernel without a human approving each change. Voss points out that humans wrote the tests, feedback systems, and checking infrastructure. Bun's agent-built Zig-to-Rust port passed 99.8% of its test suite, but reviewers found 13,044 unsafe blocks, compared with about 74 in a comparable human-written Rust codebase. Tests can check public behavior while missing assertions about memory safety.
Humans remain where context, security, and accountability matter
Tests cannot tell whether a module has outside users or whether an undocumented cron job depends on it. Human checkpoints survive where correctness is hard to check cheaply, where the blast radius is large, and where someone must put their name on the result. Automated reviewers also remain vulnerable to prompt injection. Voss cites a study in which vulnerable code disguised by an innocent commit message fooled an autonomous reviewer in 88% of attempts, while human reviewers passed only 35% of those attempts.
The practical replacement for manual review is a review harness
Voss ends by telling teams to stop treating individual pull requests as the main review abstraction. Human judgment should go into a reliable harness that encodes definitions of good, company context, and domain knowledge. Agents can then work against those rules and evaluations. Review itself has not disappeared in the examples she discusses. It has been rebuilt as a system whose benchmark, classifier, rubric, test suite, and evaluation still need human oversight.