AI-generated pull requests rose from under 1% to about a quarter of reviewed pull requests over the previous year.
2
Agent-written pull requests were broadly comparable to human-written ones on reverts, issue severity, and review rounds, although their failure patterns differed.
3
Code validation should test the user contract, the risk of future violations, and whether the change matches the author's intent.
Summary
Daksh Gupta uses Greptile's review data to examine whether autonomous coding agents can produce useful pull requests in enterprise codebases. Greptile reviews more than a million pull requests each month from companies including NVIDIA, Coinbase, and Scale. Gupta estimates that about a quarter of reviewed pull requests are now largely or completely AI-generated, compared with under 1% a year earlier. Across revert rates, revert rates by pull-request size, issue severity, and review rounds, agent-written changes were broadly similar to human-written changes. The important differences were in failure modes. Claude produced more SQL injection issues, Devin produced fewer auth bypasses, and Cursor produced more N+1 queries. Gupta then describes a validation approach based on codebase context, blast-radius analysis, and sandboxed browser agents. These agents install dependencies, run the application, mock inputs, and try to break the change before merge.
Autonomous coding moved from assistance to complete pull requests
Gupta traces AI coding from tab completion in 2022 to multifile editing in 2024 and autonomous agents in 2025. These agents receive a task, modify a codebase, and open an entire pull request. He describes December as a major change, when coding agents became sufficiently autonomous for people to run them repeatedly without writing each change themselves. The central question was whether this workflow could work inside companies with real customers and large, commercially important codebases.
AI-generated pull requests now make up about a quarter of the data
Greptile initially found fewer than 1% of pull requests with an AI tool in the GitHub author field. Gupta then used additional signals, including co-author footers in pull-request descriptions and branch-name prefixes associated with tools such as Codex. Those signals suggested that about a quarter of the pull requests Greptile reviewed in a month were largely or completely generated by AI. Looking back over the prior year, the share had risen from under 1%. Gupta says the increase looked continuous rather than tied to individual model releases.
Agent and human pull requests had similar revert rates
Gupta uses reverts as one measure of a bad pull request. In his data, Codex pull requests were reverted about once per thousand pull requests, Devin pull requests about three and a half times per thousand, and human pull requests about two and a half times per thousand. He then checked whether agents were simply handling smaller, easier tasks. Revert rates showed little relationship to pull-request size for either group. That reduced the evidence for a quality gap between human and agent changes.
Issue severity was broadly similar, with humans producing more P0 issues
Greptile classifies findings as P0, P1, and P2 issues, which Gupta uses as another quality measure. Three of the four agents tested produced P0 issues at a lower rate than humans. The broader pattern for P1 and P2 findings was also close between the groups. Gupta's conclusion is that human-generated and agent-generated pull requests were broadly equal on these measures. He does not claim that agents fail in the same way as humans.
The tools differed in the kinds of bugs they introduced
Gupta searches Greptile's comments for specific failure patterns, including SQL injection, auth bypass, and N+1 queries. He normalizes the chart so that 1x means the human rate. Claude was about 1.5 times more likely than humans to produce a SQL injection issue. Devin was about half as likely to produce an auth bypass issue. Cursor produced N+1 queries at a much higher rate. Similar overall quality therefore hid meaningful differences in how each agent failed.
The median Greptile user creates about 50 pull requests a month. The 90th percentile creates about 500, and the 99th percentile is in the thousands. Gupta says the people at the high end produce pull requests at roughly the rate at which they develop new ideas. Manual review, ordinary testing, and external QA do not naturally scale to that volume. This leads Greptile to frame validation around what must be established before a change is merged safely.
Validation should test contracts, future risk, and intent
Gupta reduces code validation to three questions. Does the change violate the application's user contract? Does it make a future violation of that contract more likely? Does it do what the author intended? Greptile approaches these questions by giving agents the surrounding codebase, rather than reviewing only the changed lines. It also runs the application in a sandbox, installs dependencies, mocks inputs, and uses browser agents to interact with the result and try to expose failures.
"The second one, does it increase the propensity of a future violation of the user contract, whatever the user contract might be for that application?"11:17
Who should watch
You are deciding whether autonomous coding agents can be used in a production codebase with real customers.
You need data about agent-written pull requests rather than anecdotes about individual coding sessions.
You are building code review or validation systems for a team whose pull-request volume is too high for manual review alone.