Most software engineers work in large, old codebases that support everyday services such as banking, insurance, transport, and aviation.
2
AI coding agents are increasing code volume while causing duplicated code, inconsistent standards, brittle dependencies, and new vulnerabilities across large repositories.
3
Enterprise agents need code graphs, search, and compiler-accurate analysis to see and change code across thousands of repositories.
Summary
Dan Adler argues that AI coding agents are making large codebases harder to own. The software behind banks, airlines, insurance companies, cars, and other services is often decades old, spread across thousands of repositories, and maintained by many engineers. Agents are now producing code faster than teams can review it. That increases duplication, inconsistent coding standards, brittle dependencies, and vulnerabilities. Adler says the main limit at enterprise scale is context. An agent can search a repository, but it cannot search code that is outside its view. Sourcegraph's proposed answer is a code graph that combines search with compiler-accurate analysis. Adler also introduces Agentic Batch Changes, which can apply a change across hundreds or thousands of repositories from one prompt, using agents where judgment is needed and deterministic scripts where consistency matters. A Mercari example found 80 more potential vulnerabilities after two known cases were used as a starting point.
Adler says most software engineers work in large enterprises with long-lived codebases. He gives examples of code handling bank transactions, insurance reimbursement, warehouse deliveries, ride estimates, airplane radar, payroll, and office air conditioning. These systems are old, complex, and spread across thousands of repositories. They were built over many years by thousands of engineers, so their owners have to maintain software that is difficult to understand as a whole.
AI-generated code is increasing the maintenance burden
Coding agents are producing more code, faster than teams have seen before. Adler describes this as a flood that engineers must review and keep healthy. Code review agents and code health tools may help with local problems, but he says the underlying codebases are still beginning to decay. The faster production of features and patches creates more work for the people responsible for the whole system.
Codebase decay appears as inconsistency and hidden risk
Adler points to several forms of decay: different agents apply different coding standards, duplicate code spreads even when an existing library could be reused, and cross-service dependencies become more brittle. Small deviations can create hidden issues across the codebase. He also says agents are uncovering new vulnerabilities every day, which increases the need for constant oversight of legacy systems.
Adler says owning a codebase is harder because the volume of code has grown beyond what teams or tools can hold in context. Large systems may contain millions of lines and tens of thousands of repositories. They cannot fit into a context window, and cloning and processing the whole system in real time is not practical. The same tools that help engineers write code faster can create the conditions for these systems to fail.
Enterprise leaders need to know what generated code does
Adler recounts hearing a developer at a top-ten car manufacturer say, "I don't know what this code does. AI wrote it for me." He connects that statement to the risk of using generated code in vehicle autopilot systems, where thousands of engineers may be working across a large organization. His point is that faster code generation does not remove the owner's responsibility to understand and maintain the resulting system.
Agents need infrastructure that can expose the whole codebase
Adler says a bank executive described the scale problem plainly: Claude Code might make a change, but the bank has 90,000 repositories where the change may be needed. Agents rely heavily on search to build an understanding of code. Repository instructions and agent files only help so much when important code is outside the agent's view. Adler says enterprise tooling must make code visible across hundreds, thousands, or hundreds of thousands of repositories.
A code graph combines search with structural analysis
Sourcegraph's approach uses a code graph that includes search and compiler-accurate output. Adler presents this graph as a foundation for agents that need to understand how code is connected and how a change will affect the wider system. He says visibility is infrastructure because agents cannot safely change code they cannot locate or understand.
Agentic Batch Changes applies and audits changes at scale
Adler introduces Sourcegraph's Agentic Batch Changes beta. It lets an owner start changes across hundreds or thousands of repositories with one prompt. The system can roll changes out iteratively, respond to CI status and pull request comments, use coding agents where judgment is needed, and use deterministic scripts where the same operation must be applied consistently. It also tracks the work and provides auditability so owners can check that all required locations were covered.
The Mercari example starts with two known issues and finds 80 more
An early-access user at Mercari used the product to patch a GitHub code-injection issue involving environment variables. After running it on two repositories where the issue was known, the user asked it to explore the rest of the codebase. It found 80 other potential vulnerabilities. A deterministic script then patched the configuration files consistently across the affected locations. Adler presents this as an example of applying one verified change across a large set of repositories.