# AI Coding Agents Are Breaking Big Codebases

Dan Adler, Sourcegraph | AI Engineer World's Fair 2026 | 12:23

Source: https://www.youtube.com/watch?v=Bdrs3uAX0_M
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/ai-coding-agents-are-breaking-big-codebases
Published: 2026-10-04
Tags: code-generation, enterprise, search, security

## TL;DR
- Most software engineers work in large, old codebases that support everyday services such as banking, insurance, transport, and aviation.
- AI coding agents are increasing code volume while causing duplicated code, inconsistent standards, brittle dependencies, and new vulnerabilities across large repositories.
- Enterprise agents need code graphs, search, and compiler-accurate analysis to see and change code across thousands of repositories.

## Summary
Dan Adler argues that AI coding agents are making large codebases harder to own. The software behind banks, airlines, insurance companies, cars, and other services is often decades old, spread across thousands of repositories, and maintained by many engineers. Agents are now producing code faster than teams can review it. That increases duplication, inconsistent coding standards, brittle dependencies, and vulnerabilities. Adler says the main limit at enterprise scale is context. An agent can search a repository, but it cannot search code that is outside its view. Sourcegraph's proposed answer is a code graph that combines search with compiler-accurate analysis. Adler also introduces Agentic Batch Changes, which can apply a change across hundreds or thousands of repositories from one prompt, using agents where judgment is needed and deterministic scripts where consistency matters. A Mercari example found 80 more potential vulnerabilities after two known cases were used as a starting point.

## Key ideas
### Large codebases run much of the world
[00:35](https://www.youtube.com/watch?v=Bdrs3uAX0_M&t=35s)
Adler says most software engineers work in large enterprises with long-lived codebases. He gives examples of code handling bank transactions, insurance reimbursement, warehouse deliveries, ride estimates, airplane radar, payroll, and office air conditioning. These systems are old, complex, and spread across thousands of repositories. They were built over many years by thousands of engineers, so their owners have to maintain software that is difficult to understand as a whole.

### AI-generated code is increasing the maintenance burden
[01:58](https://www.youtube.com/watch?v=Bdrs3uAX0_M&t=118s)
Coding agents are producing more code, faster than teams have seen before. Adler describes this as a flood that engineers must review and keep healthy. Code review agents and code health tools may help with local problems, but he says the underlying codebases are still beginning to decay. The faster production of features and patches creates more work for the people responsible for the whole system.

### Codebase decay appears as inconsistency and hidden risk
[02:57](https://www.youtube.com/watch?v=Bdrs3uAX0_M&t=177s)
Adler points to several forms of decay: different agents apply different coding standards, duplicate code spreads even when an existing library could be reused, and cross-service dependencies become more brittle. Small deviations can create hidden issues across the codebase. He also says agents are uncovering new vulnerabilities every day, which increases the need for constant oversight of legacy systems.

### The volume of code is the central problem
[03:57](https://www.youtube.com/watch?v=Bdrs3uAX0_M&t=237s)
Adler says owning a codebase is harder because the volume of code has grown beyond what teams or tools can hold in context. Large systems may contain millions of lines and tens of thousands of repositories. They cannot fit into a context window, and cloning and processing the whole system in real time is not practical. The same tools that help engineers write code faster can create the conditions for these systems to fail.

### Enterprise leaders need to know what generated code does
[04:31](https://www.youtube.com/watch?v=Bdrs3uAX0_M&t=271s)
Adler recounts hearing a developer at a top-ten car manufacturer say, "I don't know what this code does. AI wrote it for me." He connects that statement to the risk of using generated code in vehicle autopilot systems, where thousands of engineers may be working across a large organization. His point is that faster code generation does not remove the owner's responsibility to understand and maintain the resulting system.

### Agents need infrastructure that can expose the whole codebase
[06:06](https://www.youtube.com/watch?v=Bdrs3uAX0_M&t=366s)
Adler says a bank executive described the scale problem plainly: Claude Code might make a change, but the bank has 90,000 repositories where the change may be needed. Agents rely heavily on search to build an understanding of code. Repository instructions and agent files only help so much when important code is outside the agent's view. Adler says enterprise tooling must make code visible across hundreds, thousands, or hundreds of thousands of repositories.

### A code graph combines search with structural analysis
[08:56](https://www.youtube.com/watch?v=Bdrs3uAX0_M&t=536s)
Sourcegraph's approach uses a code graph that includes search and compiler-accurate output. Adler presents this graph as a foundation for agents that need to understand how code is connected and how a change will affect the wider system. He says visibility is infrastructure because agents cannot safely change code they cannot locate or understand.

### Agentic Batch Changes applies and audits changes at scale
[09:21](https://www.youtube.com/watch?v=Bdrs3uAX0_M&t=561s)
Adler introduces Sourcegraph's Agentic Batch Changes beta. It lets an owner start changes across hundreds or thousands of repositories with one prompt. The system can roll changes out iteratively, respond to CI status and pull request comments, use coding agents where judgment is needed, and use deterministic scripts where the same operation must be applied consistently. It also tracks the work and provides auditability so owners can check that all required locations were covered.

### The Mercari example starts with two known issues and finds 80 more
[10:21](https://www.youtube.com/watch?v=Bdrs3uAX0_M&t=621s)
An early-access user at Mercari used the product to patch a GitHub code-injection issue involving environment variables. After running it on two repositories where the issue was known, the user asked it to explore the rest of the codebase. It found 80 other potential vulnerabilities. A deterministic script then patched the configuration files consistently across the affected locations. Adler presents this as an example of applying one verified change across a large set of repositories.

## Notable quotes
- "The software that runs the world is not pretty. It's not new. It's not clean. It's how everything actually works." (00:35)
- "The call is coming from inside the house." (03:57)
- "Sure, Claude Code can make this change, but I have 90,000 repositories to make it in." (06:41)
- "You can't grep what you literally cannot see." (07:44)
- "Visibility is infrastructure." (09:18)

## Tools & references mentioned
- Sourcegraph
- Agentic Batch Changes
- Claude Code
- OpenAI
- Anthropic
- Cursor
- Mercari
- GitHub
- npm

## Who should watch
- You own or maintain software spread across many repositories and need to understand what is deployed before making a broad change.
- Your team is using coding agents, but reviews are becoming harder because generated code duplicates existing logic or follows inconsistent standards.
- You build developer tools or agent infrastructure and need to support search, analysis, and controlled changes across large codebases.

## Related talks

- [Code Generation and Maintenance at Scale](https://aietalks.com/talks/code-generation-and-maintenance-at-scale) (Morgante Pell, Grit, 18:54)
- [How Coding Agents Change Software Development Forever](https://aietalks.com/talks/how-coding-agents-change-software-development-forever) (Hailong Zhang, 08:50)
- [Beyond the Prototype: Using AI to Write High-Quality Code](https://aietalks.com/talks/beyond-the-prototype-using-ai-to-write-high-quality-code) (Josh Albrecht, Imbue, 17:59)
- [Making Codebases Agent Ready](https://aietalks.com/talks/making-codebases-agent-ready) (Eno Reyes, Factory, 15:33)
- [No Vibes Allowed: Solving Hard Problems in Complex Codebases](https://aietalks.com/talks/no-vibes-allowed-solving-hard-problems-in-complex-codebases) (Dex Horthy, HumanLayer, 20:31)
