# The State of AI in Software Development: Data from 400+ Orgs

Justin Reock, DX | AI Engineer | 19:09

Source: https://www.youtube.com/watch?v=Se8jHLliLXE
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/the-state-of-ai-in-software-development-data-from-400-orgs
Published: 2026-09-30
Tags: evals, reliability, team-adoption

## TL;DR
- Deployment frequency is rising, but AI has made change failure rates much more volatile across companies.
- AI use is associated with larger pull requests, lower change confidence, and a 10% drop in the perceived ability to deliver incrementally.
- The median increase in PR throughput was 7.7%, because code generation addressed only 14% to 16% of the software value stream.

## Summary
Justin Reock presents DX research covering about 200,000 engineers and explains how AI is changing software development. Deployment frequency is increasing, but this does not show whether teams are creating more value. Change failure rates have become unusually volatile, while maintainability has improved and confidence in changes has fallen. Pull requests grew from about 44 to 72 lines, making review and rollback harder. Juniors use AI more often, while staff-plus engineers save similar amounts of time with fewer tokens. Reock recommends measuring AI through utilization, impact, and cost, then comparing AI cohorts against established productivity and quality measures. He also argues that platform readiness depends on documentation, modular code, reliable CI, and non-flaky tests. His central point is that code generation was never the main bottleneck. Organizations should apply AI across the wider software lifecycle, as shown by examples from Morgan Stanley, Zapier, and Spotify.

## Key ideas
### Deployment frequency is rising, but it is only a proxy for value
[01:24](https://www.youtube.com/watch?v=Se8jHLliLXE&t=84s)
DX sees deployment frequency steadily increasing, although the rate is beginning to taper. Reock cautions that this metric only shows how much work reaches production. It does not show revert rates, defect ratios, change failure rate, or the value created by the software. Regional patterns differ: North America is trending upward, while Europe pulled back in the latest quarter. Reock attributes some variation to differences in work practices, regulation, and spending on AI and tokens. He treats deployment frequency as one part of the software delivery lifecycle rather than a complete productivity measure.

### AI has made change failure rates much more volatile
[04:24](https://www.youtube.com/watch?v=Se8jHLliLXE&t=264s)
The effect of AI on quality varies sharply between companies. Reock describes the change failure rate graph as unusually volatile, with some companies increasing their rate by as much as 2 percentage points. Since the industry benchmark is about 4%, that increase could mean shipping 50% more defects than before. He says the pattern is not purely caused by AI. Release pipelines, automated testing, and the surrounding delivery process also matter. AI has changed the amplitude of the volatility, pushing some companies toward more extreme outcomes.

### Maintainability has improved while confidence in changes has fallen
[06:02](https://www.youtube.com/watch?v=Se8jHLliLXE&t=362s)
Across the study of about 200,000 engineers, perceived code maintainability rose by almost 4%, while change confidence fell by 6%. Reock describes this as a tension created by assistants and agents. Engineers find it easier to understand and modify the code in front of them, but they trust their outputs less and feel more afraid of breaking production. Pull requests grew from roughly 44 lines on average to 72 over about a year. Larger changes create more review work and more opportunities for bugs and vulnerabilities.

### Larger pull requests are weakening incremental delivery
[07:02](https://www.youtube.com/watch?v=Se8jHLliLXE&t=422s)
Reock gives several reasons for the growth in PR size. Models often generate mediocre code, and fast function generation encourages developers to pack several changes into one pull request when builds take 45 minutes or an hour. Larger PRs are harder to review, roll back, and move between systems. The perceived ability to make small incremental changes fell by 10%, making it one of the lowest-scoring developer experience drivers in the study. Reock connects this decline to the difficulty of understanding and reversing larger changes.

### Junior engineers use AI more, while senior engineers use fewer tokens
[08:46](https://www.youtube.com/watch?v=Se8jHLliLXE&t=526s)
Junior engineers have the highest AI usage, which Reock attributes partly to having less to unlearn and entering the industry with these tools already familiar. For the same use case, junior engineers spend more tokens than senior engineers. Staff-plus engineers save about the same amount of time despite using AI less, because they can spot hallucinations and understand the surrounding architecture more easily. Smaller companies also report greater time savings, partly because they have simpler release pipelines and less organizational complexity.

### AI measurement needs utilization, impact, and cost
[10:37](https://www.youtube.com/watch?v=Se8jHLliLXE&t=637s)
Reock recommends keeping established developer experience and productivity measures rather than replacing them with tool telemetry. DX's framework has three dimensions: utilization, which covers who uses AI, how often, and for what use cases; impact, which asks which business and delivery measures should change; and cost, including token spending. Teams can compare user cohorts across measures such as PR cycle time, PR size, and review pushback. Utilization alone cannot show whether an AI investment is creating value.

### Good platform foundations also improve agent experience
[13:27](https://www.youtube.com/watch?v=Se8jHLliLXE&t=807s)
Reock says many organizations gave developers coding assistants in 2024 and began building agents in 2025 before preparing their infrastructure. AI readiness depends on clear documentation, well-structured data, modular code, reliable local CI, and non-flaky test suites. These are familiar developer experience practices, and Reock argues that they also help agents work with less wasted context and token use. DX is additionally collecting qualitative feedback from agents about steering, the context they receive, and the feedback cycles required when working with humans.

### Code generation was never the main bottleneck
[14:57](https://www.youtube.com/watch?v=Se8jHLliLXE&t=897s)
Even perfectly accurate, instant code generation would address only about 14% to 16% of the overall software value stream. DX found a median 7.7% increase in the measured velocity metric, with a 13% average, while even the top performers did not reach twice the throughput. Meeting-heavy days, context switching, interruptions, and other development friction can outweigh time saved by AI. Reock uses Eli Goldratt's bottleneck argument: time saved outside the bottleneck does not change the overall result.

### The strongest applications span the wider software lifecycle
[16:27](https://www.youtube.com/watch?v=Se8jHLliLXE&t=987s)
Morgan Stanley's DevGen agent interprets legacy Natural and COBOL code and creates PRDs for engineers, saving about 300,000 hours per year. Zapier uses agents for administrative work, reduced standups to twice a week, automates about 3,000 initial code reviews per week, and reports 15% more value per engineer. Spotify built an SRE incident agent that gathers remediation steps from runbooks and incident context and posts them into communication channels. These examples apply AI to discovery, coordination, review, onboarding, and incident response rather than only code generation.

## Notable quotes
- "They're not fully representative of the value generation." (01:40)
- "Our change confidence number has gone down 6%." (06:30)
- "Code generation was never the bottleneck in the first place." (15:02)
- "An hour saved on something that isn't the bottleneck is worthless." (16:16)

## Tools & references mentioned
- DX
- DORA metrics
- SPACE framework
- DevEx framework
- MER study
- Morgan Stanley
- DevGen
- Zapier
- Spotify
- Eli Goldratt
- The Goal
- The Phoenix Project
- Natural
- COBOL

## Who should watch
- You are responsible for measuring AI adoption and need a way to connect tool usage with delivery, quality, and business outcomes.
- Your team is seeing larger PRs or more uncertainty around AI-generated changes and wants data on the trade-offs.
- You are building coding agents and need to assess documentation, testing, CI, and other platform conditions before spending more on tokens.

## Related talks

- [What Data from 20m Pull Requests Reveal About AI Transformation](https://aietalks.com/talks/what-data-from-20m-pull-requests-reveal-about-ai-transformation) (Nick Arcolano, Jellyfish, 17:57)
- [Can You Prove AI ROI in Software Engineering?](https://aietalks.com/talks/can-you-prove-ai-roi-in-software-engineering) (Yegor Denisov-Blanch, Stanford, 16:40)
- [How AI Is Changing Software Engineering](https://aietalks.com/talks/how-ai-is-changing-software-engineering) (Gergely Orosz, The Pragmatic Engineer, 26:42)
- [Leadership in AI Assisted Engineering](https://aietalks.com/talks/leadership-in-ai-assisted-engineering) (Justin Reock, DX, 18:11)
- [The State of AI Code Quality: Hype vs Reality](https://aietalks.com/talks/the-state-of-ai-code-quality-hype-vs-reality) (Itamar Friedman, Qodo, 21:15)
