# How VS Code Went from Monthly to Weekly Releases with AI

Harald Kirschner, Microsoft | AI Engineer | 19:35

Source: https://www.youtube.com/watch?v=I2LL_wd89-A
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/how-vs-code-went-from-monthly-to-weekly-releases-with-ai
Published: 2026-10-03
Tags: agents, coding-agents, evals, harness-engineering, reliability

## TL;DR
- VS Code moved from monthly to weekly releases by changing its engineering system around AI, rather than simply asking developers to use more AI.
- The team improved speed by making the codebase agent-ready, encoding expertise as skills, speeding up TypeScript builds, and letting agents run and inspect the application.
- Quality depends on feedback loops, including mandatory AI code review, automated issue triage, error-stack fixes, staged rollouts, and product evaluations through VSC-Bench.

## Summary
Harald Kirschner explains how VS Code moved from more than ten years of monthly releases to weekly releases. The change came from rebuilding the software delivery system around AI agents. The team made the repository easier for agents to understand, turned accessibility guidance into a reusable skill, moved to TypeScript Go for faster builds, and gave agents ways to launch VS Code, click through scenarios, and inspect screenshots. Faster code creation created new pressure on review, issue triage, telemetry, and releases. VS Code responded with mandatory AI review, automated issue filtering, error-stack grouping, auto-filed issues, fix pull requests, and staged rollouts. Kirschner also describes VSC-Bench evaluations, rapid prototypes, daily feedback, and smaller squads. The talk is candid about the costs: slow CI becomes much worse when many agents depend on it, and simple evaluations can expose large differences in token use between models.

## Key ideas
### Weekly releases required a new delivery system, not just more AI usage
[02:42](https://www.youtube.com/watch?v=I2LL_wd89-A&t=162s)
VS Code had shipped monthly releases for more than ten years, dating back to version 1.0. Kirschner says the move to weekly releases did not come from spending more time with AI or trying to maximize token use. The team changed the whole system around software delivery. He breaks the work into shipping faster, holding quality, and learning faster. Faster code creation alone increases the burden on review and release safety. Once the team can ship high-quality code quickly, it still needs to decide whether it is building the right thing. AI supports that learning loop through evaluations, prototypes, and faster product feedback.

### Agent-ready repositories encode the knowledge agents need
[04:52](https://www.youtube.com/watch?v=I2LL_wd89-A&t=292s)
Kirschner recommends lightweight repository guidance such as agents.md files that map the codebase and point agents toward the right places. These documents should be living material that changes as agents make mistakes and learn how the repository works. Existing developer documentation and onboarding material can help agents as well as people. VS Code also encodes specialist knowledge as skills. Its accessibility skill carries accessibility practices into everyday work, so the area owner can maintain guidance that used to require one specialist to review each contribution.

### Fast feedback becomes more important when many agents share the same build system
[07:01](https://www.youtube.com/watch?v=I2LL_wd89-A&t=421s)
Slow builds and CI are inconvenient for developers because they can switch tasks while waiting. Agents cannot hide the same delay as easily. Kirschner says VS Code moved to TypeScript Go and saw a ten-times improvement in builds. With ten or twenty agents running and depending on the same CI feedback loop, a slow part of the system compounds across all of them. Faster builds helped agents receive feedback, correct code, and continue work without turning the shared pipeline into the main bottleneck.

### Agents need to use the application to correct their own UI work
[08:31](https://www.youtube.com/watch?v=I2LL_wd89-A&t=511s)
Kirschner says an agent may claim that a UI looks right even when the opened application is misaligned. VS Code therefore invested in an automated component browser that captures screenshots whenever components change and shows visual differences. The team also uses Playwright through a launch skill. An agent can open VS Code, run a defined scenario, click through the product, collect logs, and check a fix before and after making it. This gives the agent an application-level feedback loop instead of relying only on source code and tests.

### Automated review and issue triage protect quality at higher output
[11:20](https://www.youtube.com/watch?v=I2LL_wd89-A&t=680s)
VS Code made AI review mandatory for pull requests. Humans wait until the automated review comments have been addressed, and the review effort can vary with the risk of the repository. Issue handling also moved from manual triage toward a mix of AI and human correction. AI filters spam, enriches and translates reports, and assigns area owners. A Chrome extension lets people correct duplicate issue decisions, feeding human feedback back into the process when the agent gets something wrong.

### Telemetry can turn recurring errors into proposed fixes
[13:25](https://www.youtube.com/watch?v=I2LL_wd89-A&t=805s)
VS Code collects raw telemetry and filters it down to error stacks with complete information. The remaining errors are grouped and fingerprinted, then used to file issues for specific area owners. Kirschner says the system processes 51 billion telemetry events per day before this filtering. It can open a pull request with an initial diagnosis and an attempted fix. In his example, an agent traced a cancellation problem to a missing part of an RPC protocol and opened a pull request that the team could merge after human oversight.

### Staged rollouts reduce the cost of failures on installed software
[15:24](https://www.youtube.com/watch?v=I2LL_wd89-A&t=924s)
VS Code previously released a tested build to all users at once. Kirschner calls this a form of 'yolo releases'. With more AI-generated changes, the team moved to staged rollouts. During the rollout, it watches error logs, issues, and other signals before expanding distribution. This matters because VS Code installs binaries on user machines, where recovery and rollback are more expensive than in a web application. The rollout system applies the monitoring and gradual exposure used by web services to a desktop product.

### Evaluations and prototypes make product learning faster
[15:59](https://www.youtube.com/watch?v=I2LL_wd89-A&t=959s)
VSC-Bench gives the team product-specific evaluations for its agentic features. New scenarios can begin from a GitHub issue and a template, so customer conversations and reported problems can become tests. Kirschner describes a five-character hello-world file evaluation where the most expensive model used 70 times more tokens than another model. The team also prototypes ideas in daily conversations rather than waiting for a monthly release cycle. Smaller squads own smaller areas, which lets them revise prototypes and product decisions more quickly.

## Notable quotes
- "It's about evolving the whole system, how you're shipping software to evolve to actually make better use of AI along the whole process." (03:07)
- "If your agent cannot use your application, your product directly to get this feedback loop of is it all working, then that's a really big investment that pays off every time you work on UI." (08:31)
- "The same five character file took the most expensive model 70x more tokens." (17:04)
- "We should have probably done it earlier. So now actually we do stage rollouts and as we do the roll out we monitor error locks and issues and everything else." (15:24)

## Tools & references mentioned
- VS Code
- Microsoft
- GitHub
- GitHub Copilot
- TypeScript Go
- Playwright
- Xcode MCP
- VSC-Bench
- Chrome extension
- agents.md

## Who should watch
- You are increasing AI-generated pull requests and need to keep review, CI, and release safety from becoming the next bottleneck.
- You maintain a desktop or installed application where a bad release is expensive to roll back and want practical staged-rollout and telemetry ideas.
- You are building an agentic product and need product-specific evaluations, application-use feedback, and faster prototype discussions.

## Editor's note

Harald Kirschner describes VS Code turning recurring telemetry errors into issues and attempted pull requests, including an agent tracing a cancellation problem to a missing part of an RPC protocol. Kitaru records every model call and tool result from a real agent run. Replay reruns the agent against the same inputs and tool responses, so the team can change the model, prompt, or code and inspect the failed run without touching real systems.

Written by the AIE Talks editors (the Kitaru team), not by the speaker.

## Related talks

- [Developer Experience in the Age of AI Coding Agents](https://aietalks.com/talks/developer-experience-in-the-age-of-ai-coding-agents) (Max Kanat-Alexander, Capital One, 18:20)
- [Software Engineering Is Becoming Plan and Review](https://aietalks.com/talks/software-engineering-is-becoming-plan-and-review) (Louis Knight-Webb, Vibe Kanban, 20:23)
- [Software Development Agents: What Works and What Doesn't](https://aietalks.com/talks/software-development-agents-what-works-and-what-doesnt) (Robert Brennan, OpenHands, 16:46)
- [Unlocking AI-Powered DevOps Within Your Organization](https://aietalks.com/talks/unlocking-ai-powered-devops-within-your-organization) (Jon Peck, GitHub, 22:13)
- [Don't get one-shotted: Use AI to test, review, merge, and deploy code](https://aietalks.com/talks/dont-get-one-shotted-use-ai-to-test-review-merge-and-deploy-code) (Tomas Reimers, Graphite, 05:45)
