How VS Code Went from Monthly to Weekly Releases with AI

Harald Kirschner, Microsoft19:35 · Oct 2026 · 5,813 views
Thumbnail for How VS Code Went from Monthly to Weekly Releases with AI Watch on YouTube
TL;DR
  1. 1

    VS Code moved from monthly to weekly releases by changing its engineering system around AI, rather than simply asking developers to use more AI.

  2. 2

    The team improved speed by making the codebase agent-ready, encoding expertise as skills, speeding up TypeScript builds, and letting agents run and inspect the application.

  3. 3

    Quality depends on feedback loops, including mandatory AI code review, automated issue triage, error-stack fixes, staged rollouts, and product evaluations through VSC-Bench.

Summary

Harald Kirschner explains how VS Code moved from more than ten years of monthly releases to weekly releases. The change came from rebuilding the software delivery system around AI agents. The team made the repository easier for agents to understand, turned accessibility guidance into a reusable skill, moved to TypeScript Go for faster builds, and gave agents ways to launch VS Code, click through scenarios, and inspect screenshots. Faster code creation created new pressure on review, issue triage, telemetry, and releases. VS Code responded with mandatory AI review, automated issue filtering, error-stack grouping, auto-filed issues, fix pull requests, and staged rollouts. Kirschner also describes VSC-Bench evaluations, rapid prototypes, daily feedback, and smaller squads. The talk is candid about the costs: slow CI becomes much worse when many agents depend on it, and simple evaluations can expose large differences in token use between models.

Key ideas
02:42

Weekly releases required a new delivery system, not just more AI usage

VS Code had shipped monthly releases for more than ten years, dating back to version 1.0. Kirschner says the move to weekly releases did not come from spending more time with AI or trying to maximize token use. The team changed the whole system around software delivery. He breaks the work into shipping faster, holding quality, and learning faster. Faster code creation alone increases the burden on review and release safety. Once the team can ship high-quality code quickly, it still needs to decide whether it is building the right thing. AI supports that learning loop through evaluations, prototypes, and faster product feedback.

04:52

Agent-ready repositories encode the knowledge agents need

Kirschner recommends lightweight repository guidance such as agents.md files that map the codebase and point agents toward the right places. These documents should be living material that changes as agents make mistakes and learn how the repository works. Existing developer documentation and onboarding material can help agents as well as people. VS Code also encodes specialist knowledge as skills. Its accessibility skill carries accessibility practices into everyday work, so the area owner can maintain guidance that used to require one specialist to review each contribution.

07:01

Fast feedback becomes more important when many agents share the same build system

Slow builds and CI are inconvenient for developers because they can switch tasks while waiting. Agents cannot hide the same delay as easily. Kirschner says VS Code moved to TypeScript Go and saw a ten-times improvement in builds. With ten or twenty agents running and depending on the same CI feedback loop, a slow part of the system compounds across all of them. Faster builds helped agents receive feedback, correct code, and continue work without turning the shared pipeline into the main bottleneck.

08:31

Agents need to use the application to correct their own UI work

Kirschner says an agent may claim that a UI looks right even when the opened application is misaligned. VS Code therefore invested in an automated component browser that captures screenshots whenever components change and shows visual differences. The team also uses Playwright through a launch skill. An agent can open VS Code, run a defined scenario, click through the product, collect logs, and check a fix before and after making it. This gives the agent an application-level feedback loop instead of relying only on source code and tests.

11:20

Automated review and issue triage protect quality at higher output

VS Code made AI review mandatory for pull requests. Humans wait until the automated review comments have been addressed, and the review effort can vary with the risk of the repository. Issue handling also moved from manual triage toward a mix of AI and human correction. AI filters spam, enriches and translates reports, and assigns area owners. A Chrome extension lets people correct duplicate issue decisions, feeding human feedback back into the process when the agent gets something wrong.

13:25

Telemetry can turn recurring errors into proposed fixes

VS Code collects raw telemetry and filters it down to error stacks with complete information. The remaining errors are grouped and fingerprinted, then used to file issues for specific area owners. Kirschner says the system processes 51 billion telemetry events per day before this filtering. It can open a pull request with an initial diagnosis and an attempted fix. In his example, an agent traced a cancellation problem to a missing part of an RPC protocol and opened a pull request that the team could merge after human oversight.

15:24

Staged rollouts reduce the cost of failures on installed software

VS Code previously released a tested build to all users at once. Kirschner calls this a form of 'yolo releases'. With more AI-generated changes, the team moved to staged rollouts. During the rollout, it watches error logs, issues, and other signals before expanding distribution. This matters because VS Code installs binaries on user machines, where recovery and rollback are more expensive than in a web application. The rollout system applies the monitoring and gradual exposure used by web services to a desktop product.

15:59

Evaluations and prototypes make product learning faster

VSC-Bench gives the team product-specific evaluations for its agentic features. New scenarios can begin from a GitHub issue and a template, so customer conversations and reported problems can become tests. Kirschner describes a five-character hello-world file evaluation where the most expensive model used 70 times more tokens than another model. The team also prototypes ideas in daily conversations rather than waiting for a monthly release cycle. Smaller squads own smaller areas, which lets them revise prototypes and product decisions more quickly.

"If your agent cannot use your application, your product directly to get this feedback loop of is it all working, then that's a really big investment that pays off every time you work on UI."08:31
Who should watch
  • You are increasing AI-generated pull requests and need to keep review, CI, and release safety from becoming the next bottleneck.
  • You maintain a desktop or installed application where a bad release is expensive to roll back and want practical staged-rollout and telemetry ideas.
  • You are building an agentic product and need product-specific evaluations, application-use feedback, and faster prototype discussions.