# How Software Factories Improve Themselves

Suraj Gupta, Warp | AI Engineer | 12:52

Source: https://www.youtube.com/watch?v=TN3mj92oZ8I
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/how-software-factories-improve-themselves
Published: 2026-09-27
Tags: coding-agents, continual-learning, human-in-the-loop, memory, software-factories

## TL;DR
- An outer-loop agent can inspect a triage agent's work and propose updates to its skill through a human-reviewed pull request.
- Persistent memory lets agents reuse versioned facts and prior root-cause analysis instead of repeating the same investigation.
- Model routing can assign cheaper or better-performing models to specific task types, based on rules and workflow-specific evaluations. 

## Summary
Suraj Gupta describes three ways a software factory can improve after it has been deployed. Skills give agents procedural memory, but those instructions become dated as agents run and people provide feedback. Warp uses an outer-loop agent to inspect the triage agent's decisions, collect feedback, update the skill, and open a pull request so the change is tracked in Git and reviewed by a person. Persistent memory stores facts, outcomes, and learnings from prior runs. Gupta demonstrates this with a Sentry agent that can reuse root-cause analysis and inspect where memories came from. The memory store is versioned and works across harnesses. The final approach is model routing. Warp offers automatic routing and lets users define rules for task classes. Gupta says Warp's internal evaluations found that UI tasks worked well with GLM, and describes plans to let customers run their own workflow-specific evaluations.

## Key ideas
### Software factories need their own improvement loops
[00:48](https://www.youtube.com/watch?v=TN3mj92oZ8I&t=48s)
Gupta separates two ideas that are often discussed independently. Self-improvement lets agents or models improve over time, while a software factory automates work from triage through production. His focus is what happens to the factory itself as engineers maintain it. He says factories should become better, more efficient, and faster over time, especially as software engineers shift from building individual products toward building and maintaining these automated systems.

### An outer-loop agent can keep procedural skills current
[02:04](https://www.youtube.com/watch?v=TN3mj92oZ8I&t=124s)
Skills give an agent procedural memory by describing how to perform a repeated task. Gupta uses a triage agent as the example, with a skill that explains how to reproduce issues. The skill can become stale after humans give feedback or the agent learns from its own runs. An outer-loop agent watches the inner-loop agent, looks for mistakes and feedback, and improves the skill instead of leaving the original instructions unchanged.

### Skill changes should pass through Git and human review
[03:15](https://www.youtube.com/watch?v=TN3mj92oZ8I&t=195s)
Warp's open-source client repository includes a triage workflow that runs when an issue is submitted. The outer-loop agent reviews the triage decisions and collects signals such as thumbs-up or thumbs-down reactions, user comments, and comments from Warp employees. It synthesizes that feedback, updates the triage skill, and opens a pull request. Git provides a history of how the skill changed, while human review can catch an update that would make the triage agent worse.

### Persistent memory prevents repeated investigations
[05:29](https://www.youtube.com/watch?v=TN3mj92oZ8I&t=329s)
Skills remember procedures, while persistent memory stores facts, learnings, and outcomes from earlier runs. Gupta describes a Sentry agent that investigates an issue, gathers context, finds a root cause, and fixes it. Without memory, a later run may need to rediscover the same cause and spend more tokens gathering context. An outer-loop agent extracts useful information from completed runs so future investigations can use what the system already learned.

### Memory stores need versioning and traceable sources
[06:35](https://www.youtube.com/watch?v=TN3mj92oZ8I&t=395s)
In Warp's cloud agent platform, agents can have attached memory stores made up of facts. The Sentry example shows memories about previous root-cause analysis. People can create, edit, delete, and version memories, while agents primarily create them. Each memory can be traced back to its source run, which lets a person decide whether it captures a useful fact or an overly local conclusion. Later runs report which memories they used.

### Persistent memory can travel across harnesses
[08:25](https://www.youtube.com/watch?v=TN3mj92oZ8I&t=505s)
Gupta says persistent memory in Warp works across harnesses rather than being tied to one agent runtime. He names Warp's proprietary harness, Claude Code, and Codex as examples. Memories are created automatically, while users retain control over reviewing, editing, deleting, and versioning them. The resulting memory history is traceable.

### Routing avoids using expensive models for every task
[08:45](https://www.youtube.com/watch?v=TN3mj92oZ8I&t=525s)
Running every factory agent on Opus can become expensive when the work is simple, such as triage or fixing a basic CI failure. Warp offers out-of-the-box auto models that are evaluated as new models become available. Users can also define routing rules themselves. Gupta's example assigns database migrations to GLM and runbooks and API documentation to Qwen, allowing task classes to use different models.

### Workflow-specific evaluations can guide routing choices
[10:50](https://www.youtube.com/watch?v=TN3mj92oZ8I&t=650s)
Gupta calls custom model selection more art than science when it is based only on general impressions. Warp is working on customer-facing evaluations where teams can define the qualities they care about and test models against their own workflows. Its internal routing setup uses an evaluation sidecar and a best-at-k approach, running agents across several models for a prompt. Gupta says Warp found that UI tasks worked well with GLM and did not need Opus.

## Notable quotes
- "The idea is that your inner loop agent is the thing that's actually applying the skill." (02:42)
- "All of the improvements to the inner loop skill are going to be tracked through Git." (04:56)
- "You can think of this as a fact store scoped to an agent." (06:11)
- "It can become prohibitively expensive to run all of your agents that are doing simple things like triage or fixing simple CI failures with Opus." (08:44)
- "UI tasks are really well done with GLM. We don't really need to run those with Opus." (11:59)

## Tools & references mentioned
- Warp
- Oz
- Sentry
- Claude Code
- Codex
- Opus
- Haiku
- GLM
- Qwen
- Git

## Who should watch
- You are building a factory with repeated triage, issue investigation, or CI work and need a way to improve its instructions without silently changing production behavior.
- Your agents repeatedly rediscover the same facts or root causes, and you want those findings stored with sources and version history.
- You are paying for a powerful model on simple tasks and want routing rules or evaluations based on your own workflows.

## Related talks

- [Building your own software factory](https://aietalks.com/talks/building-your-own-software-factory) (Eric Zakariasson, Cursor, 1:23:37)
- [Software Engineering Is Becoming Factory Engineering](https://aietalks.com/talks/software-engineering-is-becoming-factory-engineering) (Zach Lloyd, Warp, 20:37)
- [Agentic Engineering: Working With AI, Not Just Using It](https://aietalks.com/talks/agentic-engineering-working-with-ai-not-just-using-it) (Brendan O'Leary, Kilo Code, 27:03)
- [Agents Building Agents](https://aietalks.com/talks/agents-building-agents) (Alfonso Graziano, Nearform, 30:14)
- [From Coding to Knowledge Work Agents](https://aietalks.com/talks/from-coding-to-knowledge-work-agents) (Karan Vaidya, Composio, 20:42)
