A software factory runs the whole software lifecycle autonomously, from collecting signals and prioritizing work through building, validation, production testing, and iteration.
2
The hard parts are model routing, long-running execution, reliable validation, context management, and preparing the codebase for AI, rather than generating code.
3
Humans should decide what software to build while agents handle more of the implementation and administrative work around it.
Summary
Tereza Tížková defines a software factory as an autonomous loop covering the whole software lifecycle. It collects signals such as user feedback and logs, prioritizes work, orchestrates agents, builds and validates software, tests it in production, and improves from the results. She argues that writing code is relatively easy compared with deciding when work is done, validating behavior, managing long-running tasks, and keeping context under control. Her approach has three parts: remain agnostic to models and existing team workflows, run agents autonomously for long periods, and keep improving the system and codebase. Factory Missions use an orchestrator, sequential workers, and validators, including one that clicks through the application. Tereza also describes model routing, deferred context, agent-readiness checks, and plugins for capturing team knowledge. She expects humans to spend more time deciding what to build while agents take on implementation and repetitive coordination work.
A software factory runs the whole software lifecycle, not just code generation
Tereza defines a software factory as the complete loop of software development with autonomy. It collects signals such as user feedback and logs, decides what matters, orchestrates the work, executes it, validates the result, tests in production, and iterates. The system also gains knowledge and skills over time. She says this became practical for enterprise customers such as EY and Adobe only after models, context handling, reasoning quality, and agent environments improved enough.
The difficult work starts after an agent writes code
Tereza says a software factory is not a single coding agent or a swarm of coding agents. Even thousands of agents generating code would leave the harder work untouched. Engineers spend much of their time on the surrounding tasks, including deciding priorities, checking behavior, and handling the rest of the software lifecycle. She also rejects treating the factory as a consultancy project that can be dropped into an unchanged organization. In her view, the organization needs to be rebuilt around the new way of working.
Model routing should fit existing teams and choose the cheapest model that can do the task
The factory should work across tools such as Slack and GitHub and let teams keep the subscriptions they already use. Tereza describes Factory's automatic model routing as a system that classifies a task using its prompt, codebase, difficulty, and tools. It then selects the cheapest model predicted to clear the required quality threshold. The router can switch models when a task fails, and model choice can also affect speed and provider reliability. Tereza gives a conservative benchmark of about 25 percent savings.
Long-running agents need explicit, testable definitions of done
Loops are easy to describe, but open-ended tasks make completion difficult to define. Tereza gives an example involving agents building a 3D-printed Factory logo, where the system must decide whether the real-world result is complete. Agents can also optimize for passing the stated tests instead of accomplishing the intended task if the completion conditions are weak. Factory Missions address this with long-running sessions that can continue for weeks, using an orchestrator to assign work and validators to check the result.
Sequential workers give each agent fresh context and separate validation from authorship
In Factory Missions, workers operate in sequence rather than as one large swarm. Each worker completes part of the task and passes it to the next worker, while smaller subagents can work in parallel on research or files. Tereza says the sequence gives agents fresher context, similar to having another colleague review work. Validators judge code they did not write. A validation contract is written before implementation and includes code scrutiny plus user testing.
A validator that uses the application can catch results that only look correct
The user-testing validator runs the application in a virtual computer and clicks through it. Tereza contrasts this with systems that create a product which looks complete in the code but is not interactive. In one migration example, the agent needed to click through the application to confirm that it actually worked. She connects this capability to better computer-use models and persistent virtual-machine environments for agents.
Deferred context keeps large tool collections from overwhelming agents
Enterprise teams may connect hundreds of tools, including Figma, Notion, Gmail, Drive, and Slack. Loading every tool schema, parameter, and description into every task can fill the context window and cause agents to choose similar tools incorrectly. Factory's deferred context engine initially exposes only a short list and brief descriptions. A tool's full details become available when the agent needs it. Tereza says this can save 50 percent of tokens or more at scale because information is hidden rather than repeatedly loaded.
Tereza describes AI adoption as having a power-law pattern: teams can gain a lot, or the codebase can degrade. She points to code structure, documentation, reproducible developer environments, tests, linters, and consistent style as factors in the outcome. Factory's agent-readiness check examines these conditions and recommends fixes before a team turns more work over to agents. She also describes plugins as reusable packages of skills and context that capture team rules that would otherwise be learned only by observation.
Humans should decide what to build while agents handle how to build it
Tereza expects people to move up the abstraction ladder as agents take over more implementation work. She compares this with earlier shifts from human computers to programming languages and then coding agents. In her model, humans monitor structured groups of orchestrator, worker, and validator agents and decide what software to build. She also expects factories to take on organizational chores such as status syncs, meetings, and the work of collecting context from different people.
"I would define the software factory as the whole loop, the whole life cycle of developing software with autonomy, which doesn't mean just coding and generating code."01:17
Who should watch
You are deciding whether an agent system should handle more than code generation and need a concrete description of the surrounding lifecycle.
Your team wants to reduce model costs or run agents for hours or weeks, but you are unsure how routing, validation, and completion criteria should work.
Your codebase has weak tests, documentation, or developer environments and you want to understand how that affects AI adoption.