Software engineers are moving from writing code directly to governing agents that turn signals such as bug reports and user feedback into production code.
2
A software factory should be model-agnostic, support sovereign deployment, connect the full software development lifecycle, and make agents available across the tools teams already use.
3
Teams should roll out agent capabilities gradually and measure signal-to-production time, human interventions, repair time, code shelf life, and cost per pull request instead of token usage.
Summary
Davis Palmie describes a software factory as a system of agents that takes signals such as bug reports, user feedback, and incident alerts through planning, coding, review, testing, deployment, and monitoring. He argues that coding was never the main bottleneck. Review, debugging, documentation, testing, and the handoffs between teams take more time and create more failure points. The factory should choose models by task and cost, support deployment from managed environments to air-gapped systems, and share context across the full software development lifecycle. Palmie recommends starting with a narrow task such as incident triage, where engineers can check the agent's diagnosis before expanding its authority. He also describes eight areas of agent readiness, including validation, build systems, feedback loops, documentation, reproducible environments, modular code, observability, and security. Humans still set architecture, strategy, risk limits, and direction. The useful measures are delivery and quality outcomes, not generated tokens or lines of code.
AI engineering is moving from code completion to system governance
Palmie describes three eras: autocomplete and next-token prediction, generation of whole files, and agents that gather context, call tools, and debug their own output. The abstraction level has moved from tokens to files, then to context and entire systems. In the next stage, engineers govern agents, keep guardrails tight, catch code drift, and set priorities while the system turns inputs such as bug reports and user feedback into production code.
Palmie argues that coding was never the bottleneck. Teams spend substantial time reviewing pull requests, reproducing user bugs, debugging, updating documentation, and maintaining tests. Documentation becomes stale, dead code remains, and communication layers separate product, engineering, testing, and other teams. He says the time from code being written to reaching production can dwarf the time spent making the engineering change itself.
A software factory connects many engineering functions through agents
The software factory is a system of agents that takes an input signal to production deployment. Its agents can triage incidents, make plans, create and update documentation, execute code, review it, test it, roll it out, and monitor the release. Palmie also places customer support, product engineering, and deployment operations within this system. The aim is to connect these activities so an improvement in one group can improve the wider system.
Incident triage is a safer starting point than full automation
Palmie recommends beginning with a narrow, verifiable task. An agent could read a Sentry alert, gather traces and other context, and post a diagnosis in Slack for engineers to review. Repeated correct diagnoses can build trust and understanding before the agent receives more authority. Each part of the factory needs clear ownership and monitoring before the pipeline is allowed to run as a loop.
The factory should choose models and deployment options by task
Palmie gives three design principles. The system should be model-agnostic because models differ in capability and cost. He says Factory found GPT-5.2 could match the latest Opus models for code review at half the price, while open-source alternatives could reduce cost by 10 to 30 times. Deployment should be sovereign, with options from fully managed systems to bring-your-own-machine and air-gapped setups. The factory should also connect documentation, code, planning, testing, product, and engineering across the SDLC.
Agents need to work across the surfaces where teams already work
Palmie says agents should be surface-agnostic and available through remote machines, the command line, Slack, and other team tools. Humans and agents may use different services, but they should share an underlying harness and context. A unified system can expose tribal knowledge, manual processes, and outdated information that obstruct both agents and human engineers. For larger organizations, shared governance and context can reduce the management burden of fragmented teams.
Agent readiness depends on the surrounding engineering system
Factory assesses readiness across eight areas: validation, build systems, feedback loops, documentation, reproducible development environments, modular code, observability, and security scanning. Examples include linters and formatters, documented build and CI commands, unit and integration tests, README files and agents.md files, clear code boundaries, fast diagnosis of failures, and checks for leaked secrets or vulnerabilities. Palmie says paving these paths helps agents and human engineers who use the same systems.
Human judgment remains a control point as automation expands
Palmie says teams cannot govern what they cannot see. Agent actions should be auditable and follow organizational standards, role-based access control, and least privilege. Teams should automate individual pillars before turning the whole system into a loop. Humans remain responsible for architecture, strategy, prioritization, and direction, and they act as validation gates when code drifts or guardrails need tightening.
Outcome measures are better than token leaderboards
Palmie rejects lines of code and tokens generated as measures of engineering value because teams can optimize for those numbers without improving results. He calls token leaderboards an example of Goodhart's law. Factory's suggested measures are signal-to-production time, human intervention count, median time to repair, code shelf life, and cost per pull request. These measures connect agent activity to delivery, maintenance, quality, and spending.
"Engineers are going from typing code to governing agents and eventually you'll organize those agents into a system."01:52
Who should watch
You are considering agents that can act across coding, operations, testing, documentation, and release workflows, and need a model for how those pieces fit together.
Your organization has slow handoffs between product, engineering, testing, and operations, and you want a concrete starting point such as incident triage.
You need to set controls and outcome metrics for agent adoption instead of judging progress by generated tokens or lines of code.