AI performs well on code because code is self-documenting, modular, single-domain, and easy to benchmark, while production work crosses code, infrastructure, telemetry, and teams.
2
Production agents break through anchoring bias, poor context selection, coherent answers without causal evidence, unsafe actions, and a lack of learning from earlier investigations.
3
A production agent harness needs model orchestration, context engineering, causal reasoning, governed actions, learning systems, and evals.
Summary
Varun Krovvidi explains why AI that writes code well still struggles with production operations. Code is structured, modular, and usually confined to one domain. Production work crosses code, infrastructure, observability data, deployments, and several engineering teams, often with one correct answer required. He describes five failure modes that appear as agents scale: anchoring on an early theory, choosing the wrong model for a task, having too much or too little context, producing plausible explanations without causal evidence, taking unsafe actions, and repeating investigations without learning. Resolve AI's proposed harness has six parts: model orchestration, context engineering, causal reasoning, governed actions, learning systems, and evals. A demo shows the system investigating a Grafana alert from Slack, tracing failures to stale integrations, ruling out a simultaneous GCP outage, and resisting a suggested alternative explanation when the evidence does not support it.
Production work takes place across several systems and teams
Krovvidi says engineering time is split between generating software and running and fixing it in production. He groups this work into regular on-call alerts, major incidents that need broad attention, and recurring maintenance such as infrastructure, cost, and platform work. These tasks cross multiple domains and tools. Production is also a multiplayer problem because engineers from service, backend, platform, and other teams need shared context while resolving an issue.
A single large model does not cover every production task
A model can look effective in a focused demo, then fail as the product expands across use cases and teams. Krovvidi says a workflow contains many task types, including reasoning, deterministic operations, image reasoning, SQL, and log analysis. Different models work better for different tasks, so the system must keep evaluating new models and route each task to an appropriate one. Resolve AI describes this as a model treadmill that the harness has to manage.
Context selection is harder than simply enlarging the context window
Krovvidi describes two context failures. With too much information, an agent overexplores and invents hypotheses that do not help. With too little, it cannot see relevant paths or alternatives. Production telemetry can grow through logs, metrics, dashboards, code, and infrastructure data. The harness therefore needs to choose the precise context for each step, using combinations of retrieval methods and narrowly defined tool calls.
Production investigations need causal evidence rather than plausible explanations
Models are designed to produce coherent answers and may agree when a user pushes them toward a theory. Krovvidi says an incident investigation needs a causal chain, like a detective reconstructing the steps that led to the failure. Resolve AI bases a root cause on that chain. If it cannot establish the evidence beyond a point, it should lower its confidence and direct the user elsewhere instead of presenting a patch as the answer.
Actions need restrictions that match the team's operating rules
Krovvidi points to cases where an AI system could decide that deleting a file, database object, or code section is the cleanest fix. The model may be following its objective, so the system around it must restrict what it can do. Those restrictions include least-privilege access and rules for read and write operations. Each team or organization can define different conditions based on its access levels and operating environment.
A learning system carries investigation context forward
Production incidents involve several teams and repeated investigations. Krovvidi says the system should preserve the context of an earlier investigation rather than starting from the same point every time. It should also learn from user interaction. Positive and negative feedback, guidance, and attempts to steer an investigation can all become signals about how the agent performed.
Krovvidi presents evals as both the starting and ending point of the architecture. The system can record positive and negative reinforcement, trace how an answer was reached, score the investigation against how an engineer might perform it, and assess the agent's own confidence calibration. These evals need to be rerun when a model, use case, or part of the architecture changes.
The demo traces an alert to stale integrations and rejects tempting alternatives
Resolve AI takes a Grafana alert arriving in Slack and starts an investigation across code, infrastructure, knowledge bases, and observability systems. It identifies a log skill with a high failure rate and traces the issue to stale integrations. The interface also shows theories it rejected, including a GCP outage and traffic spikes that happened at the same time. When Krovvidi presses the agent on those alternatives, it keeps to the causal evidence it established.
"If you're not able to establish that chain of evidence, we have to give out that information with a low confidence level and point the user in a different direction."13:48
Who should watch
You are building an agent for incident response, on-call work, or production maintenance and need a design beyond prompting a general model.
Your agent produces plausible explanations but struggles with evidence, context size, model choice, or safe actions.
You want a concrete example of how an incident agent can investigate an alert and defend a conclusion when a user suggests another cause.