# Brains vs Hands: How to Run AI Agents Safely in Production

Viren Baraiya, Orkes | AI Engineer | 16:50

Source: https://www.youtube.com/watch?v=NaOkR3VSfR4
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/brains-vs-hands-how-to-run-ai-agents-safely-in-production
Published: 2026-10-08
Tags: harness-engineering, human-in-the-loop, reliability, workflows

## TL;DR
- Production agents include scheduled, event-driven, coordinating, and long-running systems, so their surrounding harness needs to behave like an application.
- The LLM should plan what happens next, while deterministic harness code executes the plan, handles approvals, and records side effects.
- Agent harnesses are late-bound sagas that compile runtime plans into durable workflows, as shown in the SRE remediation demo.

## Summary
Viren Baraiya argues that production agents need more than an LLM and a tool call. They may run on schedules, react to alerts, coordinate other agents, or wait for events and human decisions for months. The harness around them therefore becomes the application. It combines non-deterministic reasoning from an LLM with deterministic workflows for actions that must be repeatable, such as restarting a Kubernetes cluster. The LLM proposes what should happen next. The harness executes the approved plan, applies human gates, handles idempotency, and records side effects. Baraiya describes these systems as late-bound sagas because the agent assembles a workflow from a finite set of tools at runtime. His demo uses an SRE agent that gathers evidence, analyzes logs and metrics, performs a rollback when needed, verifies recovery, then replans from the new state. The workflow runs through Conductor and completes across two remediation loops.

## Key ideas
### Production agents include workers, listeners, coordinators, and multi-agent systems
[01:20](https://www.youtube.com/watch?v=NaOkR3VSfR4&t=80s)
Baraiya starts with the narrow hello-world pattern of an LLM calling a tool, then expands the scope. An agent can run every hour to inspect a schedule, listen for production alerts and logs, or monitor another long-running agent and nudge it forward. Agents can also talk to other agents. He says splitting responsibilities reduces the risk of one agent being responsible for too many things and hallucinating an unsuitable action.

### The harness becomes the application around the agent
[02:58](https://www.youtube.com/watch?v=NaOkR3VSfR4&t=178s)
Baraiya compares agent systems with microservice applications. A single microservice does not normally deliver an entire business goal, and multiple services coordinate through choreography or orchestration. His SRE harness can work with agents connected to logging systems, Kubernetes, metrics, and customer-service dashboards. The harness controls their execution and brings the surrounding systems, tools, databases, and humans into one application.

### Deterministic workflows must contain the actions that matter
[05:02](https://www.youtube.com/watch?v=NaOkR3VSfR4&t=302s)
A production harness combines the LLM's non-deterministic reasoning with deterministic execution. Baraiya uses payments and Kubernetes operations as examples of work that should not change unpredictably. If a cluster needs restarting, the harness should define the exact sequence of steps and run that sequence consistently. The LLM can reason about the situation, but it should not improvise the mechanics of a production restart.

### Durability is required for harnesses that wait and resume
[06:02](https://www.youtube.com/watch?v=NaOkR3VSfR4&t=362s)
A harness may run for seconds, days, months, or longer. It may wait for a third-party service, a human approval, or an event before taking its next action. Because the infrastructure around it can fail, the harness needs to recover from outages, network failures, and partitions. Baraiya calls durability a cost of admission rather than a special feature. The harness must preserve enough state to continue after a failure.

### The LLM plans while deterministic code executes
[08:03](https://www.youtube.com/watch?v=NaOkR3VSfR4&t=483s)
The harness keeps a state of the world, including completed work and recorded side effects such as a cluster restart or an email. The LLM uses that state and the goal to decide what should happen next. The harness then executes the plan. For an unhealthy cluster, the harness determines the defined restart procedure. This separation also lets the system require a human approval, such as sending a Slack message before restarting production.

### Approval, idempotency, and side effects belong in the execution layer
[09:21](https://www.youtube.com/watch?v=NaOkR3VSfR4&t=561s)
Baraiya says a production action should not depend on an LLM hallucinating whether approval is needed. The harness should guarantee the approval step. Ideally, actions are idempotent, so repeating them has a controlled result. If an action is not idempotent, the system should record that fact and track its side effects separately. This is his brain-and-hands split: the LLM is the brain, while the harness is the hands.

### Agent harnesses are late-bound sagas
[10:21](https://www.youtube.com/watch?v=NaOkR3VSfR4&t=621s)
Traditional sagas define a deterministic sequence of workflow steps in advance. Baraiya describes agentic harnesses as late-bound sagas because they have a finite set of tools and tasks, while the agent proposes how to combine them at runtime. The workflow still provides visibility and control. The agent can plan one step or several steps ahead, then revise its plan from the resulting state. This avoids manually defining every possible combination of tool calls.

### The SRE demo compiles a plan into a durable workflow
[11:51](https://www.youtube.com/watch?v=NaOkR3VSfR4&t=711s)
The demo uses an SRE agent in a remediation loop. The agent first gathers evidence, analyzes logs, and decides whether a rollback is needed before verifying recovery. Its proposed sequence goes to a plan-and-compile tool, which turns it into a runnable Conductor workflow. After the first execution, the agent examines the new state, checks downstream systems, and plans the next work. In the second iteration it does not roll back again because the first loop already did so.

## Notable quotes
- "Your harness is the application and vice versa." (04:22)
- "The responsibility of a non-deterministic agent is to plan what should happen next, not really to do things." (08:03)
- "Harness is the hands, the brain is the LLM." (09:56)
- "When you think about agentic harnesses, they are essentially late bound sagas." (10:21)

## Tools & references mentioned
- Orkes
- Conductor
- Kubernetes
- LangChain
- OpenAI

## Who should watch
- You are building an agent that must react to production events, wait for people, or continue after infrastructure failures.
- Your LLM can propose operational actions, but you need repeatable execution, approval gates, and a record of what already happened.
- You want a concrete pattern for compiling an agent's runtime plan into a durable workflow rather than letting it call production tools directly.

## Related talks

- [Build Agents That Run for Hours](https://aietalks.com/talks/build-agents-that-run-for-hours) (Ash Prabaker & Andrew Wilson, Anthropic, 1:15:40)
- [The 6 Pillars of an Agentic Harness for Production](https://aietalks.com/talks/the-6-pillars-of-an-agentic-harness-for-production) (Varun Krovvidi, Resolve AI, 21:05)
- [Harness Engineering: Building the Production Cage for Powerful Domain Agents](https://aietalks.com/talks/harness-engineering-building-the-production-cage-for-powerful-domain-agents) (Mike Chambers, AWS, 20:46)
- [Harnesses in AI: A Deep Dive](https://aietalks.com/talks/harnesses-in-ai-a-deep-dive) (Tejas Kumar, IBM, 20:27)
- [Your Agent Didn't Fail. Your Harness Did.](https://aietalks.com/talks/your-agent-didnt-fail-your-harness-did) (Vinoth Govindarajan, OpenAI, 18:26)
