A harness is everything around a model, including memory, skills, tools, MCP servers, runtime, identity, and evaluation.
2
Agents built for large numbers of users need their components to scale separately instead of living in one container.
3
A simple JSON configuration with a model and system prompt can deploy an agent without custom agent code.
Summary
Mike Chambers separates agents people use from agents teams build. A coding assistant can rely on a harness containing memory, skills, tools, MCP servers, and shared engineering standards. A production agent needs more: loop management, scaling, payments, identity, runtime, context management, observability, and evaluation. Chambers argues that these parts should not all be placed in one container when the system must serve many users. His live coding moves from a small Strands agent with a calculator and time tool to an agent with session state and memory, then to an AWS Bedrock AgentCore deployment where memory and runtime infrastructure can scale independently. He also argues against agents creating cloud resources through ad hoc console actions. Infrastructure should remain code that the team owns. At the end, he shows a harness configuration that contains only a model and system prompt, with no agent code.
Agents people use and agents teams build have different constraints
Chambers divides agents into two groups. The first includes coding assistants such as Claude Code, Cursor, and Kiro, along with productivity tools that people use directly. The second includes agents that teams build for other users. He says token maximization and similar choices may be fine for an agent someone uses personally, but a built agent needs to be assembled for its audience and operating conditions. The distinction still applies when a team builds an agent that other people use. His talk concentrates on the second group, where deployment, scale, and operational behavior become part of the design.
A harness is everything left after removing the model
Chambers defines a harness by subtraction. Start with an agent, remove the model, and everything remaining is the harness. For a coding assistant, that includes memory, skills, tools, and MCP servers that connect to sources such as documentation. It also includes the standards an engineering team wants to distribute across its coding assistants. The definition gives the surrounding software and operating rules equal attention with the model itself. Chambers uses the image of straps and fastenings controlling an animal, then replaces the animal with a model to explain why the term fits.
Production agents need separate systems around the model
An agent built for broad deployment still needs memory, skills, tools, and MCP, but Chambers adds several operational concerns. The builder must manage the loop, scaling, payments, identity, runtime, and context. He puts observability and evaluations first in his own priority, even though he mentions them last. Chambers rejects putting all of these pieces into one container and scaling that container when the agent may serve thousands of users. Each component should be considered separately so it can scale on its own. That separation is what he calls harness engineering.
A bare SDK agent is useful but has little production support
The first live example uses the Strands Agents SDK. Chambers creates an agent with a system prompt and two tools, a calculator and a time tool. The framework manages the loop, so the code is small and already has a basic harness. It runs on his laptop, though, and does not operate at meaningful scale. Chambers says it lacks several attributes he wants in agents deployed to users. The example establishes the starting point for the rest of the demonstration: a working model-and-tools loop that still needs session state, memory, deployment, and other infrastructure.
Session management gives the local agent memory between invocations
The second example adds a session manager. It maintains session state between invocations and rehydrates the conversation history when Chambers returns to the agent. He describes this as short-term or medium-term memory rather than complete long-term memory, although the setup also stores long-term memories in files. In the demonstration, the agent remembers that Chambers wants Australia to win the World Cup because that preference appeared in an earlier conversation. The example remains local, so it shows what memory does without solving deployment or scale.
Cloud deployment moves memory and runtime outside the agent process
Chambers uses Amazon Bedrock AgentCore to deploy the agent and its supporting infrastructure. The AgentCore command-line tool can create a starter agent, with choices for Python or TypeScript, HTTP or other connection styles, and different agent frameworks. It can also provision short-term and long-term memory as cloud infrastructure separate from the running agent. This lets memory operate asynchronously and scale independently. The resulting agent includes connections for the runtime, session manager, memory, tools, and MCP, while the framework still permits custom code and models outside Amazon's own model family.
Infrastructure should stay owned code instead of agent-driven console actions
Chambers uses the term "slop ops" for asking an agent to create cloud resources directly, such as an S3 bucket or EC2 instance. He compares this with the earlier rejection of click ops, where people click around a cloud console to deploy systems. Console actions can help someone understand a service, but Chambers says production deployments should be built as infrastructure as code. That keeps the deployment under the team's ownership. The AWS Agent Toolkit is offered as a way to guide agents toward this approach when teams are deploying on AWS.
A configuration file can replace custom agent code for simple cases
At the end, Chambers returns to the harness option he skipped in the deployment wizard. He argues that many agent use cases may need little more than a system prompt and connections to MCP tools. In that case, AgentCore can provide the harness itself. The configuration he shows is a JSON file containing the model and the system prompt. There is no agent code in this version, and the configuration can still be deployed with the AgentCore command. He also says the components are composable, so an existing production agent can adopt a separately managed capability such as long-term memory without adopting everything else.
"We can just deploy it straight out. If you want to know any more about any of this then please do come and see us down on the booth or see me after this session."18:47
Who should watch
You are building an agent for many users and need to decide which runtime, memory, identity, and evaluation pieces should scale independently.
Your local agent already has tools and prompts, but you need a path to session memory and a managed cloud deployment.
You want agents to help write infrastructure while keeping production deployments as code that your team owns.