An agent is a language model surrounded by a harness that provides memory, tools, perception, and control over context.
2
Harness engineering makes nondeterministic model behavior more repeatable by controlling the data, tools, workflows, and limits around the model.
3
Agent memory should separate short-lived context from durable knowledge, then turn successful workflows into reusable skills.
Summary
Ignacio Martinez argues that the language model is the least controllable part of an agent. Its weights are usually fixed, and its outputs can vary for the same input. The controllable part is the harness around it, which includes storage, memory engineering, semantic knowledge, retrieval, context management, tools, skills, and an agent loop. He compares files and databases, recommending a hybrid approach where files hold short-term state and databases hold durable information. He explains context rot, where adding more material to a context window reduces attention to each item. He also introduces the semantic layer as the institutional knowledge an agent needs in order to interpret work correctly. The workshop builds these components step by step, starting with a bare model call. Later sections cover tool retrieval, skill promotion, model routing, and a hysteresis variable that limits how long the harness keeps trying a failing task.
Martinez separates the language model from everything built around it. The model provides reasoning, but its weights usually do not change, and its output can differ even when the input stays the same. The harness adds memory, tools, perception, and control logic. Its purpose is to turn that variable reasoning into outputs that are more reliable and repeatable. He describes an agent as a large language model plus a harness, with the model acting as the reasoning core and the harness handling the surrounding systems.
The agent stack concentrates engineering control in the data layer
Martinez describes five layers in an AI application: the application, data, model, infrastructure, and compute. He says the application, model, infrastructure, and compute are becoming increasingly commoditized. Data remains the layer where teams have the most control. The harness sits close to that layer and includes a gateway, memory, semantics, retrieval, context, tools, and skills. MCP connects the model to external programs and data. He uses Outlook as an example of a program that a model could access through an MCP interface.
A hybrid storage system combines file-friendly workflows with database guarantees
Files are easy for models to create, append to, and use through normal operating-system semantics. They do not provide transactional consistency, backups, or convenient hybrid search. When several agents edit the same file, work trees isolate their changes until they can be merged. Martinez argues that databases already solved many of these problems through atomicity, consistency, isolation, durability, replication, and search. His proposed hybrid keeps short-term memory in files and promotes durable information, such as user preferences, into a database. He also describes Oracle DBFS as a way to store files inside a database.
Agent memory separates temporary context from information worth keeping
Martinez defines agent memory as the systems that let an agent retain, reuse, refine, and recall information. Short-term memory holds ephemeral work, such as a coding agent's current to-do list. Episodic memory can preserve previous conversations, while procedural memory can preserve workflows that worked well. Shared memory lets agents communicate with a parent agent or collaborate on a task. He presents Oracle Agent Memory Package as a managed way to create context cards containing topics, a summary, facts, preferences, memories, unanswered questions, and recent messages.
Martinez rejects the idea that putting everything into a very large context window solves memory. The context window is itself a form of short-term memory, and adding more material reduces the attention available to each item. He compares this to a conversation that continues for eight hours, when attention becomes harder to sustain. In the model, the attention matrix grows across both rows and columns as the number of tokens increases. His practical advice is to keep the context window as small as possible and retrieve only information that matters for the current step.
An agent's semantic layer contains the knowledge people usually leave unstated
Martinez uses the concept of Umwelt from biologist Jakob von Uexküll to explain how an agent perceives work through the information available to it. An organization's semantic layer contains the vocabulary and assumptions that colleagues normally share without explaining. This includes institutional knowledge, data models, query patterns, metadata, and other forms of tribal knowledge. The agent's perception is filtered through this layer, so an agent needs more than documents or raw retrieval. It needs the context that explains how the organization interprets its data and processes.
Context engineering should load tools and skills only when needed
The agent loop observes, reasons, and acts repeatedly. At each iteration, the harness can assemble a fresh context and decide which tools and skills belong in it. Martinez presents the toolbox pattern and skillbox pattern as ways to store available tools and skills and retrieve them when the task requires them. This avoids placing every possible capability into the context window. He also describes tool retrieval with HNSW indexes, where a graph structure connects vector indexes and lets the system search a large tool collection in a database.
Skill promotion turns successful work into reusable behavior
Martinez describes continual learning that does not require changing model weights. A workflow that took several hours can be stored, distilled, and promoted into a better skill. The old skill is retired and replaced by a new skill.md that captures the improved procedure. This can encode preferences such as a favored library, database engine, tone, or working style. The result is a workflow that can be retrieved and reused when a similar task appears later.
A hysteresis variable limits how long the harness trusts a failing model
The harness needs a limit on repeated attempts because model calls cost money and may continue producing hallucinations. Martinez says the right cutoff depends on the model and the task. For one frontier model, he found that allowing roughly 8 to 12 tool calls was a reasonable maximum before giving up, although another example succeeded after 16 steps. He calls the control a hysteresis variable, which stores how much patience the harness gives the model. The harness can also route harder tasks to a frontier model and easier tasks to a smaller specialist model.
"The more things that you put in the context window, the less attention there will be for each one of the things that are in the context."29:09
Who should watch
You are building an agent and need a practical breakdown of the systems around the model, including storage, retrieval, context, tools, and loops.
Your team is deciding whether agent state belongs in files, databases, or both, and you want a concrete argument for separating short-term and durable memory.
You are trying to make agents improve over time without fine-tuning model weights, through workflow storage, skill promotion, model routing, and retry limits.