# Giving AI Agents Memory That Learns

Jake Broekhuizen, LangChain | AI Engineer | 16:39

Source: https://www.youtube.com/watch?v=KGFyOtl5ktI
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/giving-ai-agents-memory-that-learns
Published: 2026-10-08
Tags: continual-learning, human-in-the-loop, memory, observability

## TL;DR
- An agent trace becomes memory only when a lesson is converted into durable context that the agent can read on a future run.
- Procedural memory, including instructions, skills, and rules, is the memory type that most directly changes visible agent behavior.
- Most trace data should remain history, memory updates must account for caching, and changes to important procedural rules should receive human review.

## Summary
Jake Broekhuizen explains how agents can improve across runs by turning selected experience into durable context. He uses a financial services agent whose tone sometimes shifted from advisory to directive. Reviewing its traces and manually changing its instructions worked as a correction, but the process did not scale. Broekhuizen separates memory into semantic facts and preferences, episodic experiences and patterns, and procedural instructions, skills, and rules. He also distinguishes temporary working memory from long-term memory. His proposed read-write cycle loads relevant context, runs the agent, captures the resulting trace, filters evidence for useful signal, and writes that signal back into durable context. He describes LangSmith for capturing and analyzing trajectories and Context Hub for storing context the agent can use later. He is clear that most traces should stay history. Caching can prevent updates from reaching later runs, and procedural changes need human review because they strongly affect behavior.

## Key ideas
### Observability gives evidence, while memory changes future runs
[02:43](https://www.youtube.com/watch?v=KGFyOtl5ktI&t=163s)
Broekhuizen separates watching an agent from helping it improve. A trace records the tools called, the reasoning produced along the way, and the artifacts the agent referenced. That information is useful for understanding what happened, but it does not alter the context used on the next run. If a misguided instruction remains unchanged, the same mistake can persist. Memory begins when a lesson from the trace becomes durable context that the agent can reference later. In his terms, observable agents tell you what happened, while adaptive agents can change how they act on following runs.

### A financial agent needs durable tone rules
[00:54](https://www.youtube.com/watch?v=KGFyOtl5ktI&t=54s)
The example is a spending, budgeting, and non-investment advice agent for a financial services company. Its tone is supposed to be advisory and helpful. In some cases, it became pointed and directive, telling users that they needed to cancel subscriptions and move money into savings. The team treated this as a tone violation. A person had to inspect the traces, identify where the tone slipped, change the context the agent used, and check for regressions. Broekhuizen uses this example to show why an agent needs a system that can learn from interactions instead of relying on repeated manual corrections.

### Agent memory has semantic, episodic, and procedural forms
[04:22](https://www.youtube.com/watch?v=KGFyOtl5ktI&t=262s)
Broekhuizen groups memory into three types. Semantic memory is what the agent knows, including facts and preferences that guide responses. Episodic memory is what the agent has experienced, including learned patterns, past interactions, and examples. Procedural memory is how the agent should behave, including instructions, skills, and rules. The financial agent's tone correction belonged to procedural memory. The team was not adding a fact or banning a particular word. It was adding a rule for how the agent should engage with users. Broekhuizen says this type produces most of the visible behavioral gains.

### Working memory and long-term memory have different jobs
[06:20](https://www.youtube.com/watch?v=KGFyOtl5ktI&t=380s)
Working memory contains what the agent needs during its current run, such as intermediate scratch pads, tool results, and retrieved files. Long-term memory lives in a separate store or data structure and contains instructions, skills, and other context for future turns. Different agent systems inject long-term memory in different ways. They may place it directly in the prompt, retrieve it through tools, load it through files, or change runtime state. The implementation can vary, but the distinction matters because some context is temporary and some is meant to guide continued behavior.

### The read-write cycle filters experience before storing it
[08:02](https://www.youtube.com/watch?v=KGFyOtl5ktI&t=482s)
Broekhuizen describes memory as a read-write cycle. At the start of a run, the agent reads relevant skills, instructions, or other context from long-term memory into short-term memory. The run then produces evidence, including retrieved context, tool calls, decisions, and sub-agent activity. A filter decides which evidence is useful signal and which should remain history. Only the selected signal is written back into durable context. For the financial assistant, that signal could be a rule to keep the tone measured rather than pointed or directive. Memory informs the run, the run creates evidence, and filtered evidence updates memory.

### LangSmith and Context Hub divide the memory workflow
[10:48](https://www.youtube.com/watch?v=KGFyOtl5ktI&t=648s)
Broekhuizen gives LangSmith as one concrete implementation of the cycle. LangSmith captures the agent trajectory, including tool calls, sub-agent calls, and retrieved context. An analysis process then examines traces for the signals that matter to the team. LangChain's Context Hub stores the durable context that the agent will use on future runs. In the financial assistant example, that context can include instructions, skills, markdown files, policies, and rules. The system detects patterns such as tone drift, identifies which context files should change, updates them, and lets the next run use the revised material.

### Most traces should remain history
[13:17](https://www.youtube.com/watch?v=KGFyOtl5ktI&t=797s)
Agents produce a large amount of exhaust and feedback, but Broekhuizen warns against turning all of it into memory. Most trace data should remain available as history. Some of it may support offline evaluations. Only a small subset should change the agent's memory. Storing every observation as durable context would make future reasoning harder because the agent would have too much material to process. The design task is therefore deciding what matters to the application and creating a filter that promotes only useful signal. Memory quality depends on this selection step as much as on the storage system.

### Caching can hide memory updates from long-running agents
[13:42](https://www.youtube.com/watch?v=KGFyOtl5ktI&t=822s)
Broekhuizen describes a problem he encountered with long-running background agents. The system could update memory early in an operation, but runtime state and caching meant later work did not see the updated material. The agent's behavior therefore failed to change even though a memory write had occurred. Teams need to understand what can be cached and what cannot, then check that updates are available to the future runs that should use them. Memory in the hot path must be designed alongside the runtime's access and caching behavior.

### Procedural updates need human review
[14:47](https://www.youtube.com/watch?v=KGFyOtl5ktI&t=887s)
Self-updating agents should not receive unrestricted control over every part of their memory. Broekhuizen points to procedural memory as especially sensitive because instructions, policies, and tone guidelines drive much of the agent's behavior. Changes to those materials should have a place where a person can review them before they enter the agent's hot reload path. The system can still automate detection and propose changes, but important behavior needs human interaction before the update is committed. His financial example makes the reason concrete: a small tone rule can change how the agent gives advice to users.

## Notable quotes
- "A trace, a transcript, or a log is evidence of what happened. It only becomes memory when that lesson is converted into durable context the agent can reference on future runs." (04:14)
- "Procedural memory is where actually most of the visible gains come from." (05:39)
- "Memory informs the run. The run makes the evidence. We filter that evidence to make signal, and then we use that signal to update memory." (09:52)
- "Most trace data actually should stay as history as referencable history." (13:21)
- "The changes that you make to the parts of the agent's memory that are procedural before actually committing those and putting those in the agent's sort of like hot reload path." (15:07)

## Tools & references mentioned
- LangChain
- LangChain Labs
- LangSmith
- Context Hub
- Hermes agent
- sleep-time compute
- progressive disclosure

## Who should watch
- You are building an agent and already collect traces, but its instructions and behavior do not improve between runs.
- Your agent uses long-running background processes, cached context, or a memory store in the hot path.
- You need a way to propose updates to agent instructions while keeping humans involved in changes to policies, skills, or tone rules.

## Editor's note

Jake Broekhuizen describes a financial agent whose tone drift required manual correction and regression checks. Kitaru records the full run, including model calls and tool results, so the failed interaction can be replayed after a prompt or code change. The replay uses the recorded inputs and responses, allowing the team to check whether the tone rule now holds.

Written by the AIE Talks editors (the Kitaru team), not by the speaker.

## Related talks

- [Memory Masterclass: Make Your AI Agents Remember What They Do!](https://aietalks.com/talks/memory-masterclass-make-your-ai-agents-remember-what-they-do) (Mark Bain, AIUS & Vasilia Marovitz, Cognee & Daniel Chalef, Graphiti and Zep & Alex Gilmore, Neo4j, 51:25)
- [Architecting Agent Memory: Principles, Patterns, and Best Practices](https://aietalks.com/talks/architecting-agent-memory-principles-patterns-and-best-practices) (Richmond Alake, MongoDB, 17:37)
- [Turning Agent Memory Into Skills That Work](https://aietalks.com/talks/turning-agent-memory-into-skills-that-work) (William Lyon, Neo4j, 18:41)
- [Total Recall: Agent Memory and Harness Engineering](https://aietalks.com/talks/total-recall-agent-memory-and-harness-engineering) (Ignacio Martinez, Oracle, 1:00:47)
- [Stop Renting Your AI's Memory](https://aietalks.com/talks/stop-renting-your-ais-memory) (Dylan Couzon, Qdrant, 15:17)
