Turning Agent Memory Into Skills That Work

William Lyon, Neo4j18:41 · Oct 2026 · 4,619 views
Thumbnail for Turning Agent Memory Into Skills That Work Watch on YouTube
TL;DR
  1. 1

    Retrieval finds relevant text, but agents also need connected, typed, canonical, and traceable knowledge.

  2. 2

    A context graph can combine short-term, long-term, and reasoning memory, including the evidence and tool calls behind agent decisions.

  3. 3

    Skills can be distilled from grounded memory graphs as governed execution graphs, with checks for grounding, coverage, coherence, and staleness.

Summary

William Lyon argues that agent memory built from embeddings and retrieval does not give agents actionable knowledge. Retrieved text may mention the same entity in different ways, and it does not by itself explain how facts connect or how a previous decision was made. Neo4j's approach stores short-term, long-term, and reasoning memory in a context graph. The graph resolves entities, applies a shared ontology, and records decision traces with evidence, policies, tool calls, results, timing, and token use. Lyon then explains how this memory can produce reusable skills. Neo4j research models skills as typed execution graphs rather than prose files. Its Agent Memory Service distills skills from selected conversations or graph scopes, checks grounding, coverage, and coherence, and tracks changes that could make a skill stale. In the demo, a healthcare conversation produces a patient-intake skill whose steps are linked to the tool calls and entities behind them.

Key ideas
00:38

Retrieval gives an agent relevant text, not a usable representation of knowledge

Lyon describes an amnesiac loop in which an agent completes a task and then forgets what it learned. Current memory systems often embed text, retrieve similar chunks, and place them in the context window. That does not resolve whether references such as "Dr. Nwin," "Robert N," and "the cardiologist" identify the same person. Agents need knowledge that is connected, typed, and traceable. Retrieval remains part of the system, but it cannot by itself provide a canonical representation or explain how information should be used.

02:13

Agent memory combines short-term, long-term, and reasoning memory in a context graph

Neo4j's model has three parts: short-term memory, long-term memory, and reasoning memory. Messages and extracted entities form part of the graph, while reasoning memory records what the agent actually did. The system extracts entities and their relationships from user and assistant messages, then gives those relationships strong types. A shared ontology describes the domain and its data model. This turns unstructured conversation data into a graph that can be queried and governed.

03:06

Entity resolution and shared ontologies give graph data a consistent meaning

Entity resolution is one of the most important steps in building the knowledge graph. The system needs a canonical representation for an entity and should know attributes such as a provider's name and role. Lyon also stresses the need for a shared ontology that describes the domain being modeled. In the healthcare example, the ontology includes encounters and providers. These choices determine whether later skills refer to stable entities and well-defined relationships instead of disconnected text mentions.

04:01

Decision traces preserve the evidence and actions behind an agent's choices

Reasoning memory captures more than stored facts. Each decision can be linked to the reasoning and evidence that supported it, along with explicitly modeled policies, the execution plan, tool calls, and tool results. Lyon also mentions recording token use and timing. These decision traces let a system inspect how an agent decided. They can help the same agent repeat a task and can also be shared with other agents that use the same tools.

05:22

A shared context graph lets many agents learn from previous runs

Persisted agent runs and decision traces can be stored in one shared context graph for an organization. Lyon describes a setting with hundreds or thousands of agents that share groups of tools. In that setting, one agent's recorded execution can become useful context for another agent. The purpose is broader than improving one agent on its next attempt. The graph gives multiple agents access to the decisions, evidence, and outcomes accumulated by their peers.

06:07

Skills turn recorded memory into executable procedures

Lyon presents skills as the next step after representing what happened in memory. The Agent Skills open standard uses metadata, descriptions, and progressive disclosure, with more detailed information loaded when needed. Skills can then be packaged and reused across agents. Lyon says prose-based skills still make it difficult to identify canonical entities, debug individual steps, and understand expected outputs. A memory graph can provide the underlying data for creating skills that are grounded in observed work.

08:26

Typed execution graphs make skills easier to describe and govern

Neo4j research on APE, described as an extension of the Agent Skills protocol, models skills as typed execution graphs. Steps become graph nodes, and a strict schema describes how those steps can be executed. Lyon refers to a SkillsBench benchmark in which applying the APE protocol to human-curated skills produced a significant increase in successfully completed tasks. The graph representation supplies structured steps and governed descriptions instead of relying only on a markdown document.

10:05

Skill distillation checks grounding, coverage, coherence, and staleness

The proposed distillation process takes a scoped context graph and produces a skill grounded in observed data. Lyon wants the resulting procedure to have well-understood steps and to remain governable as its source data changes. Grounding checks whether skill content comes from the memory system. Coverage checks whether the steps and descriptions are grounded throughout the skill. Coherence can reveal that a selected subgraph mixes several topics, suggesting that the material should become multiple skills. If source data changes, disappears, or becomes contradictory, the system can flag the skill as stale or drifted.

14:39

NAMS distills a patient-intake skill from a healthcare conversation

The Neo4j Agent Memory Service dashboard exposes REST APIs and MCP tools for ingesting and retrieving short-term, long-term, and reasoning memory. Lyon shows a healthcare ontology and ingested conversations, then selects a recent conversation as the scope for distillation. In a previously completed run, the service extracted the steps for patient intake through charting. Each step is grounded in the tool calls and entities that produced the underlying skill components. The result is packaged as a skill MD file with additional references that follow progressive disclosure.

"Each one of these steps is grounded in the actual tool calls and the entities that constructed the underlying components of the skill."17:09
Who should watch
  • You are building agent memory with embeddings and retrieval, but agents still confuse entities or fail to reuse what they learned.
  • You need several agents to share decisions, tool results, and domain knowledge from previous runs.
  • You are designing reusable agent skills and want them grounded in source data, represented as executable steps, and checked for staleness.