# Stop Renting Your AI's Memory

Dylan Couzon, Qdrant | AI Engineer | 15:17

Source: https://www.youtube.com/watch?v=apyrzaWj0Z4
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/stop-renting-your-ais-memory
Published: 2026-10-02
Tags: edge, embeddings, memory, privacy

## TL;DR
- Owning model weights gives an AI autonomy, but owned memory gives it continuity across sessions and devices.
- Useful memory needs write, retrieve, and forget operations, with retrieval controlling which past information reaches the model.
- Qdrant Edge can keep vector memory local and offline, while optional cloud sync lets users share selected memories with consent.

## Summary
Dylan Couzon argues that local inference is becoming practical, while the memory that makes an assistant personal remains mostly controlled by cloud providers. He separates autonomy from continuity: owning model weights prevents someone else from disabling or changing the model, but persistent memory lets an assistant learn preferences, corrections, and past failures over time. He describes memory as a system that writes, retrieves, and forgets information. Retrieval lets an application select relevant memories instead of placing an entire archive in the prompt. In a live offline demo, a drone turns detected objects into embeddings and stores them with Qdrant Edge. The demo searches those memories in under a millisecond, with a 15 MB engine footprint. Couzon also describes optional cloud synchronization for shared memory between devices, agents, robots, or family members. His condition is that memory should start locally and sharing should require consent.

## Key ideas
### The question is who owns superintelligence
[00:12](https://www.youtube.com/watch?v=apyrzaWj0Z4&t=12s)
Couzon says the important question is not when superintelligence arrives, but who it belongs to. In one version, it runs in someone else's data center and users pay by the token. In the other, it runs on hardware they own, learns from their lives, keeps their context private, and cannot be switched off or throttled by a provider. He calls this second version the frontier coming home. The distinction matters because ownership determines whether the system remains available to its user or depends on a company's continued access and pricing decisions.

### A capable model is still a stranger without memory
[01:12](https://www.youtube.com/watch?v=apyrzaWj0Z4&t=72s)
An assistant that forgets almost everything when a conversation ends starts every session from zero. Couzon describes these systems as geniuses with no long-term memory. He says a model becomes personal through what it knows about its user, including preferences, corrections, and failed approaches. He argues that a larger prompt does not solve this problem. The model may already be powerful, but it remains a stranger if it cannot carry useful information from one interaction to the next.

### Local hardware is approaching last year's frontier
[01:52](https://www.youtube.com/watch?v=apyrzaWj0Z4&t=112s)
Couzon says frontier-class models are no longer limited to remote data centers. A machine costing under $2,500 can run what was essentially last year's frontier, although he is clear that this is not the absolute frontier. Open-weight models are closing the gap quickly. Owning the hardware and the weights means a provider cannot reach into the machine and turn the model off. It also prevents a provider from quietly retiring a version or degrading it when its economics change.

### Renting the AI stack creates control and cost problems
[02:47](https://www.youtube.com/watch?v=apyrzaWj0Z4&t=167s)
The compute, model, harness, and long-term memory in many current AI systems are rented. Couzon says providers can remove flagship models, retire versions, or throttle access. He also points out that agent loops can consume far more tokens than ordinary chat. Some developers spend over $5,000 a month on compute while using a $200 plan, which he calls a subsidy that will eventually disappear. Owning compute and weights restores control over the system, but it does not solve continuity without persistent memory.

### Memory needs write, retrieve, and forget
[05:21](https://www.youtube.com/watch?v=apyrzaWj0Z4&t=321s)
Couzon defines memory through three operations: write, retrieve, and forget. He rejects the idea that a large folder of Markdown files is enough, because finding the right memory is different from dumping every past detail into a prompt. Retrieval lets an application filter by topic, change relevance with recency and frequency, and control what the model sees. He compares the system to human memory, where information does not remain equally prominent forever. The supporting infrastructure has existed for years, including HNSW from 2016 and local embeddings from 2019.

### The model is the CPU, context is RAM, and memory is disk
[06:41](https://www.youtube.com/watch?v=apyrzaWj0Z4&t=401s)
Using a comparison associated with Andrej Karpathy, Couzon maps the AI system to a computer. The model is the CPU, the context window is RAM, and persistent memory is disk. RAM is fast but disappears when the session ends. Disk carries information across conversations. His concern is that the disk containing an assistant's record of its user usually sits in another company's data center, or in an unindexed file that is hard to query. He says that persistent record should belong to the user.

### Qdrant Edge keeps vector memory inside the application
[07:30](https://www.youtube.com/watch?v=apyrzaWj0Z4&t=450s)
Couzon describes a vector search engine embedded directly in an application. It opens a local store inside the process, writes embeddings with payloads, and supports offline queries in sub-millisecond time. He says it uses the same Rust core as Qdrant in the cloud. With quantization, a million memories can fit in less than a gigabyte, which makes local storage practical on a phone or Raspberry Pi. Each application has an isolated store that remains intact after the application closes and can follow the user between devices.

### The drone demo turns observations into searchable memory
[08:10](https://www.youtube.com/watch?v=apyrzaWj0Z4&t=490s)
The live demo uses a drone video running fully offline on Couzon's laptop. An object detection model labels what appears in each image, and those labels become embeddings stored as the drone's memory. The display shows the vector space and groups similar memories together. The demo had recognized 92 objects and stored over 300 vectors, while the Qdrant engine footprint was 15 megabytes. Searching for a coffee table returned every matching observation in less than one millisecond, along with images, timestamps, a definition, and the number of times it had been seen.

## Notable quotes
- "Owning inference gives you autonomy. Nobody can take the model away. But only memory gives you continuity." (04:36)
- "It's not a bigger prompt. It's a system with three verbs: write, retrieve, and forget." (05:21)
- "The model is the CPU, the context window is the RAM, and the memory is the disk." (06:26)
- "In less than one millisecond, I was able to pull up every single coffee table that this drone has seen before." (09:40)
- "Sharing should always be opt-in, never the default." (13:04)

## Tools & references mentioned
- Qdrant
- Qdrant Edge
- HNSW
- Fable
- Mythos 5
- YOLO
- Raspberry Pi
- Andrej Karpathy

## Who should watch
- You are building an agent or assistant that needs to remember user preferences, corrections, and past interactions across sessions.
- You want local vector search for an offline device, robot, drone, phone, or Raspberry Pi.
- Your team is deciding which memories should stay on-device and which can be shared with a cloud service or other users.

## Related talks

- [Memory Masterclass: Make Your AI Agents Remember What They Do!](https://aietalks.com/talks/memory-masterclass-make-your-ai-agents-remember-what-they-do) (Mark Bain, AIUS & Vasilia Marovitz, Cognee & Daniel Chalef, Graphiti and Zep & Alex Gilmore, Neo4j, 51:25)
- [Architecting Agent Memory: Principles, Patterns, and Best Practices](https://aietalks.com/talks/architecting-agent-memory-principles-patterns-and-best-practices) (Richmond Alake, MongoDB, 17:37)
- [Lessons from Studying Every Memory System](https://aietalks.com/talks/lessons-from-studying-every-memory-system) (Shlok Khemani, Independent, 19:31)
- [No Memory, No Harness: Why the Database Is the Last Line of Defense](https://aietalks.com/talks/no-memory-no-harness-why-the-database-is-the-last-line-of-defense) (Kay Malcolm, Oracle, 21:37)
- [CrabRAG: Why Automated Assistants Need Graph Memory, Not More Tokens](https://aietalks.com/talks/crabrag-why-automated-assistants-need-graph-memory-not-more-tokens) (Stephen Chin, Neo4j, 20:42)
