# Stop Fine-Tuning to Fix Retrieval Problems

Anant Srivastava, Oracle | AI Engineer World's Fair 2026 | 20:12

Source: https://www.youtube.com/watch?v=qflLT3SoVbw
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/stop-fine-tuning-to-fix-retrieval-problems
Published: 2026-10-04
Tags: fine-tuning, memory, prompt-engineering, rag

## TL;DR
- Prompt edits, indexed documents, and training samples all decide where an AI system's behavior or knowledge lives.
- Current, large, citable, or access-controlled knowledge belongs in memory, while prompts should hold small and stable behavior.
- Fine-tuning fits stable reflexes and can reduce inference cost, but it should not be used to hide retrieval failures.

## Summary
Anant Srivastava argues that enterprise teams often build AI architecture by accident. They escalate from prompt edits to retrieval changes to fine-tuning without asking where a piece of knowledge should live. He separates the jobs of prompts, memory, and weights. Prompts hold small, stable instructions about tone and behavior. Memory holds knowledge that changes, needs citations, is too large for a prompt, or requires access control. Weights hold stable patterns and reflexes, especially when a smaller fine-tuned model can reduce cost. A support assistant illustrates the danger: product information leaked into the weights through fine-tuning, so the assistant continued inventing obsolete product names after the catalog changed. Srivastava closes with a circulating architecture. Signals from work can become durable memory, repeated patterns can later become model reflexes, and fine-tuning changes what needs to be retrieved. The model matters, but the surrounding system decides how information moves.

## Key ideas
### Prompt, memory, and weights solve different problems
[00:46](https://www.youtube.com/watch?v=qflLT3SoVbw&t=46s)
Srivastava says teams often treat prompt edits, retrieval, and fine-tuning as steps on an escalation ladder. He argues that they are separate tools with separate jobs. A prompt controls behavior such as tone, persona, and instructions. Memory supplies knowledge during inference. Weights encode patterns that have become stable enough to learn. This distinction changes how engineers debug an agent. Instead of adding information wherever the latest failure appears, they should ask what kind of information it is and how quickly it changes.

### Every content change is an architecture decision
[02:06](https://www.youtube.com/watch?v=qflLT3SoVbw&t=126s)
An edited prompt determines which behavior stays in the prompt. An indexed document determines which knowledge stays in memory. A training sample puts information into the model weights. Srivastava describes these as architecture decisions even when teams make them during ordinary product work. Six months of small fixes can therefore produce a system whose information boundaries nobody designed or owns. The problem is not only technical ownership. Each team may own one change while no one owns the complete path by which knowledge reaches inference.

### Fine-tuning can make stale product knowledge persist
[02:40](https://www.youtube.com/watch?v=qflLT3SoVbw&t=160s)
Srivastava describes an internal support assistant whose prompt was changed for tone, whose retrieval system received refund policies, and whose prompt later received a product catalog. The machine learning team then fine-tuned the model on support tickets from as far back as six months. When a new catalog launched, replacing the catalog in the prompt did not fix the answers. Product names from the old catalog had leaked into the weights through the support-ticket fine-tune. The assistant continued producing names that did not exist.

### Prompts should hold small, stable behavior
[05:14](https://www.youtube.com/watch?v=qflLT3SoVbw&t=314s)
The prompt is appropriate for tone, persona, escalation rules, and other authored instructions that do not change per query or user. Srivastava's support example uses a professional tone, offers human escalation after three failures, and gives concrete next steps. These instructions are small, stable, and about how the agent should behave. A product catalog is a poor prompt payload because it adds context cost and can contribute to the lost-in-the-middle problem. The diagnostic question is whether the information is small, stable, and behavioral.

### Memory holds changing knowledge with retrieval and access control
[06:59](https://www.youtube.com/watch?v=qflLT3SoVbw&t=419s)
Memory includes agent memory and external enterprise memory retrieved through systems such as RAG. It is the right home for knowledge that is current, large, or citable. Access control is another reason to keep information in memory. A code assistant should retrieve from repositories rather than learn an organization's code through fine-tuning. Its retrieval system should use code-aware chunking, filters, and metadata such as the repository and the users allowed to commit. Without that filtering, retrieval can return too many unrelated chunks and confuse the model.

### Weights should capture stable reflexes, not changing facts
[11:23](https://www.youtube.com/watch?v=qflLT3SoVbw&t=683s)
Srivastava defines the rate of change as the main test for putting information into weights. Runbooks, procedures, and other documentation belong in external memory when the model lacks the right answers, because that is usually a retrieval problem. Fine-tuning can make sense after human corrections reveal a stable center in an ambiguous task such as claims processing or content moderation. The team should monitor that center for drift. In medical coding, the codes themselves remain knowledge to retrieve, while the model can learn the reflex of interpreting input formats and selecting among them.

### A circulating architecture moves patterns between memory and weights
[17:33](https://www.youtube.com/watch?v=qflLT3SoVbw&t=1053s)
Srivastava describes a system in which the prompt or context window produces signals that can become durable memory. Later sessions retrieve that memory into context. Over time, repeated retrieval patterns and stable formats can move from memory into fine-tuning. Once a model has learned a format reflexively, the system no longer needs to retrieve every example of that format. The flow also changes what is worth retrieving. This creates a loop in which the agent improves through normal work, provided the system has a harness that manages the movement.

## Notable quotes
- "Every edit that you make to the prompt decides what behavior lives in the prompt." (02:06)
- "The job of the prompt is behavior, is tone, is the persona of the agent." (05:14)
- "You fine-tune your model to learn reflexes, not facts typically." (09:40)
- "If your model is not giving the right answers, fix the retrieval problem. Don't just go ahead and fine-tune your model." (13:41)
- "The model is the easy part. What you build around the model, the harness around the model that helps you store the right information at the right place and circulate among them is the architecture that you got to build." (19:27)

## Tools & references mentioned
- Oracle
- RAG
- Mount Sinai
- IMO Health
- ICD-10

## Who should watch
- You are deciding whether a bad answer needs a prompt change, better retrieval, or a model update.
- Your team is fine-tuning on documentation, support tickets, or product data that changes over time.
- You are building a code assistant or another agent where repository scope, citations, and user access must affect retrieval.

## Related talks

- [The GenAI Maturity Curve, or You Probably Don't Need Fine Tuning](https://aietalks.com/talks/the-genai-maturity-curve-or-you-probably-dont-need-fine-tuning) (Kyle Corbitt, OpenPipe, 18:03)
- [How We Taught Agents to Use Good Retrieval](https://aietalks.com/talks/how-we-taught-agents-to-use-good-retrieval) (Hanna Lichtenberg, Mixedbread, 14:28)
- [Agent Reinforcement Fine-Tuning](https://aietalks.com/talks/agent-reinforcement-fine-tuning) (Will Hang & Cathy Zhou, OpenAI, 16:55)
- [On AI and Knowledge](https://aietalks.com/talks/on-ai-and-knowledge) (Pablo Castro, Microsoft, 17:35)
- [User Signal Dies at the Retrieval Boundary](https://aietalks.com/talks/user-signal-dies-at-the-retrieval-boundary) (Sonam Pankaj, StarlightSearch Inc, 15:37)
