Where RL Will Take Search

Maximilian-David Rumpf, SID.ai09:36 · Sept 2026 · 5,560 views
Thumbnail for Where RL Will Take Search Watch on YouTube
TL;DR
  1. 1

    Agentic search finds the right documents more often than classical search, but it currently costs far more and takes much longer.

  2. 2

    Reinforcement learning lets a search model choose its own searches, filters, and follow-up steps instead of following a fixed pipeline.

  3. 3

    A specialized RL search model can take about five seconds instead of two minutes and cost about one hundredth as much as a frontier model.

Summary

Maximilian-David Rumpf argues that search is a strong use case for reinforcement learning because its reward can be checked directly: a model either finds the correct document or it does not. Classical search uses a fixed chain of query rewriting, backend retrieval, reranking, and result delivery. A reranker may recognize that the results are inadequate, but it cannot search again. An RL model can interact with the database, read results, apply filters, try another search, and spend more computation on difficult questions. Rumpf compares this shift with the move from hand-designed systems to learned systems in computer vision and chess. SID.ai trains a specialized model for search and reports roughly five-second execution times instead of two minutes, at about one hundredth of the cost of a frontier model. A search sub-agent can also keep poor results out of the main agent's context window.

Key ideas
00:01

Agentic search improves retrieval while consuming too much time and money

Rumpf says agents are roughly twice as likely to find the right documents as classical search queries. The tradeoff is substantial: agentic search costs about one hundred to a thousand times more and takes minutes rather than milliseconds. Searching also uses 30 to 50 percent of an agent's tokens, usually at the start of a task, before the agent performs the requested work. His proposed direction is to pass this work to a dedicated search sub-agent trained specifically for retrieval.

02:16

Fixed search pipelines cannot recover from bad results

The classical pipeline may rewrite a question, send it to a search backend, rerank the results, and return them. Its models are optimized locally, and its decisions are fixed at design time. Each question receives a fixed amount of computation. Rumpf points out that a reranker may know the returned documents do not answer the question, yet it has no way to search again. Unexpected questions therefore create a long tail of failures, which teams patch with edge cases that cannot cover every case.

03:17

Learned systems can replace layers of hand-designed search logic

Rumpf compares search with computer vision and chess. Computer vision moved from primitive edge detection to narrow object detectors and then to models that handled more of the full task. Chess moved from human-written rules in IBM Deep Blue, through Stockfish, to AlphaZero and MuZero, which learned strategies inside the model. He expects search to follow a similar path, from BM25 and PageRank through vector search and rerankers toward a model that makes its own search decisions.

04:21

An RL search model can adapt its work to each question

In Rumpf's description, one model interacts with a database, searches, reads results, iterates, sets metadata filters, and searches again until it is satisfied. It then produces a ranked result list. The model can use more computation for a difficult question instead of applying the same budget to every query. Rumpf says search fits reinforcement learning because the outcome is verifiable and the training environment can run thousands of attempts per second.

05:48

Specialization makes the search model faster and cheaper

Rumpf says a general language model contains capabilities that are unnecessary for search. He compares this with hardware: a CPU can perform tasks a GPU can perform in theory, but a CPU is not the right choice for language-model inference. Training a specialized search model produces a system whose quality rises with compute, while latency and retrieval rewards can be included during training. The model is not given a fixed search recipe. It discovers its own retrieval strategies.

06:46

SID.ai reports five-second search at about one hundredth the cost

Rumpf presents SID.ai's results against vector and reranker baselines and frontier models. The specialized model takes around five seconds on average instead of around two minutes, making it about twenty times faster. He also says it is about one hundred times cheaper for the task. It has not reached the latency of a vector-and-reranker pipeline, but Rumpf says the team thinks it can move closer to that speed.

07:27

A search sub-agent keeps poor retrieval out of the main context

In a normal agent trace, good and bad search results enter the same context window. Rumpf says the bad results pollute the agent's context. A separate search sub-agent can perform the searching, reading, and iteration, then pass only strong results to the main agent. This gives the main agent more useful material and moves the 30 to 50 percent of tokens spent on searching to a search system that Rumpf says is about one hundred times cheaper.

08:35

Rumpf expects RL search to reach private data and new interfaces

Rumpf predicts that scaling reinforcement learning will produce better search across domains, while faster and cheaper models will make search useful in settings such as voice and e-commerce. He also argues that the web contains only a small share of the available information. Operational knowledge, such as how to run JPMorgan, is not published on the web and instead exists inside the company's databases.

"The re-ranker might know that the results are insufficient at answering the question, but the re-ranker can't take action."02:58
Who should watch
  • You are building an agent whose context window is filling with search attempts and irrelevant documents.
  • Your retrieval stack relies on fixed query rewriting, vector search, or reranking and cannot retry when results are inadequate.
  • You want to understand where a specialized RL search model might fit when latency and retrieval cost limit agent use.