# Distill the LLM, Don't Serve It: Search & Personalization at DoorDash

Raghav Saboo, DoorDash | AI Engineer | 22:11

Source: https://www.youtube.com/watch?v=ACPEpji5NV4
Channel: AI Engineer (https://www.youtube.com/@aiDotEngineer). Summarised by AIE Talks.
Page: https://aietalks.com/talks/distill-the-llm-dont-serve-it-search-personalization-at-doordash
Published: 2026-09-25
Tags: distillation, embeddings, memory, search

## TL;DR
- Marketplace discovery depends on understanding item meaning and shopper intent, since engagement signals can rank popular products that violate a shopper's constraints.
- DoorDash uses LLMs offline to create relevance labels, semantic IDs, and consumer memory, then distills those outputs into smaller models for retrieval, ranking, and personalization.
- LLM-generated personalized collections, built from consumer memory and catalog semantics, increased order rate by close to 1% in DoorDash's pets category.

## Summary
Raghav Saboo argues that marketplace search is mainly a semantic understanding problem. Engagement signals can make regular spaghetti outrank gluten-free pasta because the regular product sells more, even though it fails the query constraint. DoorDash uses LLMs to reason about relevance and shopper intent offline, then distills that reasoning into production systems. The talk covers four primitives: graded LLM supervision for retrieval and ranking, semantic IDs that give catalog items learned hierarchical relationships, consumer memory stored as text, vectors, and graphs, and steerable content generation for personalized collections. These shared representations support several models without requiring an LLM call at serving time. Saboo describes improvements including a 2.3% gain in relevance NDCG for retrieval, a 4–5% MRR gain in ranking from semantic IDs, and close to a 1% order-rate increase for personalized collections in pets. The approach keeps existing ranking and business objectives while adding semantic signals.

## Key ideas
### Semantic understanding is the bottleneck in marketplace discovery
[00:36](https://www.youtube.com/watch?v=ACPEpji5NV4&t=36s)
Saboo says DoorDash historically treated discovery as an engagement optimization problem, but the harder problem is understanding what an item means in context and what the shopper intends. This matters as DoorDash expands beyond restaurants into grocery, retail, pets, and gifting. A shopping mission can start with a query such as "I just adopted a puppy. What do I need to get started this week?" and continue through searches, collections, related products, and several sessions. The system therefore needs shared representations that connect the shopper's broader journey.

### LLMs provide scalable relevance supervision
[03:17](https://www.youtube.com/watch?v=ACPEpji5NV4&t=197s)
For a query such as "gluten-free pasta," engagement ranking can favor regular spaghetti or retrieve gluten-free bread. Saboo's target is graded relevance: gluten-free pasta is highly relevant, chickpea pasta is a possible substitute, and regular spaghetti fails the constraint. Human labels are expensive and become stale, while behavioral signals are affected by exposure, position, price, promotions, and earlier model choices. DoorDash builds a seed set with three relevance levels, audits suspicious cases with a stronger LLM, reconciles labels with category models, and fine-tunes a lightweight LLM labeler for the full catalog.

### Offline LLM reasoning can improve retrieval without replacing the serving stack
[05:59](https://www.youtube.com/watch?v=ACPEpji5NV4&t=359s)
DoorDash runs the fine-tuned labeler offline to generate graded query-item pairs, then uses those labels for retrieval and ranking. Retrieval is trained in two contrastive stages. The first shapes global geometry with two-tower encoders and multi-level supervised contrastive loss. The system then mines hard negatives and trains on them in a second stage. Saboo reports a 2.3% improvement in relevance NDCG. For ranking, DoorDash adds an ordinal relevance tower beside click, add-to-cart, and conversion towers, then blends relevance and engagement through a value function.

### Semantic IDs give the catalog a learned, fine-grained structure
[09:39](https://www.youtube.com/watch?v=ACPEpji5NV4&t=579s)
DoorDash has billions of store-level items, while store IDs carry no semantic meaning and its human-curated taxonomy can be too coarse. Semantic IDs assign each item a short hierarchical code. Early tokens capture broad neighborhoods, while later tokens capture fine distinctions, such as different regional types of hot sauce. These IDs allow models to compare items across taxonomy branches, add new items with sparse ID features, and transfer signal to long-tail products. They also let DoorDash audit how its human labels compare with the learned structure.

### Semantic IDs improve ranking and query reformulation
[12:55](https://www.youtube.com/watch?v=ACPEpji5NV4&t=775s)
Saboo says semantic IDs improved ranking MRR by 4–5% and also produced conversion gains. They support query reformulation by grounding related queries in catalog neighborhoods. A query for sriracha can lead to garlic chili sauce or sambal oelek when those products occupy a related semantic region. The system can avoid suggesting queries that have no inventory on DoorDash, since those suggestions would not lead to useful shopping results.

### Consumer memory is stored at several timescales and in several formats
[14:03](https://www.youtube.com/watch?v=ACPEpji5NV4&t=843s)
DoorDash applies memory ideas from agent systems to recommendations. Long-term memory captures durable preferences from orders, searches, browsing, and support interactions. Session context captures the active search and cart state, while stated preferences capture constraints that consumers explicitly give through interactions such as Ask DoorDash. Memory blocks can cover dietary preferences, dining preferences, and substitution preferences. Each block is materialized as readable text, an embedding for retrieval and ranking, and graph or tree relationships linking consumers to brands, taxonomies, and concepts.

### Context graphs connect sparse shopper interactions
[17:04](https://www.youtube.com/watch?v=ACPEpji5NV4&t=1024s)
Saboo describes context graphs that connect consumers with extracted memory concepts. They fit shopping journeys where interactions are sparse and can span multiple steps. The graphs create links between consumers and fine-grained taxonomy concepts that were previously disconnected. DoorDash uses graph-based embeddings in retrieval, where Saboo says they outperform existing taxonomy-based embeddings. The memory framework also feeds personalized collections, agentic personalization, and downstream retrieval and ranking models.

### Offline generation makes personalized collections steerable
[18:20](https://www.youtube.com/watch?v=ACPEpji5NV4&t=1100s)
DoorDash combines semantic IDs, memory blocks, and graded relevance with LLMs, small language models, and traditional models to generate ranked item lists, collection titles, and supporting copy. For store pages, an offline LLM process uses consumer memory and semantic IDs to generate a collection title, subtitle, and item set. Existing retrieval and ranking systems still hydrate the items and rank the collections. The system can create plant-based pantry rows, cat food rows for cat-only households, or pantry-staple rows during a restocking session. In pets, early tests produced close to a 1% order-rate increase and a 6% increase in active users.

## Notable quotes
- "The real bottleneck is semantic understanding." (00:36)
- "Use expensive reasoning once offline then distill it into models that can be served cheaply and quickly." (05:59)
- "We're not replacing the retrieval and ranking system with an LLM, but rather distilling the reasoning and understanding into some production ranking architecture." (08:52)
- "The online LLM call is often not the product architecture you need." (21:23)

## Tools & references mentioned
- DoorDash
- Ask DoorDash
- LLMs
- GPT-4o mini
- semantic IDs
- NDCG
- MRR

## Who should watch
- You are building search or ranking for a marketplace where clicks and conversions reward popular items that do not satisfy the shopper's actual constraints.
- Your catalog is large, frequently changing, or sparse in its long tail, and you need semantic relationships that traditional taxonomy and item IDs do not provide.
- You want to use LLM reasoning in production while keeping online latency and existing retrieval, ranking, and business-objective systems under control.

## Related talks

- [Transforming search and discovery using LLMs](https://aietalks.com/talks/transforming-search-and-discovery-using-llms) (Tejaswi Tenneti & Vinesh Gudla, Instacart, 21:10)
- [Improving Recommendation Systems and Search in the Age of LLMs](https://aietalks.com/talks/improving-recommendation-systems-and-search-in-the-age-of-llms) (Eugene Yan, Amazon, 20:54)
- [Why LLM Recommenders Will Be AI's Biggest Consumer App](https://aietalks.com/talks/why-llm-recommenders-will-be-ais-biggest-consumer-app) (Devansh Tandon, Meta, 18:01)
- [360Brew: LLM-based Personalized Ranking and Recommendation](https://aietalks.com/talks/360brew-llm-based-personalized-ranking-and-recommendation) (Hamed Firooz & Maziar Sanjabi, LinkedIn, 22:00)
- [Personalization in the Era of LLMs](https://aietalks.com/talks/personalization-in-the-era-of-llms) (Shivam Verma, Spotify, 20:12)
