Recommendation systems follow scaling curves similar to LLMs, and most of the industry is still early on those curves.
2
An LLM recommender can be built by creating semantic IDs, training the model across English and the catalog language, then post-training it to rank recommendations.
3
Feeds can produce more engagement per inference cost than chat apps because they decode pointers to existing content instead of generating every content token.
Summary
Devansh Tandon argues that recommendation systems scale with more data, compute, and model capacity in much the same way as LLMs. He describes four stages of this progression: traditional recommenders, LLM-inspired models, LLM-native recommenders, and agentic systems that plan, retrieve, rank, and critique recommendations. His proposed LLM recommender uses semantic IDs to compress content, pre-training to connect those IDs with English, and post-training to rank candidate items with readable reasoning. This creates feeds that can explain their choices and respond to natural-language instructions, such as adding or removing interests. Tandon also compares feeds with AI chat apps. A feed returns a semantic ID that points to creator-supplied content, while a chat app generates each output token. He argues that this makes feeds much cheaper per hour of engagement and gives LLM recommenders a path into some of the largest consumer applications.
Recommendation systems improve along predictable scaling curves
Tandon says recommendation systems follow a power-law scaling pattern similar to LLMs. Increasing data, compute, and model size improves offline recommendation quality, measured through metrics such as net entropy or AUC, and can lead to higher engagement or revenue in production. He points to Meta research showing this relationship in recommender models. The claim matters because recommendation models already run at the scale of more than a billion daily active users, so improvements in their scaling curve can affect large consumer products.
Meta connects larger recommendation models with higher Reels engagement
Tandon points to Meta's recent earnings reports as a production example. Instagram Reels had 30% year-on-year watch-time growth, and the recommendation improvements included a simpler ranking architecture that allowed more efficient model scaling and longer interaction histories. Meta doubled the length of the user interaction sequences used for Instagram training and made each interaction richer. Tandon presents these changes as the product-level version of the same data, compute, and model scaling pattern seen in LLMs.
Recommendation products are moving through four model paradigms
The first paradigm uses feature engineering, user and content embeddings, and two-tower or sparse-network rankers. The second scales models more end to end, with HSTU and OneRec as examples. The third adapts a foundation model that understands and reasons for recommendation tasks, which Tandon calls LLM-native. The fourth is agentic: models plan, retrieve, rank, and critique recommendations in a loop, then call models again to refine the result. Tandon says companies are climbing these curves in parallel.
Semantic IDs compress content and give models a stable vocabulary
Tandon's first construction step is to replace unstable hash IDs with semantic IDs. A semantic ID gives the model a stable representation of content and lets it reason over long interaction histories more efficiently. He says a three-minute Instagram Reel could occupy about 10,000 tokens without compression, but can be represented in about 10 semantic tokens. Similar tennis Reels can share an initial token prefix, while later tokens distinguish the individual videos.
The recommender must learn both English and the catalog's semantic language
After creating semantic IDs, the model is pre-trained to connect them with natural language and to understand sequences of user interactions. One example pairs a video's semantic ID with a missing description that the model must complete in English. Another masks items in an interaction sequence and asks the model to predict them. Tandon describes this as making the model bilingual: it understands ordinary language and the token language used to represent the recommendation catalog.
Post-training turns the foundation model into a readable re-ranker
The post-training stage teaches the model to rank content for a user. Tandon shows an input containing user information and 30 candidate videos, with the model selecting the top five. He says the model's chain-of-thought reasoning exposes interests such as comedy, food, DIY, and wellness, along with engagement style and creator affinity. That reasoning gives the product a way to inspect why items were ranked, although the talk does not describe a separate evaluation of the explanations.
Natural-language controls can make feeds steerable
Tandon describes Instagram's Your Algorithm as an example of a feed that users can discuss while consuming content. Users can see what the algorithm thinks they are interested in, add or remove interests, and communicate in natural language. He expects recommendation systems to move away from being black boxes toward systems that users can direct toward goals expressed in language. He also describes an example where adding an interest in the FIFA World Cup changes the personalized recommendations.
Feeds spend inference on pointers while chat apps generate content
Tandon compares content feeds with AI chat apps using the same training, inference, engagement, and monetization flywheel. He says feeds typically use models with around 1 to 10 billion active parameters, while the chat apps he names serve models with around 10 to 100 billion active parameters. A feed outputs a semantic ID that points to existing creator-supplied content. A chat app generates every output token, often producing a few thousand tokens in a turn. Tandon therefore estimates that feeds can cost up to 100 times less than chat apps for an hour of engagement.
LLM recommenders target some of the largest consumer applications
Tandon notes that four of the ten most-used apps are content feeds. He says recommendation and advertising models drive much of the engagement and monetization in those apps, so replacing them with LLM recommenders could affect a very large consumer surface. The product changes he expects include steerable feeds, explanations, interactive recommendations, and consumer agents. His argument rests on both audience scale and the lower inference cost of returning pointers to existing content.
"And so we're going to see this shift I think from blackbox recommendations algorithms to giving users more control over their algorithm and algorithms becoming more interactive and steerable."12:45
Who should watch
You are building recommendation systems and want a model-level path beyond feature engineering and embedding-based rankers.
You work on a consumer AI product and need to compare the inference economics of generated answers with feeds that retrieve existing content.
You are exploring natural-language controls, explanations, or agentic loops for recommendation products.