Teaching LLMs to Speak Spotify

Yves Raimond, Spotify, Jacqueline Wood, Spotify19:40 · Sept 2026 · 6,494 views
Thumbnail for Teaching LLMs to Speak Spotify Watch on YouTube
TL;DR
  1. 1

    Spotify is moving from ranked recommendations toward generative personalization, where users can steer experiences and inspect or edit what the system believes about their taste.

  2. 2

    Spotify's NEO training recipe adds semantic IDs to an open-weight LLM, grounds those tokens while freezing the backbone, then instruction-tunes the full model across Spotify tasks.

  3. 3

    Grounding LLM judges with user profiles and behavioral data made their decisions agree more closely with human preferences, especially for ambiguous queries.

Summary

Yves Raimond describes Spotify's move from manually curated playlists, through systems such as Discover Weekly, to generative personalization. The Large Taste Model powers experiences that users can steer with natural language, including DJ sessions, prompted playlists, editable taste profiles, and personal podcasts. Jacqueline Wood explains NEO, Spotify's four-stage method for adapting open-weight LLMs to its catalog. Semantic IDs represent catalog items, and a frozen-backbone grounding stage teaches the model those tokens without removing its language ability. Multitask instruction tuning then teaches recommendation, retrieval, explanation, and related tasks. The team found useful transfer to newer content such as audiobooks. Jacqueline also covers decoding choices and evaluation. Beam search was preferred over top-P sampling, and grounded LLM judges agreed more often with humans when given listening histories or behavioral signals. The talk gives a practical account of adapting language models to a large recommendation catalog while preserving language capability and measuring generated results.

Key ideas
00:58

Spotify's recommendation problem spans a huge catalog and many content types

Yves Raimond says Spotify has about 760 million monthly active users across roughly 184 markets. Its catalog contains more than 100 million music tracks, along with videos, podcasts, and audiobooks. That makes matching a user with content more complicated than ranking music alone. Spotify began with manual curation, where people assembled playlists for particular tastes. It later used creation and listening signals to produce recommendations at scale, including Discover Weekly, launched in 2014. Raimond presents generative personalization as the next phase, where the system also creates an experience that changes around the individual user.

02:40

Generative personalization lets users steer and inspect the recommendation system

Raimond describes a shift from personalization as guessing to personalization as reasoning. Instead of only producing a ranked list, the system can inspect whether a result fits the user and context. Spotify also wants personalization to be transparent and steerable. The Spotify DJ lets a listener redirect a music session while it is playing. Prompted playlists accept broad or detailed requests, such as bands playing in San Francisco that night or music for different parts of a run. The taste profile exposes what the system thinks it knows about a listener and lets that person correct it, add goals, or ask to explore a new genre or topic.

06:58

The Large Taste Model combines recommendation, reasoning, and generation

Spotify uses the name Large Taste Model for one system behind these experiences. Raimond says it understands historical interactions and content, combines prediction with reasoning, and lets users shape or generate experiences in real time. It supports recommendations, explanations, and generated content. The talk mentions a personal podcast experience that produces a daily brief about what is happening in the listener's community. Raimond says about one in four US Spotify Premium subscribers interact with the system each day. Deploying it across existing surfaces also produced gains in autoplay, podcast discovery, and interaction with DJ messages.

08:18

Semantic IDs give an open-weight LLM a language for Spotify catalog items

Jacqueline Wood explains that Spotify creates semantic IDs from existing content embeddings, such as podcast episode embeddings, by applying quantization to produce discrete tokens. Spotify adds those special tokens to the vocabulary of an open-weight LLM such as Qwen. The model is then trained to understand natural language and the new tokens together. For a request about a podcast on morality, the prompt can include the user's listening history as semantic IDs. The model returns a relevant podcast episode ID and a natural-language explanation of why it recommended that episode to the user.

09:33

NEO grounds new catalog tokens before tuning the model for Spotify tasks

NEO has four stages. The semantic foundation stage creates meaningful semantic ID tokens and adds them to the model vocabulary. Domain grounding learns mappings between semantic IDs and text in both directions while freezing the original model weights and embeddings. Only the new semantic ID embeddings are trained at this point. Capability induction then unfreezes the model for multitask instruction tuning on Spotify tasks such as next-item recommendation and retrieval. Spotify can optionally add post-training such as reinforcement-learning fine-tuning. Wood says the frozen stage helps prevent catastrophic forgetting of the LLM's language abilities.

11:32

Multitask tuning transfers useful knowledge to newer content types

Spotify compared a multitask model with models trained on individual tasks. Across the evaluated tasks, the multitask model matched or exceeded single-task performance, suggesting that learning transfers between tasks. The effect was especially noticeable for audiobook recommendations, which involve newer content at Spotify. The model could use knowledge learned from other catalog types, including podcasts, to make meaningful audiobook recommendations. This gives the training approach a way to help with cold-start entities, where the system has less history for a particular type of content.

12:32

Frozen-backbone grounding preserves language ability better than continued pre-training

Wood says removing the domain-grounding stage or combining it with capability induction degraded performance. A randomly initialized backbone produced the largest performance drop. Continuous pre-training caused only a small decline on task-specific measures, but it nearly removed the original model's natural-language and world-knowledge abilities. Frozen-backbone domain grounding retained those abilities while teaching the model semantic IDs. Spotify reproduced the finding with both Qwen and Llama, so Wood presents it as a property of the training method rather than of one particular backbone.

14:02

Beam search is preferred, while constrained decoding is reserved for targeted cases

Spotify tested beam search with and without constrained decoding, along with top-P sampling. Without constrained decoding, the model still produced valid semantic IDs 98% of the time. Constrained decoding adds latency, but it is useful when the system must restrict recommendations to a type of content, such as new releases. Top-P sampling significantly reduced accuracy. Spotify therefore accepted the extra latency of beam search. The result is a decoding policy that relies on the model's learned validity most of the time and applies stricter constraints when the request needs them.

16:11

Grounded LLM judges are more useful for ambiguous recommendation evaluations

Traditional offline metrics can show whether a user interacted with recommended content, but they do not fully assess whether the recommendation fits the user's intent or whether its explanation is accurate. Spotify grounds LLM judges with data about the user and task. Textual profiles summarizing listening history produced 75% agreement with human preferences for podcast recommendation evaluation. For search, adding similar queries and the user's past behavior raised overall agreement by 5%, while agreement on ambiguous queries rose by 91%. Spotify also used grounded judges to scale Cranfield-style evaluation collections, reaching an agreement value of 0.87 with human system rankings.

"The phase that we're entering now is something that we call generative personalization where we are not only solving a matching problem from the user to the content. We're also solving the ability to generate an experience that is interactively and dynamically shaped around each user."Yves Raimond02:40
Who should watch
  • You are adapting an open-weight LLM to a recommendation catalog and need a concrete training recipe for grounding catalog entities.
  • Your product needs natural-language controls, explanations, or generated experiences on top of an existing personalization system.
  • You are evaluating generative recommendations and need to understand how user history and behavior can improve LLM-judge agreement with people.