Spotify is moving from separate recommendation pipelines toward a unified generative model that can respond to user requests.
2
Semantic IDs compress catalog vectors into short token sequences, allowing an LLM to predict the next track, artist, or episode autoregressively.
3
A soft tokenization layer projects each user's embedding into the LLM's token space, and this approach is already used for podcast next-episode recommendations in production.
Summary
Shivam Verma explains how Spotify is adapting open-weight LLMs for personalized recommendation. The system starts with user embeddings trained from listening history. Spotify then represents tracks, artists, and episodes as Semantic IDs, short token sequences derived from catalog vectors. This lets the model treat recommendation as sequence generation, where the next song or episode is the next token-like item to predict. Since the LLM cannot be trained directly on every Spotify user, Spotify projects an individual user embedding into the model's token space as a soft token. The frozen or adapted model can then attend to a representation of that user's taste while combining it with catalog and world knowledge. Verma describes this as a move away from separate candidate-generation and ranking systems. Podcast next-episode recommendations are already running on the approach, while Spotify is also developing more user control through its taste profile product.
Spotify is replacing separate recommendation pipelines with a more unified generative system
Traditional recommendation systems first reduce a catalog of millions of items to a few hundred candidates, then rank those candidates, sometimes with several ranking stages. Spotify uses this pattern across home shelves, playlists, search, podcasts, and ads, with separate teams and models. Verma says the company is moving toward a single model with an LLM backbone. The user can steer this system toward the kind of recommendation they want. Natural-language features such as AI DJ, prompted playlists, and the taste profile point toward this model of interaction.
User embeddings turn a person's long-term listening history into a reusable representation
Spotify's user representations team builds embeddings from a user's history across sessions and years. These vectors become the foundation for downstream recommendation and search models. The embedding pipeline generates representations for more than a billion users every day, which Verma describes as expensive and large-scale. Spotify has moved from generalized representations, including an autoencoder that compresses and reconstructs user features, toward sequential transformer models. These models put user interactions into the model context alongside the request, product surface, and recommended item.
A shared embedding space places users, tracks, and episodes alongside one another
Spotify's newer sequential models embed users, tracks, and podcast episodes in the same space. Verma shows a visualization with tracks in blue, episodes in pink, and users in green. A user's position indicates which content is nearby and which parts of the catalog form a useful neighborhood. In his example, his embedding is close to a large technology podcast because he follows machine learning and technology topics. The model can use these relationships to connect a user's taste with different content types.
Semantic IDs let an LLM learn Spotify's catalog through short token sequences
Spotify converts a content vector, such as the vector for a track or episode, into a Semantic ID made from four or six tokens. This compresses a large vector into a form that can be used in ordinary LLM training. The model can then generate the next item autoregressively, treating the next song or episode like the next token in a sequence. Verma gives Ariana Grande and Bruno Mars as an example. Their six-token IDs share the first two tokens because both are pop artists, while the remaining tokens differ and capture more specific properties.
Semantic ID sequences give the catalog a hierarchical structure
The shared prefixes in Semantic IDs are more than a storage trick. Verma describes the representation as hierarchical: early tokens capture broader similarities and later tokens capture narrower distinctions. Spotify uses the user's context and tokenized listening history to train the LLM to work with these IDs. In an example involving an Italian listener and an Italian podcast, the model converts Spotify's URI for an episode into Semantic ID tokens and uses them to predict the next episode.
Soft tokens add a specific user's taste to an LLM without training on every user
Spotify cannot train the LLM directly on all 750 million users, so the model needs a separate personalization step. Spotify takes a user's embedding and projects it into the LLM's representation space. The result is a soft token that represents that user inside the model. This token changes from user to user and is inserted into the prompt alongside the normal input. When the model generates a recommendation, it can attend to this user-specific representation rather than relying only on generalization from its training data.
The three components form a sequential generative recommender
Verma brings together user embeddings, Semantic IDs, and soft tokenization. User embeddings represent the listener. Semantic IDs provide a compressed token representation of catalog items. Soft tokenization projects the listener into the LLM's token space. Together, these components move Spotify from conventional recommendation pipelines toward sequential modeling, where the system generates a next item. Verma says the approach has shown positive internal metrics, and podcast next-episode recommendations are already productionized on a system of this kind.
"That allows the LLM to kind of auto-regressively generate the next token or in this case the next token is not a word, but it's the next song or it's the next episode that you're going to be listening to."13:43
Who should watch
You build recommendation systems and want to understand how an LLM can replace parts of candidate generation and ranking.
You work on personalization and need a practical pattern for injecting user representations into an otherwise general language model.
You are designing catalog-aware generation for tracks, episodes, or other large item collections where item IDs need to work like model tokens.