Generative Recommendation with Semantic IDs: A Practitioner’s Handbook

4 Aug 2025 · 17 min · 8 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Generative recommendation using Semantic IDs (SIDs), focusing on Snap Inc.’s open-source GRD framework. It explains a two-phase pipeline: semantic ID tokenization (encode item features into dense representations, then cluster/tokenize into discrete SID sequences) and next-item generation (train transformer sequential models to autoregressively predict future SIDs from user history; decode with beam search).

Key claims

simpler SID tokenizers (RKMeans/RVQ) can outperform or match complex RQVAE while needing ~5x fewer training iterations; scaling semantic encoders (e.g., Flan T5 XXL) yields only marginal gains; adding SID residual layers can hurt beyond an optimum (~3 layers, 256 tokens/layer); removing “user tokens” can improve performance; encoder-decoder models outperform decoder-only; sliding-window augmentation is essential; unconstrained beam search is similarly effective but more efficient than constrained.

Notable examples

Amazon Reviews (beauty/sports/toys) with item features (title, categories, description, price).

Guests

No guest names or backgrounds are provided in the transcript.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Introduction to Semantic IDs

0:46 to 2:42

Understanding semantic IDs and their role in generative recommendations.

“Okay, so at the core of this new wave, we have something called semantic IDs or SIDs.”

Phase One: Semantic ID Tokenization

2:43 to 4:00

Explaining the process of semantic ID tokenization and its importance.

“And if we sort of connect this to the bigger picture, GRAD simplifies generative recommendation with SIDs into two main modular phases.”

Phase Two: Next Item Generation

4:01 to 7:06

The process of generating recommendations based on user behavior.

“And what's really clever here, from what I understand, is GRID's flexibility.”

Insights from GRID's Experiments

7:07 to 7:58

Exploring key insights and surprising findings from the experiments.

“Complexity often implies better performance, or so we think.”

Counterintuitive Findings in Recommendations

7:59 to 11:15

Discussion of unexpected results regarding complexity and user tokens.

“You'd instinctively think bigger is better, right?”

Practical Decoding Choices

11:16 to 14:00

Practical insights on deduplication and beam search for deployment.

“Okay, speaking of architecture, you mentioned a decoder versus decoder only.”

Comparing Beam Search Strategies

14:00 to 15:17

Learn about the differences between constrained and unconstrained beam search in generative recommendation.

“Okay, and the second thing, constrained versus unconstrained beam search.”

Insights on Simplifying AI Solutions

15:17 to 16:38

Discover how simpler solutions can often yield comparable or better results in AI systems.

“We've seen that generative recommendation with semantic IDs is this powerful new paradigm.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to the Deep Dive, where we cut through the noise and get straight to the insights. If you're like most of us, your digital life is shaped daily by recommendation systems, whether you're buying something online, choosing a movie or connecting on social media. They are truly ubiquitous. Absolutely. They're everywhere. And what's particularly exciting, I think, is that these systems are currently undergoing this fascinating evolution. Evolution how? Well, we're witnessing a shift towards what's called a generative paradigm. Yeah. It's very much inspired by the breakthroughs we've seen in AI that generate text or images.

0:37Right, like GPT or stable diffusion, but for recommendations? Exactly. It's fundamentally changing how these systems think about and deliver recommendations. Okay, so at the core of this new wave, we have something called semantic IDs or SIDs. Can you maybe break down for us how these SIDs change what a recommender system understands? They sound, I don't know, like a kind of secret code. That's actually a good way to put it. Yeah. Think of SIDs as a special kind of numerical code or maybe like a structured language. They take incredibly complex semantic representations. That's the rich, nuanced understanding an AI model gets from an item's image or its description, its features.

1:16And they translate it. Translate it into what? Into discrete, learnable sequences of numbers. And what's crucial here, the really key part, is that this allows the system to grasp the meaning of an item. while also leveraging your historical interactions. It's a powerful fusion. That fusion sounds incredibly potent. It really does. But, you know, as with any cutting-edge tech, getting it into the hands of researchers and practitioners, that can be a challenge. Oh, definitely. We've heard there's been a real lack of open-source tools for this kind of generative recommendation using SIDs. How big of a hurdle has that actually been?

1:52It's been a significant bottleneck. Yeah. A real problem. without unified open source tools, comparing models, debugging, iterating quickly. Yeah. It all becomes incredibly difficult. Right. It's like trying to build, I don't know, a complex machine without a shared blueprint or a common set of tools. Everyone's kind of working their own silo. Yeah, I can see that. So to tackle this head on, researchers at Snap Inc. developed an open source framework. It's called GRD. That's Generative Recommendation with Semantic IDs. That's the one. So for us today, our mission is to deep dive into the paper introducing GRD.

2:25We'll unpack how generative recommendation with SIDs actually works. And explore the, frankly, surprisingly impactful design choices revealed by GRD's experiments. Some real eye-openers in there. Exactly. And understand what this all means for the future of recommender systems. So get ready for some aha moments that might challenge what you thought you knew about building these advanced AI systems. Yeah, definitely. And if we sort of connect this to the bigger picture, GRAD simplifies generative recommendation with SIDs into two main modular phases. Okay, two phases. What are they? First, tokenization, then generation.

3:01It's a very clear two-step process that brings some much-needed structure to this complex task. Let's start with phase one then, semantic ID tokenization. What's the ultimate goal here? And how does Guiardi approach transforming an item's rich features, like its title, categories, description, price, into these simplified numerical semantic IDs? Right. So the goal is really to distill all those complex features down into a structured numerical format the machine can understand. How does it do that distillation? It works like this. First, you have a modality encoder. You can imagine a large language model here, something like P5 or QN or BERT.

3:41that encoder processes the item's features to create what's called a d-dimensional representation. Basically, a dense numerical fingerprint of that item's meaning. A fingerprint. Okay, I like that analogy. And then a hierarchical clustering tokenizer takes this fingerprint and converts it into a sequence of these sparse SIDs. So yeah, it's like giving a unique, meaningful numerical code to each item. And what's really clever here, from what I understand, is GRID's flexibility. It offers these plug-and-play modules, right? Exactly. That's a key part. You can easily swap out different encoders, maybe use existing models from Hugging Phase, which is super convenient.

4:15Yeah. And you can also experiment with different tokenizer algorithms, things like residual mini-batch k-means or residual vector quantization, advanced clustering techniques. So you're not locked into one specific way of doing it. Precisely. This phase is incredibly important because, like we said, it's about translating that continuous, rich information like a detailed product description into a structured, discrete language. like distilling a whole book down to a sequence of keywords you said. Exactly like that. Keywords that are memorable but still meaningful, capturing the essence. That makes perfect sense.

4:49Okay, so once we have these SIDs, these codes, what's the next step? How does the system use them to actually predict what you might want to see next? That's phase two, right? Next item generation. Exactly, phase two. So in this phase, a sequential model, often it's transformer-based, you know, similar architecture to what powers large language models. It learns from sequences of SIDs. The sequences representing a user's past behavior. Right. Their history. And then it auto-regressively predicts the SIDs of future items the user's likely to engage with. And GRED supports different architectures here too.

5:25It does. Both encoder-decoder and decoder-only model architectures. It gives researchers a lot of flexibility in how they configure these models, number of layers, attention heads, all that. Okay. And for training, they use common techniques like next token prediction objective. Yeah, that's standard. Basically teaching the model to predict the next most likely SID in a user sequence, plus things like sliding window augmentation to get more training data. And then for actually generating the recommendations efficiently, they use BeamSearch. Correct. BeamSearch helps the model sift through possibilities and generate the best potential item recommendations.

6:01So this is where the magic of generation truly happens. Using the structured SIDs to forecast future preferences, kind of like an AI predicting the next word in a sentence. That's a great analogy. And look, this is where the GRID framework really shines beyond just providing the tools. The researchers didn't just build it, they systematically tested it. Right, the experiments. What kind of data did they use? They used data sets like Amazon Reviews for beauty, sports, and toys. Yeah. And the item features included things like title, categories, description, and price. So real-world type data. And these experiments yielded some genuinely surprising aha moments, things that maybe defy common assumptions.

6:45Oh, absolutely. Some really counterintuitive stuff came out. Okay, now this is where things get really interesting. Let's dive into those aha moments. What was the first big surprise that stood out? Well, the first big one was about the SID tokenizer algorithm itself. You know, the common assumption, which you often see in the literature, is that a more complex algorithm like RQVAE, that's Residual Quantized Variational Autoencoder, is sort of the default, often seen as the superior choice. Right. Complexity often implies better performance, or so we think. Exactly. But the revelation from GRID's experiments was pretty astounding.

7:17Simpler algorithms like RKMeans and RVQ, they often led to better or comparable recommendation performance than the complex RQVAE. Better or comparable? With simpler methods. Yes. And get this. The more complex RQVAE actually required five times more training iterations to get those worse or comparable results. Five times more training time for worse results. Wow. OK, so for anyone balancing model complexity with real world efficiency, this finding is gold. Absolutely. Sometimes the simpler path leads to better results and it's five times faster to train. That's a huge optimization. It really challenges that notion that complexity always equals superiority, doesn't it?

7:58It certainly does. Okay, that's a massive takeaway. What about the next insight? Semantic encoder size. You'd instinctively think bigger is better, right? Larger language models like Flan T5 XXL with its 11 billion parameters. Yeah, the assumption is they have all this vast world knowledge embedded. So they should dramatically improve performance. What did GRAD find? Well, yet again, a surprise. When they varied the Flan T5 model size, going from large, which is about 780 million parameters, up to XXL at 11 billion, that's a 14-fold increase. 14 times bigger. They found only marginal increases in recommendation performance.

8:36Marginal after a 14x jump in size. That's a fascinating paradox. Why do you think that is? Why isn't all that extra knowledge translating into way better recommendations? That's the million-dollar question, really. It suggests that the current generative recommendation with SID pipelines, well, maybe they're just not fully leveraging the immense knowledge packed into these giant LLMs yet. So it raises a really important question for future research. Right. How can we better harness that power? Or maybe it's not always needed. Exactly. Another interesting insight came from the FID tokenizer dimension.

9:08Remember, SIDs are structured with residual layers and tokens per layer. Intuitively, you'd think adding more layers means encoding more semantic information, more detail. Right. More layers, more info, better results. You'd think so. But what happened when they actually added more layers? Let me guess. It wasn't straightforward. Not at all. Surprisingly, when they experimented with the RKMeans SIDs, performance actually dropped substantially when they used more layers than the optimal configuration. It dropped. Yeah. The sweet spot seemed to be around three layers with 256 tokens per layer. Go beyond that, add more layers, and performance tanks.

9:45So there's a crucial tradeoff. There's a point where adding more semantic information actually hurts the model's ability to learn the SID sequence. Precisely. It's like trying to drink from a fire hose, maybe. Yeah. Too much information becomes overwhelming for the model to process effectively. Okay, that makes sense. Then there's this really surprising finding about user tokens. Now, some GR models, like one called TIGER, add these user tokens into a user's SID sequence. The idea is to personalize recommendations better, right? That's the intention, yeah. More user tokens should mean more personalization.

10:18So the more the better. Well, counterintuitively, QRID's experiments showed that increasing the vocabulary size of these user tokens did not always improve performance. Didn't always improve it. In fact, completely removing user tokens, setting the quantity to zero, often led to optimal performance. Optimal performance with no user tokens, that's a shocking reversal. It's completely counter to the intuition. Isn't it? It's like trying to explicitly tell a friend exactly what you like for dinner, when maybe they already know you so well they can just infer it from your past choices. That's a great analogy.

10:53It implies that this current approach, using user tokens for personalization, it might not be achieving its goal effectively. It could even be detrimental, adding noise. So perhaps the collaborative signals within the item SIDs themselves are already capturing enough personalization implicitly, making explicit user tokens redundant. It's a strong possibility. Maybe they're just not needed, or at least not in the way they're currently used. Fascinating. Okay, speaking of architecture, you mentioned a decoder versus decoder only. Was there a clear winner there? Oh, yes. A very clear winner emerged.

11:26Encoder-decoder models significantly outperformed the decoder-only models across all the datasets they tested. Significantly better. But decoder-only models are getting so much hype in other AI areas, like text generation. Does this mean they're just inherently less suited for recommendation? Or maybe there's a way for them to catch up eventually? Well, the paper hypothesizes that the encoder's dense attention mechanism, its ability to look closely at the entire user history, is really crucial here. Okay. This mechanism lets it capture richer, more comprehensive sequential patterns. and that deep understanding of the past is vital for the decoder part to then generate effective recommendations for the future.

12:07So if your goal is truly personalized recommendations based on a user's whole journey, the encoder-decoder structure isn't just a choice. It might be a critical advantage right now. It really listens to the full history. Seems that way, at least for these types of tasks and models. Okay, this next one feels foundational across AI, but it's always worth hammering home. Data augmentation. Absolutely paramount, you said. They used a technique called sliding window. What did that reveal? Yes, sliding window basically expands a single user sequence into lots of overlapping subsequences for training.

12:40And what they found was, well, proper data augmentation was just paramount, essential. Paramount for what? For achieving robust and high-performing generative recommendation models. Without that augmentation step, performance suffered significantly. It just didn't learn as well. It really underscores that fundamental principle, doesn't it? Sometimes how you prepare and expand your training data like cultivating fertile soil for a plant can be just as important, maybe even more important than the fancy model architecture itself. Absolutely. It helps the model learn generalizable patterns and avoid just memorizing the training examples, avoids overfitting.

13:17Crucial stuff. Okay, finally, let's touch on practical decoding choices. Things relevant for actually deploying these models efficiently. They looked at two things, deduplication of SIDs and constrained versus unconstrained beam search. What did those findings suggest for real-world use? Right, so first, deduplication, making sure each recommended item is unique. They compared TIGERS method, which involves appending a digit to the SID, versus a simpler random selection if a duplicate occurs. Both performed pretty comparably. TIGERS had maybe a slight edge, but it comes at a cost. It increases the sequence length, makes decoding more complex, and critically, it requires knowing the global SID distribution.

13:57Which is impractical for really large item catalogs, right? Exactly. Millions or billions of items. That's tough. Okay, and the second thing, constrained versus unconstrained beam search. Yeah, so constrained beam search guides the HODL's output, forcing it to only pick valid existing SIDs. Unconstrained lets it explore any sequence, even potentially invalid ones, hoping the model has learned well enough. What a performance difference. Again, surprisingly, both yielded similar recommendation performance. But, and this is key for deployment, unconstrained beam search was significantly more efficient and computationally cheaper.

14:33Ah, so similar results, but one is much faster and cheaper to run. Precisely. These findings offer really valuable lessons for anyone putting these systems into production. They show that in some cases, you can get comparable performance with much simpler, much more efficient methods. Why do you think that is? Why does unconstrained work so well here? It's likely thanks to the inherent structure of the SID generation task itself, and the fact that the models, when well trained, learn the patterns of valid SIDs pretty effectively on their own. They don't need as much hand-holding during decoding.

15:06It's incredible how often that simpler is better, or at least simpler is just as good and way cheaper, mantra holds true, even in advanced AI. It really does seem to. So let's wrap this up. What does this all mean? We've seen that generative recommendation with semantic IDs is this powerful new paradigm. And GRID is clearly an invaluable open source tool that's accelerating its development. Definitely. And its experiments have revealed these truly surprising, sometimes counterintuitive insights about how best to build these systems. Yeah, we learned that complexity isn't always the answer, right?

15:40That even much larger models might not be fully utilized yet. And that some seemingly minor design choices like how you augment your data or your beam search strategy can have really major impacts. And the value of having an open source framework like GRID is just immense here. It provides that unified, reliable platform for doing systematic experiments and proper benchmarking pushes the whole field forward. It's clear that these overlooked architectural components and sometimes simpler alternatives can be just as impactful, maybe even more so than the most complex resource intensive solutions, which I think raises a fascinating question for you, the listener.

16:19In any complex system you're designing or trying to understand, whether it's a business process, a creative project, maybe even a personal habit, are there simpler alternatives or maybe overlooked components lurking there that could unlock disproportionate benefits if only you had the right framework or maybe just the mindset to systematically test them? Something to think about. Thanks for diving deep with us.

From the publisher

The research paper "Generative Recommendation with Semantic IDs: A Practitioner’s Handbook" introduces **GRID**, an open-source framework designed to standardize and accelerate research in **Generative Recommendation (GR) with Semantic IDs (SIDs)**. GR models leverage advancements in generative AI to recommend items, while SIDs convert continuous semantic representations of items into discrete sequences, allowing these models to incorporate both semantic information and collaborative filtering signals. The authors identify a current lack of unified, open-source tools in this field, making direct comparisons and systematic experimentation challenging. GRID addresses this by offering a modular platform for **tokenization-then-generation architectures**, enabling easy swapping of components like semantic encoders and tokenizers. Through experiments using GRID, the paper provides surprising insights into the **performance impact of various architectural choices**, such as the tokenizer algorithm, the size of the language model encoder, and the use of data augmentation, ultimately validating GRID's utility for robust benchmarking and research advancement.

More from Best AI papers explained

All 475 episodes
Generative Recommendation with Semantic IDs: A Practitioner’s HandbookBest AI papers explained · 17 min
Listen in VO