LLM-based Conversational Recommendation Agents with Collaborative Verbalized Experience

23 Aug 2025 · 17 min · 9 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How LLM-based conversational recommendation agents can learn user preferences from dialogue history, reflect on past interactions, and improve diversity via multi-agent debate.

Guests

No guests are mentioned; it’s a host-led “Deep Dive” discussion.

Guest backgrounds

N/A.

Key claims

A zero-shot LLM can beat many traditional recommender systems, but conversational recommenders struggle with implicit feedback and a semantic gap between user intent and literal wording. Crave (Yao Chen Zhu et al., University of Virginia + Netflix) uses verbalized experience banks plus a debater-critic agent system to outperform baselines on Redial and Reddit V2.

Notable examples

“cozy movie for a rainy evening” needing divergent genre/mood exploration; lighthearted preferences inferred from prior enjoyment (e.g., The Princess Bride). Crave’s retriever is fine-tuned with item content parameterized multinomial likelihood; removing verbalized experience collection or collaborative retrieval reduces performance.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding AI Recommendations

0:45 to 1:30

Exploration of how AI learns to improve recommendations by understanding user preferences.

“It's about truly understanding your unique preferences, learning from past interactions, and evolving over time to become genuinely helpful.”

Limitations of Traditional LLMs

1:30 to 3:30

Discussion on the challenges faced by LLMs in providing effective recommendations.

“It's a fascinating step towards more, I guess, human-like intelligence in AI.”

Introducing Crave: A New Approach

3:30 to 6:02

Detailed explanation of Crave's innovative method using verbalized experience banks.

“But not necessarily the feeling or the underlying intent behind them.”

Building Experience Banks

6:02 to 7:57

How Crave creates experience banks through trajectory sampling and reflection.

“It's about capturing the essence of what worked, not drowning in detail.”

The Debater-Critic Agent System

7:57 to 10:19

Description of DCA system that promotes divergent thinking through collaboration.

“But what it means, essentially, is that it's trained to understand not just the semantic similarity between queries like how similar the words are, but also the similarity of the actual items involved.”

Performance of Crave Compared to Baselines

10:19 to 12:14

Insights on how Crave outperforms traditional models and establishes a new standard.

“You might naturally assume, well, a team of AI agents debating would always be better than just one, right?”

Key Findings and Implications

12:14 to 14:01

Discussion of critical insights regarding experience retrieval and recommendation diversity.

“It really sets a new bar for conversational recommendation.”

The Role of Item Content in AI Recommendations

14:01 to 16:27

Learn how incorporating item content improves AI recommendation systems significantly.

“And that very technical phrase you mentioned earlier, the item content parameterized multinomial likelihood for tuning the retriever.”

Future of AI Learning Paradigms

16:27 to 16:41

Discover the potential of AI agents learning through reflection and debate.

“And it raises an important question for all of us looking forward.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Hey there, and welcome to the Deep Dive. You know that feeling when your favorite streaming service or online store just mails a recommendation? Yeah, it's great when that happens. It's like magic, right? Right. But then, just as often, maybe it suggests something wildly off the mark. Oh, definitely. Leaves you wondering if it knows you at all. Today, we're diving deep into how AI is learning to get, well, smarter at recommending things to you. Mm-hmm. Especially those back-and-forth conversations we have with our, you know, digital assistants. It's a really critical area, absolutely. As conversational AI becomes truly ubiquitous, think smart speakers, chatbots, assistants everywhere, the challenge isn't just generating any recommendation.

0:45Right. It's about truly understanding your unique preferences, learning from past interactions, and evolving over time to become genuinely helpful. Exactly. And to unpack this, we're exploring some brilliant new research. which is from Yao Chen Zhu and colleagues from the University of Virginia and also Netflix. Ah, interesting mix. Yeah. They've developed something called Cravy. That's conversational recommendation agents with verbalized experience. Bit of a mouthful, but crave. Crave. Got it. So this deep dive is really all about how we can teach large language models, LLMs, to do more than just chat eloquently.

1:20More than just chat. Yeah, to actually reflect on their own past interactions, learn from them, maybe even collaborate with themselves to give you truly personalized recommendations. Collaborate with themselves. Okay, that sounds cool. It's a fascinating step towards more, I guess, human-like intelligence in AI. So our mission today is to explore the ingenious ways Crave tackles some pretty sticky problems in AI recommendations. We'll reveal some truly surprising findings from their work and show you why this is a significant step forward, making our digital assistants feel genuinely intelligent and really attuned to your needs.

1:57Sounds good. Okay, let's unpack this. So we all know LLMs like, you know, the one powering chat GPT. They're incredibly powerful. Undeniably. Amazing at generating answers, interacting in natural language. And the research paper even points out something interesting. What's that? That a basic LLM, a zero-shot one, right, without any special training for recommendations, can often perform better than many traditional recommender systems. That is interesting. It really shows their baseline capability. But what's fascinating here is that despite that vast knowledge, LLMs face specific hurdles when it comes to these conversational recommender systems or CRSs.

2:37CRSs, right. These are the systems aiming to suggest items based on a dialogue, letting you express preferences naturally, not just, you know, clicking buttons. And what exactly are those hurdles? What stops them from being perfect recommenders out of the box? Well, primarily it's about effectively using historical conversations and understanding the implicit feedback you give. Implicit feedback, like what you don't say or how you react. Exactly. Or just subtle cues in the conversation. LLMs are often what we call black box models. They're incredibly complex. That makes them hard to fine tune with past dialogues in a really precise way.

3:12Simple approaches, like just giving them a few conversation examples, a few shot prompt. They often fall short. Why? Why don't those examples work well? Because there's often a significant semantic gap. A semantic gap. Okay, break that down. So the AI might understand the words you're using. Right. But not necessarily the feeling or the underlying intent behind them. Precisely. Like saying, I want something light and it thinks weight, not movie mood. Exactly that kind of thing. It struggles to bridge the literal meaning to your deeper preference. And this raises another important question. How do you get an LLM to think beyond just one single straightforward path?

3:52Because another challenge highlighted in the research is that single LLMs tend to think convergently. Convergent thinking. Okay. That sounds like they get stuck, like on just one idea or one type of solution. Exactly. They follow a logical path, but maybe miss other good options. For tasks like conversational recommendations, where your queries can be quite vague, you might say, uh, recommend me a cozy movie for a rainy evening. Yeah, I do that. You actually need divergent thinking. The system needs to explore various possibilities, different genres, moods, themes. Because if you knew exactly what you wanted.

4:27You'd just search for the movie title. Right. The system needs to be creative, not just logical, to truly surprise and, well, delight you. OK, so if single LLMs struggle with these things, especially that convergent thinking problem, what's Crave's innovative solution? How does it push for more personalization, more creativity? This is where it gets really interesting, you said. Right. Crave introduces a novel approach. It creates what they call verbalized experience banks for the LLM agents. Verbalized experience banks. Yeah, think of it like teaching an AI to keep a detailed, reflective journal, like of its own successes and failures in past conversations.

5:04So it's learning from experience. It's essentially teaching the AI to reflect on its past actions and user feedback, much like, you know, a human expert learns from their own experiences. OK, so it's definitely not just storing old chat logs then. It's much deeper than that. Much deeper. What exactly goes into these experience banks? How does it work? Well, it involves three key steps. First, trajectory sampling. Here, the system samples its own actions on historical conversations. Now, this is clever. Unlike some AI tasks where the right answer is crystal clear, in conversational recommendations, it's tough to pinpoint every single item a user truly liked from a whole chat.

5:45It's messy. Yeah, people mention lots of things. Exactly. So what the researchers found was ingenious. Simply taking one key learning or trajectory sample from each successful conversation was enough. Just one? Just one key takeaway. It was enough to build a remarkably effective experience bank. It's about capturing the essence of what worked, not drowning in detail. Okay, that makes sense. Capture the essence of success. Then what? Then comes verbalized experience collection. Here, the AI actively reflects on these sampled actions against actual user feedback. the ground truth items, the things the user genuinely liked.

6:21And then, crucially, it summarizes these useful experiences in plain natural language. So it's synthesizing the data. Ah, so it writes it down for itself, basically. In a way, yes. Instead of just storing raw data points, it's creating actionable insights. Like, maybe it writes, when a user asks for lighthearted, they often prefer comedies with romantic elements if they previously enjoyed, say, The Princess Bride. That's a great analogy you used before, like a sports coach reviewing game footage. Right. Identifying key plays and then creating specific actionable advice for the future. Exactly that.

6:56And the third step. Okay, what's third? The third crucial step is collaborative experience retrieval. So for a new query you make, a specialized collaborative retriever network looks through these rich experience banks. It finds the most relevant preference-oriented experiences. Preference-oriented, not just keyword matching. Not just keywords. It's trying to match your underlying taste. Okay. Preference-oriented sounds incredibly important here. Yeah. How does it actually ensure it's not just retrieving similar words, but similar tastes? Because like we said, cozy movie could mean very different things to different people.

7:30This is absolutely key. The retriever itself, it's built on sentence BERT. Okay. Sentence BERT. Heard of that. Powerful stuff for understanding sentences. Very powerful. It creates these meaningful numerical representations, embeddings of sentences. But here's the trick. It's specifically fine-tuned using something called item content parameterized multinomial likelihood. Whoa, okay. That's a lot of syllables. Huh, it is. But what it means, essentially, is that it's trained to understand not just the semantic similarity between queries like how similar the words are, but also the similarity of the actual items involved.

8:08It looks at the natural language description of all the items in the catalog. Like plot summaries, genres, actors from movies. Exactly. Plot summary, genre, cast, themes, everything. And it uses statistical likelihood to predict what you'll actually like based on past similar preferences, linking queries to item characteristics. So it's deeply connecting the words you use to the actual stuff being recommended. Precisely. It bridges that gap. Okay. So we've covered how Crave helps LLMs learn from their own past interactions, building these experienced banks. But let's circle back to that nagging convergent thinking problem.

8:43Right. Getting stuck on one track. Yeah. When we're looking for recommendations, we don't want the AI to get stuck on one narrow path. We want a breadth of good creative ideas. So what's Cray's answer to fostering truly divergent thinking? This is where Cray introduces the debater-critic agent system, or DCA. Debater-critic, okay. Instead of a single LLM trying to figure everything out on its own, you have a team of AI agents. You have debaters who evaluate each other's reasoning and recommendations. Oh, you argue. In a structured way, yes. And then you have a critic who judges the debate and provides the final ranked list.

9:18Like a mini-committee making decisions, but with AI agents. That sounds surprisingly human. It does, doesn't it? But with a powerful twist. Each debater maintains its own independent experience bank from TTEF. Ah, so they draw on their own learning. Exactly. It guides their arguments. So their debate is informed by their own learned reflections, their individual sort of expertise. The debaters are explicitly instructed to find issues in previous recommendations and offer new perspectives, new reasoning. Trying to poke holes and suggest alternatives. Right. Then the critic comprehensively evaluates all the contributions, gives numerical scores to judge the quality of the recommended items.

9:59This whole open-ended debate structure is specifically designed to promote that divergent thinking we talked about. So they're literally having an internal debate powered by their individual learned experiences to brainstorm the best possible recommendations. That's brilliant. It's like the AI is performing its own kind of internal peer review. It really is. And here's where it gets truly surprising. You might naturally assume, well, a team of AI agents debating would always be better than just one, right? Yeah, seems logical. More minds, better ideas. Well, the research showed something fascinating.

10:31A zero-shot DCA, without Kray's verbalized experience, just debating in real time. Right. It actually doesn't outperform a single chain of thought agent. Really? What's chain of thought again? That's an LLM prompted to sort of think step by step towards an answer. It's often considered a strong baseline for complex reasoning. So the basic debating team wasn't better than one smart agent thinking carefully. Correct. It's like having a brilliant committee, but maybe without any shared institutional knowledge or history to ground their debate. Ah, here it comes. When you arm that debating committee with Crave's structured, reflected learnings, those experience banks.

11:11Yeah. That's when the magic happens. The DCA system, augmented by Crave, shows a dramatic leap forward, significantly outperforming the single agent. It really highlights how crucial that structured, reflected learning is for the multi-agent system to actually leverage its debate structure effectively. The experience gives the debate substance. So what does this all mean for us, the users, and, you know, for the future of AI recommendations? The paper shows Crave's performance is impressive, but what are the big takeaways? Indeed, the performance is strong. Crave consistently outperforms various state-of-the-art baselines.

11:49This was tested across two real-world data sets, Redial and Reddit V2. And what kind of baselines did it be? All sorts. Traditional models, fine-tuned language models, other LLM-based methods like RAG retrieval augmented generation. Right, where it pulls in external info. Yeah, like raw movie content. and even systems using global rules summarized from training data. Interestingly, the zero-shot LLM itself was already stronger than many of those older baselines. We mentioned that earlier. But Crave pushes it even further. It really sets a new bar for conversational recommendation. So it's not just a tiny improvement.

12:22It's a significant leap beyond just applying generic rules or pulling up raw information from a database. This really leans into genuine personalization. Absolutely. And the study specifically notes that just retrieving raw content like enroll or applying those global rules can actually hurt performance sometimes. Hurt it. Why? Because conversational recommendation is so highly query dependent, it requires deeply personalized insights. Generic rules just don't capture your individual preferences, right? Makes sense. And raw content, like a full movie plot, can be overwhelming or just irrelevant noise if the AI doesn't have the right interpretive lens, the right context from experience.

13:04That makes perfect sense. What else stood out to you in their findings? Any specific details that were particularly insightful or surprising? Well, one fascinating detail is the importance of the number of retrieved experiences, what they call K in the paper. K? Okay, the number of memories it pulls up? Exactly. It's a delicate balance. Too few. And the AI lacks sufficient guidance. It's like trying to make a big decision with only one data point. Not enough context. Right. But too many and less relevant information can risk biasing the recommendations. It can overwhelm the model with noise. Interesting.

13:37So more isn't always better. Definitely not. It shows the incredible precision system designers need to optimize these learning banks for peak performance. Also, their ablation study. Where they take pieces out to see what happens. Exactly. That confirmed that both parts, the verbalized experience collection and the collaborative retrieval module, are essential for Crave to work well. Both are crucial. Take either one away, and performance drops significantly. They work together. And that very technical phrase you mentioned earlier, the item content parameterized multinomial likelihood for tuning the retriever.

14:12How critical was that specific piece? Extremely critical. The research found that if you only use semantic overlap just comparing the words or any non-content-based metric without considering the actual item content during that fine-tuning, performance actually dropped. It was worse than no fine-tuning at all. Wow. So understanding the items is non-negotiable. Absolutely. It powerfully highlights the indispensable role of incorporating both collaborative learning and deep item content information. It's not enough for the AI to just understand your words. It has to understand the characteristics of the items themselves to truly capture your preferences accurately.

14:52And for us, the users, a big win here is also in recommendation diversity, isn't it? Getting more interesting suggestions. Precisely. Not just variations of the last thing you watched, maybe pushing you slightly out of your comfort zone, but in a good way. Exactly. The DCA system, powered by Crave, doesn't just improve recommendation accuracy how often it gets it right. It also significantly enhances the diversity of the recommended items. That's great. Which means you're more likely to discover new things that truly align with your nuanced preferences, rather than just getting slight variations of what you've already seen.

15:24It broadens your horizons while staying relevant. This deep dive into Crave, it really shows how thinking differently about AI can unlock incredible potential. We're moving beyond just simple keyword matching or database lookups. Way beyond. To systems that truly learn and reflect on preferences, much like a human expert would, you know, evolving your understanding over time. It's a significant leap. It's about leveraging the implicit knowledge, all that stuff hidden within our interactions. It provides a powerful pathway for LLMs to become not just conversational, but genuinely insightful and personalized recommender agents that adapt to your unique tastes.

16:03Think about it. An AI that actively debates with itself, drawing on nuanced past experiences and detailed item knowledge to give you the perfect movie or book or even that perfect obscure product suggestion you didn't know you needed. This really makes our digital world feel a lot more personal, doesn't it? A lot less like just interacting with a giant impersonal database. It certainly points in that direction. And it raises an important question for all of us looking forward. If AI agents can learn through reflection and collaborative debate for recommendations, where else could this paradigm go?

16:36Right. What other complex AI tasks could benefit? Exactly. Tasks where clear-cut answers are rare, where nuanced understanding is paramount, and where avoiding that narrow, convergent thinking is crucial for innovation and discovery. What stands out to you from today's discussion? We hope this deep dive gave you some new insights into the cutting edge of AI and how it's constantly learning to serve you better, not just with more accuracy, but maybe with a bit more wisdom. Thanks for joining us on the deep dive.

From the publisher

This paper introduces CRAVE (Conversational Recommendation Agents with Collaborative Verbalized Experience), a novel framework designed to enhance Large Language Model (LLM)-based conversational recommender systems (CRSs). The core idea is to improve recommendation accuracy by leveraging implicit, personalized, and agent-specific experiences derived from historical user interactions. CRAVE achieves this by sampling trajectories of LLM agents on past queries and creating "verbalized experience banks" based on user feedback. A collaborative retriever network then helps identify relevant, preference-oriented experiences for new queries, further augmented by a debater-critic agent (DCA) system that encourages diverse recommendations through a structured debate. The research demonstrates that this approach significantly outperforms existing zero-shot LLM methods and other baselines, particularly when augmented with collaborative verbalized experience.

More from Best AI papers explained

All 475 episodes
LLM-based Conversational Recommendation Agents with Collaborative Verbalized ExperienceBest AI papers explained · 17 min
Listen in VO