In short
The episode surveys how AI personalization is added to Retrieval-Augmented Generation (RAG) and how that evolves into personalized agent systems, using the academic paper “Survey of Personalization: From RAG to Agent” by researchers from City University of Hong Kong and Huawei Noah’s Ark Lab.
Guest backgrounds
No guests are named in the transcript; it’s a two-host discussion.
Key claims
RAG reduces hallucinations by grounding answers in retrieved external sources, but personalization is needed for context-aware, user-aligned responses. Personalization data (explicit profiles, behavioral history, and persona-based user simulation) is used across pre-retrieval, retrieval (e.g., user embeddings, personalized knowledge graphs), post-retrieval (re-ranking/summarization), and generation (explicit prompting, summary-augmented prompting, adaptive prompting, implicit fine-tuning like LoRA, and reinforcement learning).
Notable examples
personalized e-commerce query rewriting/expansion; dense vs sparse retrieval for chatbots; health and travel agents using user-specific memory; shopping agents calling APIs to compare prices and apply loyalty points. Challenges: scalability, evaluation metrics, privacy (on-device + cloud), and ethical coherence/bias.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Personalization and Its Importance
0:57 to 3:18
Discover why personalization is essential for improving AI effectiveness and user experience.
“So our mission today really is to unpack what personalization means here.”
How AI Learns About User Preferences
3:18 to 5:10
Explore the methods AI uses to gather and understand user information for personalization.
“ARAG improves the facts, yes, but personalization, that's what makes the AI adaptive, context aware.”
Integrating Personalization into RAG Workflows
5:10 to 6:53
Learn how personalization is incorporated into the Retrieval Augmented Generation process.
“Here, you actually use LLM-based agents to simulate different kinds of users with specific personas or goals.”
Enhancing Retrieval Methods for Personalization
6:53 to 10:40
Understand how different retrieval methods are adapted to provide personalized results.
“First up, pre-retrieval, making the search smarter from the get-go.”
Post-Retrieval Strategies for Tailored Outputs
10:40 to 13:44
Examine the strategies used after retrieval to refine and personalize AI responses.
“We can break them down into a few main types.”
Generating Custom Content Through Personalization
13:44 to 14:00
Discover the techniques used to create personalized content in AI responses.
“Super important for making search, dialogue, everything feel personal.”
Personalization Techniques in AI
14:00 to 17:30
Learn about explicit and implicit methods for personalizing AI interactions.
“It's about generating content that doesn't just answer the query, but aligns with your individual preferences, your style, your needs.”
Personalized Agents: Understanding and Roles
17:30 to 18:50
Explore how personalized agents understand users and their roles in interactions.
“Okay, so we've walked through personalized RAG.”
Active Personalization and Planning
18:50 to 20:01
Discover the process of personalized planning and execution in agents.
“How does the agent's role adapt based on the specific user it's interacting with?”
Challenges in Personalization and Privacy
20:01 to 22:22
Identify the main challenges of scaling personalization while ensuring privacy.
“It could orchestrate APIs to compare prices across sites, check inventory, apply your loyalty points, and even complete the purchase, all based on your preferences.”
Show all 11 chapters
Ethical Considerations in AI Personalization
22:22 to 24:52
Understand the ethical implications of personalizing AI systems and user trust.
“Personalized AI often deals with incredibly sensitive data.”
Transcript
Automatic transcript. May contain errors.0:00Imagine an AI, one that doesn't just spit out information, but actually gets you. You know, your preferences, what you're doing, what you're trying to achieve. Yeah, like an assistant that knows what you need before you ask, recommends things you'd actually like, maybe even talks in a way that just feels right for you. Exactly. Sounds a bit sci-fi, but that's kind of what we're digging into today. It really is. This deep dive is all about personalization in modern AI, specifically how we're tailoring these big, large language models, LLMs, using something called Retrieval Augmented Generation, or RA.
0:37And how that's sort of morphing into these really sophisticated agent systems. That's the one. And we're basing our discussion on a really comprehensive academic survey. Right. It's called the Survey of Personalization, from RAG to Agent. It comes from researchers at City University of Hong Kong and Huawei's Noah's Ark Lab. Seems pretty cutting edge. Definitely. It gives a great view of where things are headed. So our mission today really is to unpack what personalization means here. Yeah, how it gets baked into rags, step by step. And then how these ideas kind of bloom into these more autonomous, personalized AI agents.
1:10We'll also touch on, you know, the hurdles and what's next. So why should you care? Well, it's about getting AI that feels less generic and more like it truly gets you. Makes tech way more intuitive, hopefully. Okay, so let's start at the beginning. Why is personalization suddenly so critical? I mean, LOMs have been amazing, right? But they're not perfect. They have issues. Oh, absolutely. Big issues sometimes, like the information can be outdated or they just mix stuff up, the whole hallucination problem. That really hurts accuracy. Right. And that's where Adlai comes in. Retrieval augmented generation.
1:44How does that help? Well, Arlay is a really promising framework precisely because it tackles those LLM limits. It works by integrating information that's retrieved from outside sources, external up-to-date sources. Okay, so it checks its facts first. Kind of, yeah. Before generating an answer, it looks things up, maybe from an API or a scientific database like Archive or even like specific company databases for e-commerce or healthcare. Ah, so it grounds the LLM's output in reliable current knowledge, like giving it a fact checker. Exactly. Gives it a solid reference library. Interesting. You know, as I was reading about RAG, the structure, the way it works, it felt a lot like how you'd imagine an autonomous agent operates.
2:25Is that connection real? That's a really sharp observation. And yes, it's definitely a convergence we're seeing. It's not just a coincidence. OK. If you break down the R workflow, the parallels are pretty clear. Take the first step in RAG, understanding the query, maybe rewriting it. That's very much like an agent's comprehension phase, right? Right. Understanding the instruction. Okay. Yeah, I see that. And then RAG retrieves documents. Right. Which is like an agent's planning and execution, figuring out what info it needs and then going to get it. And the final generation step in RAG. Like the agent's action execution, delivering the final result or taking the action based on the plan and the retrieved info.
3:01So, yeah, it shows this move towards more intelligent, autonomous systems. Okay. So RAG boosts accuracy, makes things more grounded. But you're saying even that isn't enough for truly advanced AI. Personalization is the key next step. This is where it gets really interesting for the actual user experience for you. ARAG improves the facts, yes, but personalization, that's what makes the AI adaptive, context aware. Without it, the AI might be correct, but still feel generic. Exactly. It might give you a perfectly factual answer that just isn't quite right for your specific situation or need. Personalization is fundamental if you want AI that can reason in a way that aligns with how you think.
3:44Or make decisions that consider your preferences. Right. Or generate content that actually resonates with you personally. Or build interactive systems that feel, you know, genuinely conversational and understanding. It's really essential if we're thinking about moving towards AGI, artificial general intelligence. Truly general intelligence has to understand and adapt to individuals. Okay. So let's define it then. When we say personalization here, what are we actually talking about? How does the AI learn about me? Good question. Formally, it means tailoring the model's predictions or the content it generates so that it aligns with a specific individual's preferences.
4:20And how does it gather that info about me? It happens in a few ways. Think of it like building up layers of understanding about the user. Okay. Layer one is probably the obvious stuff, the explicit user profile. Things I directly tell at my age, location, maybe job title, that kind of thing. Exactly. Biographical details, attributes, even social connections if you provide them. That's the explicit layer. Then there's what it observes, my behavior. Precisely. That's user historical interactions. Your browsing history clicks, purchases all that behavioral data. The AI infers your interests from what you do, even if you never spell it out.
4:56And presumably the stuff I write, my user historical content, like old chats, emails, maybe reviews I've left. Absolutely. The text you generate is a goldmine for personalization. It shows your communication style, your topics of interest, your perspective. It's rich with implicit signals. Wow. OK. And there's one more. Something about simulation. Yeah. This one's a bit more advanced. Persona-based user simulation. Here, you actually use LLM-based agents to simulate different kinds of users with specific personas or goals. Well, how would you do that? It helps the AI essentially practice interacting in a personalized way.
5:31It generates these simulated personalized interactions to refine its own ability to tailor responses later on with real users. That makes sense. Yeah. So the AI gathers all this info, explicit, historical, simulated. The paper calls it P. How does this personalized data actually get used in the ARG process? Right. That P is crucial. It's not just an afterthought. it influences multiple stages of the R pipeline directly. Like where? It gets used right at the start, in the pre-retrieval phase. When the system is processing your initial query, trying to understand it better or refine it, then it plays a big role in the retrieval phase itself, guiding the system to fetch documents that aren't just relevant to the query, but relevant to you.
6:15And finally, I assume, in the generation phase. You got it. It shapes the final output, making sure the response produced is tailored to your preferences, your style, whatever the bit indicates. So it's woven throughout the whole process. And you mentioned earlier that agent systems are kind of like an evolution of this. Personalized RGA plus Karai? Yeah, you can definitely see agents as a specialized application of this whole personalized RGA framework. They integrate personalization in very similar ways, just often taking it to a more advanced, more interactive level. Okay, this is fascinating.
6:48Let's really unpack that how. We know the what and why of personalization. Now, how does it actually get integrated into that RGA workflow? Let's go stage by stage. First up, pre-retrieval, making the search smarter from the get-go. Right. This is that vital first step. Before the system even starts looking for documents, it takes your original query and tries to enhance it, maybe modify it, to improve the relevance and quality of what comes next. It's about ensuring the search starts off on the right foot, sharpening the focus, basically. And a key technique here is query rewriting, changing my original query.
7:21How does personalization fit in? Well, there are a couple of main flavors. First, there's direct personalized query rewriting. Okay. Here, the model directly rewrites your query based on what it knows about you. Think about searching on a big e-commerce site. A system might look at your past browsing and adapt your search query, even if it was a bit vague, to show you stuff you're more likely to buy. It's about customizing the search for your habits. Clever. Makes sense. What's the other type? That's auxiliary personalized query rewriting. This is a bit more involved. It uses extra mechanisms alongside the rewrite, maybe reasoning strategies or accessing external memory.
7:59Like what kind of reasoning? For instance, some systems use something called least to most prompting. If you ask a really complex question, the AI might break it down into smaller steps, figure out each step, then combine the answers. It helps the AI grasp complex intentions more accurately by reasoning through it step by step. That sounds super useful, especially if I'm just talking to an AI, not typing a perfect search term. Yeah. What about query expansion? That's about adding stuff to my query. Exactly. Query expansion adds related terms, synonyms, maybe structures it a bit differently. The goal is to capture your intent more fully, broaden the search scope effectively.
8:38How does personalization help that? Early methods, even back in 2009, used tags. Systems would build these personalized networks based on what you and others bookmarked or tagged. More recently, models look at your search history, maybe your social network connections. Ah, so it uses my past behavior, maybe what people like me search for, to suggest related terms or concepts. Precisely. It combines, say, your local behavior with social strategies from the network. This allows for really dynamic, real-time expansion of your query, making the search much richer. Okay, so let me see if I've got this.
9:12For me, as a user, query rewriting is probably best when my original query is a bit fuzzy, maybe in a conversation. Whereas query expansion is more for when my query is okay, but maybe too narrow, and needs more semantic breadth to find the really good stuff. You've nailed it. That's a great way to think about the distinction. Cool. So once the query is sharpened up, personalized, we hit stage two, retrieval. Finding what I need in that sea of data. Exactly. Now we're talking about sifting through potentially massive amounts of information, the corpus, to pull out the documents that are most relevant, but crucially relevant specifically to you.
9:52That sounds like a huge organizational challenge. How do they index all that data so it can be retrieved in a personalized way? It is a challenge. Indexing is fundamental how you structure the knowledge base for efficient lookup. And personalization can be built right into this indexing stage. Well, some systems generate a unique user embedding, like a numerical fingerprint of your preferences, and encode that directly into the index using your personal history data. Others integrate knowledge graphs. Think of these as networks of interconnected facts. These graphs can be personalized, maybe even made editable.
10:25So the system has this dynamic map of information that reflects your profile and updates over time. That's incredibly sophisticated, like a living knowledge base tailored to me. Okay, so the data is indexed smartly. how does the system actually match my personalized query to this knowledge? What are the retrieval methods? We can break them down into a few main types. First, there's dense retrieval. Dense. Yeah, it uses vector embeddings. Basically, it turns your query and all the documents into numerical representations vectors in a high-dimensional space. Then it finds the documents whose vectors are closest to your query's vector using similarity metrics.
11:03So it's looking for conceptual similarity, not just keywords. Exactly. it understands the meaning. Systems use this for things like personalized text generation on your phone, making dialogue systems better, personalizing product search. It's about that deeper semantic match. Okay. What about sparse retrieval? Sounds like the opposite. It kind of is. Sparse retrieval is more traditional, relying on term-based matching. Think keyword matching, like the classic BM25 algorithm. It looks for exact or similar words. So one's meaning, one's words. Roughly, yeah. And interestingly, some of the most advanced systems actually combine sparse and dense retrieval.
11:39They get the best of both worlds. For instance, a personalized chatbot might use your exact words, sparse, and the overall conversation context, dense, maybe even images you shared, to give a really nuanced, personalized reply. Makes sense. And you mentioned prompts earlier. Can prompts guide retrieval too? They can. That's prompt-based retrieval. Here, specific prompts, potentially containing information about your preferences or past interactions, are used to guide the search process. Like reminding the AI what I liked last time. Sort of. Imagine a multi-session chatbot. It could store your preferences and the dialogue history within the prompts it uses for retrieval, helping it maintain context and personalization over time.
12:18This is great for recommendations or ongoing search tasks. Any other method. Yeah. A few others are popping up. Some use reinforcement, learning the system learns the best retrieval strategy based on user feedback. Others use parameter-based retrieval, where user-specific info is sort of implicitly baked into the model's parameters. Lots of innovation happening there. Okay, wow. So the system retrieves a bunch of potentially relevant stuff using one or more of these methods. But it's not done yet, right? There's a post-retrieval step. Correct. You've retrieved candidates, now you need to polish the results.
12:51A key part is re-ranking. Ordering them by relevance to me. Exactly. Putting the best stuff at the top. Some systems use specialized agents just for this re-ranking task, focusing on user-centric relevance. Others use clever scoring mechanisms during the process. What else happens post-retrieval? Summarization is another one, condensing the key info from retrieved documents. We're seeing cool applications, like using a role-playing agent to summarize your interaction history to give you a personalized opinion digest. And then there's compression, making the retrieved content smaller, more efficient to process.
13:26The survey notes this is still kind of underexplored for a personalized system, so maybe an area for future breakthroughs. Okay, so the big takeaway for this retrieval stage is it's this balancing act between finding stuff efficiently and understanding it deeply, all through the lens of the individual user. Super important for making search, dialogue, everything feel personal. Perfectly put. And that naturally leads us to the grand finale of the RH pipeline. Generation, stage three, creating your custom content. Right. This is where the magic happens, presumably. Taking all that understanding, all that retrieved info, and actually producing the tailored response.
14:03Exactly. It's about generating content that doesn't just answer the query, but aligns with your individual preferences, your style, your needs. How does it actually weave my preferences in? Does it just stick my profile info into the prompt? That's definitely one major approach. Using explicit preferences. This involves directly incorporating user input demographics, past behaviors, maybe examples of your writing style, into the process as the LLM generates the text. Okay, how does that look in practice? One common way is direct integrated prompting. Literally putting user preferences right into the prompt.
14:36If you're interacting with, say, a role-playing AI character, the prompt might include details about your profile so the character can tailor its responses to you. It helps the AI predict intent and create that personalized feel. But what if my history is huge and messy? Just dumping it all in the prompt sounds inefficient. Good point. And that's why we have summary augmented prompting. Instead of feeding the LLM a long, possibly noisy interaction history, the system first summarizes it. It extracts the key preferences or habits. Ah, so it gives the LLM a concise summary of me. Pretty much. Like in recommendation systems, an LLM might generate a summary of your tastes based on past behavior, and that summary then helps it generate better, more personalized recommendations, much cleaner.
15:21And I guess writing these personalized prompts manually would be a nightmare. Are there automated ways? Absolutely. That's adaptive prompting. Using automated methods to generate the personalized prompts themselves. Techniques like prompt tuning can fine-tune prompts specifically for personalized recommendations, or systems might look at graphs of your interactions to generate prompts reflecting those connections. Okay, so that's all about explicit preferences, stuff the AI has directly told or shown about me. What about the other side? Implicit preferences. Sounds more subtle, baked in somehow.
15:57It is more subtle. With implicit preferences, the personalization isn't fed in through the prompt. It's embedded within the generator model's own parameters, usually during a training or fine-tuning phase. How do they do that? Fine-tuning the whole model sounds expensive. Often they use more efficient methods. Fine-tuning-based methods are common, especially techniques like PEFT parameter efficient fine-tuning like LoRa. These let you adapt large models without retraining everything. So you can add user-specific knowledge efficiently. Right. You might combine general task fine-tuning with user-specific LoRa layers, or even use something simple like your user ID as a factor during fine-tuning.
16:33People are even exploring fine-tuning smaller personalized LLMs right on your device to keep data private. Very cool. And reinforcement learning, can that be used here too? Definitely. Reinforcement learning-based methods are another powerful approach. The idea is to align the generated text with your preferences by optimizing the model based on feedback, either explicit feedback you give or inferred feedback. How does that work? Well, a system might learn a reward model specific to you, guiding the generation towards text that matches your style or preferences. Or it might apply personalized rewards at a really granular level, like for each word generated, to steer the output very precisely.
17:13Okay, so let me summarize the generation part. Explicit preferences give you transparency. You can see the input driving the personalization. Implicit methods allow for maybe deeper, more nuanced personalization, but are heavier computationally. The best choice depends on the specific application. That's a perfect summary. It's always a trade-off. Okay, so we've walked through personalized RAG. Now, zooming out, what does this all mean for the big picture? You said earlier that when RAG gets really good at personalization, tracking users, adapting tools, generating context-aware stuff, it starts to look and act like an intelligent agent.
17:48Is that the personalized RAG plus A idea in action? Precisely. That's the natural evolutionary path. A RAG system that deeply integrates personalization inherently starts adopting agent-like capabilities. We can break down how this looks in agents, starting with understanding. Let's call it personalized understanding. The agent's empathy. Empathy. Wow. So this is like Raiji's query understanding, but deeper, more attuned to the person. Exactly. It goes beyond just understanding the request, understanding the user. It involves dynamic user profile, understanding modeling your preferences, your current context, your underlying intentions.
18:23Think of a health agent. It needs to understand your specific health profile to give relevant advice. Makes sense. And it also needs to understand its own role, right? Like role understanding. Is it a tutor, a travel agent, a friend? Yes. It's not just about understanding you, but understanding the role it needs to play in the interaction. There's research benchmarking how well LLMs can adopt and maintain specific personas. And I guess the magic happens when it combines both. That's the goal. User role joint understanding. How does the agent's role adapt based on the specific user it's interacting with?
18:56How does it maintain social appropriateness? Frameworks are emerging that incorporate personality info, even multimodal data, to make this interaction truly dynamic and tailored. Okay, so the agent gets me, gets its role, then it needs to act. That leads to personalized planning and execution. Smart actions. This must be like RAG's retrieval, but way more active. Absolutely. It's not just about fetching static documents. it involves real-time memory management. Agents need to integrate your history preferences, past behaviors, context. So it remembers our past conversations, my likes, dislikes.
19:30Exactly. Think of that AI travel planner again. It doesn't just look up flights. It remembers you prefer window seats, avoid red eyes, and like boutique hotels. It uses that memory to plan a truly personalized trip. This allows for more human-like behavior with memory streams and reflection. And the other part is tool and API calling. Yeah. This lets the agent actually do stuff for me, right? Precisely. This expands the agent's capabilities beyond just talking. It can interact with the outside world via tools and APIs to perform personalized tasks. Imagine a personalized shopping agent. It doesn't just suggest products.
20:04It could orchestrate APIs to compare prices across sites, check inventory, apply your loyalty points, and even complete the purchase, all based on your preferences. That's powerful. Okay, so after understanding and planning comes the final output. Personalized generation in agents. Tailored responses every time. Making sure what it says is right and feels right. That's the essence. It builds on the planning. The generated output needs alignment with user fact. It has to be accurate, grounded, avoid making things up, even while staying in character or being personalized. So getting the facts right while still sounding like me or the persona it's supposed to be.
20:39Exactly. That's one key alignment. The other is alignment with user preferences. This is about ensuring the output reflects your individual personality, values, interaction style. It means picking up on those subtle, implicit cues. How do they achieve that? Well, researchers are working on it, benchmarking role-specific alignment using psychological data sets to improve personality, fidelity, and agents. It's about making the agent's generated language genuinely resonate with the specific user's preferences and style. OK, this journey from REA to personalized agents is incredible, but it can't all be smooth sailing.
21:17What are the big roadblocks, the challenges ahead? Oh, there are definitely significant challenges. It's still an evolving field. I imagine one huge one is just the sheer scale of it all. Balancing personalization and scalability, all this personal data must add a ton of computational overhead. You hit the nail on the head. It absolutely does. Making this work efficiently for millions of users is tough. Future work really needs to focus on things like lightweight, adaptive embeddings, maybe hybrid frameworks that can manage the complexity without killing performance. And how do we even know if it's working well?
21:49Evaluating personalization effectively. Standard metrics for text quality probably don't cut it. Exactly. Metrics like BLEU or Rouget, they tell you if the text is fluent or overlaps with a reference, but they don't capture that nuanced alignment with user preference or long-term satisfaction. We need new benchmarks, new ways to measure if the personalization is actually good for the user. Right. And the massive one, preserving privacy. Especially if you combine on-device and cloud processing, that device-cloud collaboration. This is absolutely critical. Personalized AI often deals with incredibly sensitive data.
22:26The idea of using smaller LLMs on your device for processing local, private data while leveraging bigger cloud LLMs for general knowledge, that's a very promising direction. Like a personal AI vault on my phone. Kind of, yeah. Keeping your secrets safe locally while still benefiting from the power of the cloud. It's a complex balancing act, but essential for trust. What about the agent side? Personalized agent planning. You said that's still early days. It is relatively early. A lot of agent research is still focused on the core planning frameworks, integrating deep, personalized support seamlessly into those frameworks to really enhance how agents help users achieve complex goals.
23:06There's still a lot of work to be done there. It's about making the agent not just capable, but capable for you. And overarching all of this is the need for ensuring ethical and coherent systems. It's not just privacy, is it? It's bias, fairness, making sure the system behaves predictably. Absolutely. Addressing potential biases, creeping in from the data, ensuring fairness in how personalization is applied, maintaining coherence across all the different stages so the agent doesn't contradict itself or act erratically. These ethical safeguards and the need for holistic cross-stage optimization are paramount for building trustworthy systems.
23:41So quite a road ahead, but the direction seems clear. Okay, let's wrap this up. We've journeyed from the basics of why LLMs need help through how retrieval augmented generation provides grounding and then how layering personalization onto RO leads us towards these sophisticated, tailored AI agents. It's really about AI learning from everything, what you explicitly tell it, what it observes about your behavior to genuinely serve you. Yeah, the real promise here, the power of personalized AI, is moving beyond just generic facts or capabilities. It's about delivering information, interactions, and assistance that are uniquely relevant, that resonate with your context, your needs, your style.
Read the full transcript
24:20It's AI that strives to truly know you. Which brings us to a final thought to leave you with. As these AI systems get ever more personal, learning our habits, preferences, maybe even mirroring our personalities, how does that change our fundamental relationship with technology? Will it be purely empowering, giving us these incredibly intuitive tools that anticipate our needs and help us achieve more? Or is there a risk of creating personalized echo chambers, just reinforcing our existing views and biases? And critically, as AI becomes more intimately woven into our lives, how do we ensure it remains ethical?
24:52How do we guarantee it respects our privacy, adapts to our changing values, and ultimately upholds the trust we need to place in systems that know us so well?
From the publisher
We cover the comprehensive survey on the integration of personalization within Large Language Models (LLMs), specifically focusing on the evolution from Retrieval-Augmented Generation (RAG) frameworks to agent-based architectures. It systematically examines how personalization is incorporated across the pre-retrieval, retrieval, and generation stages of RAG, and extends this analysis to the more advanced functionalities of Personalized LLM-based Agents, including user understanding, planning and execution, and dynamic content generation. The survey also highlights key datasets, evaluation metrics, challenges, and future research directions in this rapidly evolving field, providing a valuable resource for researchers.




