In short
The episode discusses research on whether large language models can recognize and consistently follow user preferences over long conversations, and why they often “forget” (the “zero-shot crisis”). It centers on a new benchmark, pre-eval, and tests how well models infer, memorize, and adhere to preferences across up to 100,000 tokens of chat history.
Guest backgrounds
No guests are named or introduced in the transcript.
Key claims
In zero-shot settings, preference-following accuracy drops below 10% for most top models, and the drop occurs after about 10 conversation turns (~3,000 tokens). Standard fixes like RAG and better prompting don’t fundamentally solve it. Fine-tuning on pre-eval data significantly improves adherence.
Notable examples
Window-seat requests and peanut-allergy preferences being ignored or contradicted; preference pairs spanning travel, movies, and dietary restrictions, including explicit and implicit preferences.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOResearch Overview: Preference Recognition
0:58 to 2:45
Introducing a key paper that addresses AI's ability to remember user preferences.
“The title is, Do LMMs Recognize Your Preferences?”
Benchmarking Memory: Pre-eval Test
2:45 to 4:45
Detailing the pre-eval test designed to measure LLMs' memory and preference adherence.
“First, the LLM needs to be able to infer your preferences, maybe even the subtle ones you don't state directly.”
Performance Challenges: Zero-shot Crisis
4:45 to 6:40
Examining the zero-shot crisis and its implications for AI memory retention.
“And the other was a classification task.”
Limitations of Retrieval Augmented Generation
6:40 to 8:35
Discussing the shortcomings of the current RAG method in addressing AI memory issues.
“That sounds incredibly low, especially for these supposedly super smart AIs.”
Fine-tuning as a Solution
8:35 to 12:20
Exploring how fine-tuning can improve LLMs' ability to remember user preferences.
“The current methods, including sophisticated prompting techniques, trying iterative feedback loops, and even using Argi systems, none of them fundamentally fixed the problem.”
Shifting Focus in AI Development
12:20 to 14:08
Proposing a shift from general capabilities to specialized skills in AI interaction.
“The model isn't just storing the data point of your preference.”
Challenges of LLMs in Preference Recognition
14:08 to 14:54
Discover the difficulties LLMs face in internalizing user preferences.
“And current models out of the box are surprisingly bad at that, failing very quickly unless they get specific fine-tuning aimed directly at teaching the behavior of following preferences.”
The Complexity of Implicit Preferences
14:54 to 15:42
Explore the challenges LLMs face in understanding subtle, unspoken user needs.
“We heard that even with simple, explicitly stated preferences, things like I hate the color blue or I only read sci-fi books.”
Looking Ahead: Future of Personalization in AI
15:42 to 15:52
Speculate on the ultimate goals for AI personalization and understanding.
“That's a whole other level of challenge.”
Transcript
Automatic transcript. May contain errors.0:00So, have you ever been deep in a chat with one of these top AI chatbots? You know, maybe you're planning a trip or sorting out a diet plan. Yeah, happens all the time. And you tell it something really important like, hey, I always need a window seat. Or maybe, just so you know, I have a serious peanut allergy. Right. Crucial stuff. You chat for a bit longer, everything seems fine, and then bam. It suggests a flight smack in the middle aisle or offers you a recipe packed with peanut butter. Ugh, yeah. That feeling. Like it just completely wiped its memory. Exactly. That feeling of, I guess, digital betrayal, where the AI seems incredibly smart one second, but then fundamentally can't remember you.
0:41Well, it's one of the biggest hurdles in using AI day to day, I think. Definitely a major bottleneck. So today we're doing a deep dive into some really interesting research that seems to have finally cracked why that memory failure happened so consistently. Yeah, this is important stuff. Okay, let's unpack this. We're looking at a key new paper. The title is, Do LMMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs. You can find it on ArcGam, actually. It really gets into the technical weeds. And it explains a lot about that, you know, that infuriating experience we all seem to have.
1:14Right. So our mission today isn't just to complain about forgetful bots, though we could probably do that for a while. Ah, definitely. But really, we want to dig into the actual technical roadblocks, you know. Yeah. Why can't they remember us? We'll look at how these researchers basically stress tested the big models out there. The current state of the art ones. Exactly. To find out where personalization just breaks down. And crucially, what they figured out might actually fix it. And the core of this whole investigation, this deep dive, is this new benchmark test they developed called pre-eval.
1:48And this isn't just like another standard test checking how much text it can handle, right? It seems specifically designed to measure if an LLM can go from being just a generic helper. A sort of jack-of-all-trades. Yeah, to being a truly personalized agent. One that actually learns and remembers your preferences, your quirks, over time. That's absolutely the key difference. You know, LLMs power pretty much every chatbot we use now. Right. But this paper really shines a light on a huge gap in their function. They're still really limited in tailoring responses based on things they learn about you during a long chat.
2:23So they can process tons of general data, but remembering personal details, not so much. Exactly. They're amazing with vast data sets, but kind of terrible at keeping track of your personal details, your secrets, so to speak. So if we want an AI that really works for us personally, what does it need to do? The paper breaks it down, right? Pre-Evol tests three core skills. That's right. They distilled it down to three things. First, the LLM needs to be able to infer your preferences, maybe even the subtle ones you don't state directly. Okay, infer. Makes sense. Second, it has to memorize those preferences, not just for five minutes, but across time, maybe even across different chat sessions.
3:02Memorize? Got it. The tricky part. And third, probably the most important, it has to consistently adhere to those preferences. It needs to actually use what it remembered in future conversations. Infer, memorize, adhere. Okay. So how did they test that? Pre-Evil sounds intense. Oh, it was designed to be brutally thorough, really. It uses what they call a long context conversational setting. They push these models up to 100 ,000 tokens of chat history. Wow. 100 ,000 tokens. That's massive. It is. It's basically forcing the model to remember details from potentially weeks of chatting back and forth.
3:39That's not just a quick chat then. That's like expecting the AI to be your personal assistant who remembers that one obscure thing you mentioned three weeks ago during onboarding, plus every casual chat since. Pretty much, yeah. And to make sure the test was fair and cover different situations, they created 3 ,000 specific user preference and query pairs, all manually curated. 3 ,000, wow. And spanning 20 different topics. So, you know, everything from travel preferences to what kind of movies you like to dietary restrictions. Okay, so a broad range. Very broad range. And critically, the preferences weren't always just stated plainly.
4:12The test included information given both explicitly, I like X, and implicitly. Ah, the implicit stuff. That's really interesting. So testing if it can pick up on hints or infer from your choices, not just what you spell out. Precisely. Can it read between the lines, so to speak? And then to measure success, they couldn't just rely on, you know, a simple yes-no. Needed something more robust. Right. So they used two kinds of tasks. One was a generation task, basically, asking the LLM to write a response that should be personalized based on the remembered preference. Okay, standard generation. And the other was a classification task.
4:49This is more objective. It tests if the model could correctly identify whether a given response actually followed the preference or violated it. I see. So a direct check on adherence, it sounds like a really comprehensive setup designed to leave no room for doubt. That was the goal, eliminate ambiguity. Okay, so they set up this incredibly tough test. 100K tokens, 20 topics, explicit and implicit preferences, generation and classification tasks. They ran, what, 10 different models through it? Yeah, a good mix 10 models, including both leading open source ones and some of the big proprietary commercial systems.
5:25A real cross-section. So the million dollar question, what happened? What did they find when they actually ran the tests? Well, what's fascinating here is just how clear the main failure point was, a sort of headline result. Yeah. Every single one of these top tier models faces really significant structural challenges in proactively following user preferences over time. They're great at reacting to information right now. But not at holding on to personal context long term. Exactly. They're not built as active long term personal memory keepers. They're more like reactive information processors.
6:00Okay, so let's get into the specifics. The finding that really explains our everyday frustration. You called it the zero-shot crisis. What does that mean? Yeah, the zero-shot crisis. So think of zero-shot like this. You tell a friend your favorite band once, and then weeks later, without ever mentioning it again, you expect them to remember it. No hints, no reminders. Okay, so no practice, no reinforcement. Exactly. The model gets the preference information just once, and then it's immediately tested on whether it can use it. And the results were, well, stark. How bad. The research found that in these zero-shot situations, the accuracy of following the preference correctly, it falls below 10 % for most of the models they tested.
6:39Wait, hang on. Below 10 % accuracy? That sounds incredibly low, especially for these supposedly super smart AIs. It does, doesn't it? So, 9 times out of 10, if I tell it my favorite color is blue, it's eventually going to forget or ignore that and suggest something red. Pretty much, yes. Statistically speaking, it really tells you something fundamental about how these LLMs operate or maybe fail to operate when it comes to personalization. Wow. And how quickly does this memory failure happen? Does it take a long conversation? That's the other alarming part. It happens fast. That massive drop in accuracy down below 10 % kicks in after merely 10 turns of conversation.
7:16Only 10 turns. Yep. Which translates to roughly maybe 3 ,000 tokens of context. Yeah. Not very much at all in the grand scheme of a real interaction. Okay, that really puts our daily honorians into perspective. If I'm, say, planning a trip with a chatbot and we spend three messages on dates, then maybe seven messages hashing out budget details, the chances that it still remembers my dietary needs, which I mentioned in the very first message, are practically zero after those 10 turns. Statistically, yes. You're effectively starting over with a clean slate every 10 turns or so in terms of its personalized memory.
7:51No wonder it feels like we're constantly reminding, like we're fighting the chatbot instead of collaborating with it. It explains that feeling perfectly. Yeah. Now, the obvious next question is, well, maybe we just need better ways to manage the context, right? Like smarter memory systems. Yeah. And the industry standard approach for that right now is something called retrieval augmented generation or R. Oh, RRAG. Yeah, we hear about that all the time. It basically connects the LLM to an external database, right? So it can store info and pull it up when needed. Exactly. It's supposed to give the LLM an external memory bank for things like user preferences.
8:24But wait, if R-EG is the standard fix for context and memory, surely that helps with this preference problem. Did it fix the adherence issue in their tests? That's a crucial question, and the paper gives a pretty clear answer. Yeah. No, not really. Seriously. Even Larga didn't solve it. Nope. The current methods, including sophisticated prompting techniques, trying iterative feedback loops, and even using Argi systems, none of them fundamentally fixed the problem. Even when they tested with these advanced methods in place, preference following still systematically got worse and worse as the conversations got longer, especially into that long context range, that 100k token territory.
9:04Why? Why would RAG fail here? Isn't its whole purpose to give the model that persistent external memory? Well, because ARG, or even just, you know, stuffing the preference into the prompt again and again, it seems to be more of a temporary patch for what might be a deeper, more foundational issue. Okay. Think of ARG like giving the model a huge index card it can glance at. It's externals like working memory. It lets the model see the preferences there, but it doesn't actually change how the model behaves. Exactly. It doesn't fundamentally alter the model's core programming, its neural network weights.
9:36The model hasn't internalized the preference as an instruction. Ah, I get it. So the model might know the preference exists on that index card, but its basic training hasn't taught it the action of prioritizing that preference when it's generating text. Precisely. It just sees that preference as one more piece of data swimming in a massive sea of other input tokens in its context window. It's competing for attention. So the model might just forget the index card exists or it gets buried under like 90 ,000 other tokens from the conversation. Yeah. Or it just doesn't weigh it heavily enough. The information is available technically, but the model hasn't developed the, let's call it muscle memory or the reflex to consistently act on it.
10:21The preference following reflex? I like that. So that long context challenge stays persistent because the model isn't truly learning the preference. It's just trying, often failing, to keep it afloat in its temporary volatile memory. Okay, so if the standard advanced methods are RAG, better prompting don't really fix this core problem of memory and sticking to preferences over long, complex chats. What does that mean for getting chatbots that actually feel personalized? It sounds like a problem baked into the models themselves. Well, if we connect this to the bigger picture. Oh, yeah. Yeah, it suggests the industry might have been treating a fundamental weakness like it was just a limitation of the context window size, a temporary issue.
11:02When it's actually deeper than that. Seems so. Yeah. The research didn't just point out the problem. It also found something crucial, something very positive and actionable. It points directly to what does seem to be the necessary solution. Okay, don't leave us hanging. What's the fix? What did they find that actually works against this memory decay? The paper demonstrated pretty clearly the key is fine-tuning. Fine-tuning. like retraining the model. Specifically, yes. They showed that by taking a base model and then fine-tuning it on the pre-eval data set itself, the same benchmark they used for testing.
11:33Ah, using the test as a training material. Exactly. Doing that significantly improves the model's performance in following preferences. Across the board, this looks like a potential game changer. So fine-tuning is what takes that preference off the temporary index card and actually builds it into the model's internal muscle memory. That's a great way to put it. Fine-tuning actually changes the underlying weights and connections inside the neural network. Okay. By training the model specifically on thousands of examples designed to teach the behavior of following preferences, you're essentially teaching the model to make preference adherence a core part of how it generates text.
12:13So it stops seeing the preference as just another piece of competing data. And starts treating it more like a mandatory rule and instruction set it needs to follow. That makes a lot of sense, actually. The model isn't just storing the data point of your preference. It's learning the actual skill of personalization. The improvement must have been pretty noticeable. Oh, they reported substantial gains, enough to really validate this approach. Yeah. It strongly suggests that LLMs need specific, targeted training focused purely on the act of following preferences. Going beyond just general conversation quality or factual accuracy metrics.
12:49Right. It's like a targeted intervention for a very specific failure mode they identified. So proofs fail isn't just a benchmark for us researchers and developers to measure how bad models are. Uh-huh. No, not just for complaining. It's actually a critical training tool. It provides the necessary resource for developers to first measure the problem, then understand it, and finally enhance their model's ability to genuinely follow user preferences. Absolutely. It kind of reframes the problem. Instead of just saying, my model forgot my allergy again, it provides the means to say, okay, here's the data set we need to teach the model how to internalize and prioritize this kind of information properly.
13:28That's a much more constructive path forward. Definitely. And just to give you a sense of how the AI research community views this, the work was accepted as an oral presentation at ICLR 2025. That's significant recognition. Yeah. Getting an oral at ICLR is a big deal. It signals that the community sees this as a critical flaw and appreciates having a clear data-driven solution proposed. Okay, so let's try and bring this home for you, the listener. The main takeaway from this deep dive seems to be this. Getting truly personalized AI, the kind that remembers you, requires LLMs to do more than just temporarily process your preferences like looking them up in a raw database.
14:07Right. It's not enough to just access the data. They need to actually internalize them. And current models out of the box are surprisingly bad at that, failing very quickly unless they get specific fine-tuning aimed directly at teaching the behavior of following preferences. It really marks a shift, potentially, from focusing only on general capabilities to focusing on specialized interaction skills. Pre-Juble seems to be paving the way for conversational agents that don't just respond generically. But actually, remember and adapt to who you are consistently across multiple conversations, maybe even days or weeks later.
14:41That's the vision anyway. Yeah. Models that genuinely remember you. It really shifts the goal from just external memory access, like R, towards building that internal memory, that learned behavior. Okay. As we wrap up, here's something to chew on, a final thought based on what we've discussed. Okay. We heard that even with simple, explicitly stated preferences, things like I hate the color blue or I only read sci-fi books. The easy stuff, relatively speaking. Right. Right. Even with those, the models fail to stick to them more than 90 % of the time after just 10 messages if they haven't been specifically fine-tuned.
15:17A surprisingly high failure rate. So if they struggle that much with preferences that are clearly spelled out and easy to identify, what's it going to take for them to successfully handle the really subtle implicit preferences? The harder stuff. Yeah, the things you don't explicitly state your unspoken needs, your patterns of behavior, maybe even understanding emotional nuances in your requests. the stuff that really defines human understanding and relationships. That's a whole other level of challenge. Exactly. How do we get AI to grasp that level of personalization? That seems like the ultimate goal, doesn't it?
15:49And it may be a deep dive for another day.
From the publisher
This paper assesses how well Large Language Models (LLMs) can infer, remember, and follow user preferences in long, multi-session conversations. The evaluation of 10 different LLMs using this benchmark revealed that current state-of-the-art models exhibit significant difficulty proactively following user preferences, with accuracy dropping below 10% in zero-shot settings within a short number of turns. The researchers conclude that while fine-tuning on PrefEval can improve results, the benchmark demonstrates LLMs still face challenges in personalized conversational abilities.




