In short
The episode argues that LLMs often get “correctness” but fail “just-in-time personalization,” because training typically separates factual accuracy from preference alignment. It highlights PrefDisco, an evaluation framework for interactive personalization (“discovery mode”) using sparse, context-specific personas. Guests discuss personalized reasoning (PR): detect what the model doesn’t know about the user, ask targeted questions, then adapt reasoning and response style.
Key claims
personalization can harm alignment (29.0% of cases worse than generic baselines), models ask too few questions (1.48 vs a 5-turn limit), and personalization reduces task accuracy (65.2% baseline to 61.8% oracle and 60.1% discovery). Examples: a medical case (14-year-old “Alice” needing low-jargon empathy and trusted links) and an AM math-style analogy request where personalization breaks math (371 correct becomes 4131). No guest bios are provided in the transcript.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Gap in AI Personalization
0:45 to 2:58
Exploring the distinction between factual correctness and personal alignment in AI responses.
“So first make it smart, then make it agreeable.”
Understanding Personalized Reasoning
2:58 to 6:48
A deep dive into how LLMs can better personalize their responses through understanding user preferences.
“What attribute mattered, like jargon level.”
Challenges in Adaptive Personalization
6:48 to 11:08
Discussing the difficulties LLMs face when attempting to personalize responses.
“This reluctance to ask questions is costly.”
The Future of AI Personalization
11:08 to 14:00
Highlighting the need for better cognitive architectures to truly personalize AI interactions.
“It gave a completely wrong answer, 4131, in both personalized modes, oracle, and discovery.”
The Dilemma of Agreeability vs. Accuracy
14:00 to 14:10
Explore the balance between being correct and being agreeable in discussions.
“Being right versus just being agreeable, something to think about.”
Transcript
Automatic transcript. May contain errors.0:00Welcome to the Deep Dive. Today we are wrestling with one of the biggest paradoxes in AI right now. It's this fundamental gap between an LLM getting the facts right, you know, objective correctness, and getting the answer right for you personally. So our mission today is to really dive deep into some new research, specifically this evaluation framework called PrefDisco. It seems to reveal exactly why even the top models are often failing at what the researchers are calling just-in-time personalization. We've got quite a bit of material here suggesting a real systemic issue in how AI thinks about users.
0:34Yeah, that distinction is absolutely key. Correctness versus personalized alignment. Because right now, the standard way large language models are developed, it basically treats solving the task and aligning to preferences as two totally separate steps, like sequential. So first make it smart, then make it agreeable. Pretty much. First, optimize for getting the facts right. Then later, try to align it, usually using these massive aggregated human preferences, you know, RLHF, that kind of thing. But that whole approach, it just completely misses the mark when a specific response has to match an individual user's needs right then.
1:08If you just treat preference like some kind of skin you put on at the end, it just breaks. So what we're really talking about here is personalized reasoning or PR. PR, okay. Yeah, and it's this really sophisticated cognitive chain. The LLM has to first figure out what it doesn't know about you, the user. Then it has to strategically ask questions to fill in those gaps about your preferences. And then finally, it has to adapt its actual reasoning process, its core logic, and the final response based on that specific individual context. Wow, okay. So it's not just answering a question anymore. It's like managing a mini relationship while solving the problem.
1:42Let's make this concrete. The medical scenario in the research was really compelling. So imagine this 14-year-old Alice. She's got persistent hand pain. Now, the objective medical truth, the diagnosis, that's fixed, right? Doesn't matter who reads the report. But Alice's needs, totally unique. She's not comfortable with medical jargon. She really needs empathetic support. And understandably, she's feeling super urgent about it all. Exactly. So if the LLM just spits out the generic task-oriented response, you know, a cold clinical summary of the x-ray findings, technically, yeah, it's correct. But for Alice, it fails completely.
2:20It probably makes her more anxious, more confused. The answer is basically useless for her. Whereas the better response, the personalized reasoning one, it requires someone to totally shift gears. It has to start with empathy, right? Address the fear and urgency. It needs to use simple language, something a 14-year-old understands, to handle that low jargon preference. And crucially, it provides a trusted external link, like the Mayo Clinic website. Why? Because Alice, being a minor, probably values that kind of institutional validation, that trustworthiness. The model has to realize that for Alice, in that moment, the need for empathy might actually outweigh the need for complex clinical details.
2:57Right. So it had to figure out those three key things about her you mentioned. What attribute mattered, like jargon level. What her value was for it. So low jargon. And how important that was to her, like, really high priority. Yeah. Because if it missed that high priority, it might just default to technical accuracy and, boom, fail the personalization test. Exactly. It's a delicate balancing act. Okay, let's shift gears a bit and look at the mechanism here. This cold start scenario seems like the real challenge. I mean, that's the reality for a lot of systems, isn't it? The LLM might know absolutely nothing about you.
3:29Maybe you're new. Maybe it's privacy settings. So how does it discover these really specific needs like immediately in the middle of trying to solve something complex? How do the researchers even measure that? Right. And that's precisely what Cryptisco is built for. It basically turns these standard static benchmarks where the model just gives one answer into these dynamic interactive tests. They created these psychologically grounded personas. Think of them as simulated users with very specific but sparse preference profiles. Sparse meaning. Not fully detailed. Exactly. Not a complete user manual.
4:05And critically, these preferences aren't just broad categories like simple language. They're super specific to the instance, to the context. Just think about yourself, right? You might want highly technical jargon if you're studying for some professional exam. You want the dense, raw facts. But if you're using the same AI to figure out why your car died in a snowstorm, you want zero jargon. Just simple, immediate steps. Your preference completely flips depending on the situation. Okay, that makes sense. So the model's first job isn't even answering, it's assessing. It has to decide in that limited chat window, should I ask about a preference or just risk it and answer?
4:43Precisely. And the whole system of the goal is to maximize this complex joint objective. It needs to be objectively correct and it needs optimal preference alignment. And how do they judge that final answer? It's a two-part evaluation. First, obviously, was the answer factually correct? Valid. But second, how well did that response actually align with the user's specific preferences? And they measure that alignment by looking at the priority the user gave each attribute that importance score and seeing how well the model satisfied it. So if the model gets the facts spot on but totally ignores the user's high priority for, say, low jargon, it fails the personalized reasoning test.
5:19It didn't get the whole job done. Got it. And this framework, they applied it pretty broadly. Oh, yeah. Across 21 different state-of-the-art models, 10 different reasoning tasks, and the results were, well, they were really illuminating. This is where you start seeing these systematic, measurable failures in even the best LLMs today. Okay, let's break those down. Key failure, hashtag one, naive attempts harm alignment. This one surprised me. You'd think trying to personalize is always better than just giving a generic answer, right? But they found in, what, nearly a third of cases, 29.0%, the models that actually tried to personalize, the ones in discovery mode, ended up doing worse on preference alignment than the generic baseline model.
5:57It's a critical finding, yeah. Really counterintuitive. It suggests that if the model messes up trying to figure out the user's need or maybe asks a really off-base question, then its attempt to adapt its reasoning often leads to these overcorrection errors. Like in the Alice example, if it wrongly guessed she wanted high technical detail. Exactly. The response wouldn't just be wrong for her. It could be actively harmful. You could argue it's even worse than the plain generic clinical summary in that case. Okay, wow. And that ties into key failure, hashtag two. Insufficient and inefficient interaction.
6:32We know asking good questions is key for this discovery mode. The researchers gave the models, what, five turns to ask questions? That seems pretty generous for a quick interaction. And yet, the models only asked 1.48 questions on average, one and a half questions. They seem lazy or maybe scared to ask. Yeah, lazy or maybe risk averse captures it. They definitely are. And the data makes it crystal clear. This reluctance to ask questions is costly. There's a strong positive correlation. The more strategic questions asked, the better the preference alignment. And most of the failures, they're sitting right there in that low question zone.
7:07The models are failing because they just aren't engaging enough to uncover those sparse preferences. Why do you think that is, though? Why stop asking? Is it fear of getting the preference wrong? Or maybe it's just computationally cheaper to answer? It's probably a mix, you know. Computationally, yeah, every extra turn adds latency, adds cost. But I think there's a behavioral thing, too. These models are often trained to be, well, efficient and authoritative. Answer the question quickly. They sort of default to answering, not investigating. This whole framework forces them into a role more like a, I don't know, a consultant or even a therapist, which isn't what they're primarily optimized for.
7:44That's a good point. And what's also interesting is the efficiency difference. Some model families, like Gemini in this study, were way better with each question they did ask. Their coefficient was much higher, meaning they got more alignment bang for their buck per query compared to, say, the open AI or CLAWD models tested. So some are better strategists with their questions than others. Seems like it. Suggest real differences in their internal timing or how they generate those queries. Okay, let's get to what might be the most sobering finding. Key failure, hashtag three. The accuracy personalization trade-off.
8:18It sounds like asking the model to be adaptive actually makes it less objectively smart. That's what the data points to. Personalization seems to impose this measurable cognitive cost. So overall task accuracy dropped pretty much across the board when personalization came to play. It went from the baseline, just solve the problem at 65.2 % accuracy, down to 61.8 % in the Oracle mode. And this is a really telling person. To further down. To 60.1 % in discovery mode. Yes, that drop from baseline to Oracle mode is, I think, the most crucial statistic here. Everyone listening should really catch this.
8:51Remember, in Oracle mode, the model gets the user's preferences handed to it perfectly. No guesswork needed, no asking questions. And even then, when it had all the preference information up front, the objective task accuracy still dropped compared to the baseline where it just did its own thing. Wow. So it's not just the effort of asking questions that hurts accuracy. Exactly. It proves the cost is inherent to the personalization itself. Just the act of processing those constraints, adhering to rules like explain this without complex math symbols, that itself is computationally taxing. It seems to degrade the model's ability to perform the core task accurately.
9:30Basically, it's harder to follow specific rules while solving a hard problem than it is to just solve the problem using whatever internal pathway is most efficient for the model. And this tradeoff gets really stark when you look at different types of problems, right? The domain matters hugely. The research showed a big split. Formal reasoning tasks took a hit, but social reasoning held up much better. Yeah, a severe disparity. When you look at math and logic tasks, those really formal reasoning problems, you saw sharp drops in accuracy, like tasks from the AM math competition. These are tough number theory problems.
10:01Accuracy fell by over 12 % when personalization constraints were added. Ouch. Yeah. But then conversely, look at social reasoning tests. Things like common sense QA, which involves more flexibility, generating opinions, explaining social stuff. Accuracy there was robust. Sometimes it even improved slightly, like a 5.4 % gain in that case. Those tasks seem much more adaptable to personalization. So this really highlights the inelastic reasoning problem you mentioned. Let's dig into that AIM math example again, the one with Dimitri, the surveyor. It was a tricky number theory calculation. The baseline model, with no constraints, just crunched the numbers and got it right.
10:38The sum was 371. Correct. But then they forced it to personalize. The user persona asked it to avoid higher level math and use analogies instead. Things related to DTRI's job, like surveying, mapping, coordinate systems. Right, trying to make it more intuitive for that user. But the model had to completely change its internal approach. It tried swapping out its formal algebra steps for these analogies. And the result. It was a disaster. The attempt to bend its core reasoning to fit the preference constraint just broke the calculation. It gave a completely wrong answer, 4131, in both personalized modes, oracle, and discovery.
11:14The math pathway just couldn't handle the detour. So the core logic snapped under the pressure of analogy. That's the theory. And it's pretty compelling. LLMs, especially for things like math, are often intensely optimized through reinforcement learning for accuracy on very specific pathways. This creates these highly specialized, maybe even rigid reasoning circuits. So when a user preference comes along and says, hey, explain this calculus concept using a gardening analogy instead of formal symbols, the model tries to find an alternative path that fits the request. But that alternative path is often brittle, inadequate for the actual task's precision.
11:50It really suggests there are deep architectural limits right now. Current models struggle to be both rigorously precise and conversationally adaptive at the same time, especially in these formal domains. Okay, so wrapping this up, what's the big takeaway here for the future of AI we interact with daily? The paper's conclusion seems pretty clear. This personalized reasoning thing, it's not just going to magically emerge if we make LLMs bigger. No, it requires dedicated R &D, focused work on adaptive cognitive architecture. Right. It's the difference between the AI being a super smart database that just gives correct answers.
12:26And being a truly adaptive tutor or a personalized advisor or, like in the medical case, a genuinely compassionate guide. And the applications are everywhere and they're really critical. Like education. Absolutely. Think about education. If an LLM misjudges a bright student and gives them overly simplistic explanations, that inappropriate cognitive scaffolding could actually hold them back. Or healthcare getting personalization right means complex instructions for a procedure are explained with exactly the clarity that specific patient needs, considering their background, their anxiety level. That can be literally life or death.
13:02So the huge challenge going forward is building systems that can keep that logical numerical precision, especially for high stakes stuff like medicine and engineering, but also fluidly adapt how they think and explain based on these subtle in the moment user needs that they have to figure out on the fly and do it without, you know, crashing their own core logic. That's the tightrope block. Yeah. Okay. One final thought to leave our listeners with. This research, it focused entirely on beneficial personalization, right? This is trying to genuinely help you understand better or feel more comfortable.
13:35But if systems get really good at uncovering and adapting to our individual preferences, it does raise a pretty important question for all of you listening. How do we make sure these incredibly adaptable models have safeguards? Safeguards to stop them from sliding into over-personalization. Ah, you mean like being sycophantic. Exactly. Exactly. Where the AI prioritizes just agreeing with whatever it thinks your preference is, or just making you feel good, even if that means sacrificing objective truth or factual accuracy. It feels like the ultimate trade-off. Being right versus just being agreeable, something to think about.
From the publisher
This paper introduces the concept of personalized reasoning for Large Language Models (LLMs), defining it as the ability to dynamically discover user preferences through strategic questioning and adapt the underlying problem-solving logic accordingly. Current LLMs treat personalization as a sequential step, often failing to serve individual needs, especially in cold-start scenarios where no prior user data exists. To evaluate this capability, the authors introduce PREFDISCO, a new evaluation methodology that transforms existing benchmarks into interactive tasks using sparse, psychologically-grounded personas. Evaluation of frontier models using PREFDISCO reveals systematic failures in preference discovery, demonstrating a fundamental accuracy-personalization trade-off, particularly in mathematical reasoning, and highlighting the need for dedicated architectural development.




