Learning Personalized Agents from Human Feedback

21 Feb 2026 · 15 min · 7 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Personalized agents from human feedback (PAHF) that adapt to changing user preferences using a read-write memory loop, avoiding “static personalization” failures like preference drift.

Guests/backgrounds

No named guests; two hosts discuss the framework and engineering details, using examples (smart home, music, coffee) and research case studies (robot manipulation, online shopping).

Key claims

Static models rely on past snapshots and update only temporarily; PAHF separates an LLM “brain” from an explicit memory module that can be written in real time. It uses a three-step loop: pre-action clarification (ask), act, then post-action feedback integration (write/update rules). A salience detector filters noisy/irrelevant inputs.

Notable examples

“Kate” drink preferences evolve from Coke to tea when sleepy, then to coffee when sleepy; rules are overwritten on conflict. Robotics: changing object-placement preferences caused catastrophic errors without PAHF. Shopping: “persona shifts” from budget to luxury; PAHF pivots faster and reduces personalization error.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Limitations of Static Personalization

0:45 to 2:15

Exploring how current models fail to adapt to changing user preferences.

“But they can't remember that you stopped drinking dairy last month or that you switched from techno to lo-fi.”

Introducing PAHF: A New Framework

2:15 to 4:33

Discussion on the framework for personalized agents from human feedback (PAHF).

“The first is what we call the cold start or the new user problem.”

The Three-Step Memory Loop

4:33 to 6:05

Explaining the three-step process of PAHF for better user interaction.

“Which I assume is just the AI asking me what I want.”

Case Studies: Preference Drift in Action

6:05 to 12:17

Analyzing how PAHF adapts to evolving user preferences through case studies.

“Okay, let's make this concrete for everyone listening.”

The Implications of Precise Memory

12:17 to 14:00

Discussing the potential effects of AI knowing users better than they know themselves.

“The personalization error dropped significantly compared to the baselines.”

The Nature of AI as an Autobiographical Mirror

14:00 to 14:23

Explore the implications of AI acting as a reflection of our identities.

“It becomes an externalized autobiographical memory.”

Reflections on AI and Human Interaction

14:23 to 14:47

Discuss the evolving relationship between humans and AI assistants.

“I think I need to go have a very serious conversation with my smart speaker.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00I have to tell you, I hit a breaking point with my smart home ecosystem this week. Oh, no. Let me guess. The digital intern that never sleeps, but also never actually learns. Exactly. I mean, it wasn't even a complex task, right? I was just trying to get it to manage my morning routine. And I asked it to start my focus mode playlist and it played the exact same high energy techno track it played like three years ago when I was training for a marathon. Oh, wow. I haven't run in two years and I've told it explicitly that I need lo-fi beats for work now. But it just it keeps reverting to the old data.

0:35It's like it has a perfect memory of who I was, but zero concept of who I am today. Yeah, that is the universal frustration right now. I mean, for everyone listening, you've probably felt this. We have these incredible models that have, you know, read the entire Internet. They can pass the bar exam. Right. But they can't remember that you stopped drinking dairy last month or that you switched from techno to lo-fi. It honestly feels like a betrayal of the word smart. It really is. And it stems from how we've historically built personalization. But that's exactly why we're digging into this today, because there is a new framework emerging that is trying to solve exactly this problem.

1:10It's called PAHF. PAHF. OK, acronyms are my love language. What are we looking at here? So it stands for personalized agents from human feedback. The core idea is moving away from what we call static personalization, which is what your music app is doing, to continual personalization. Continual. That implies it's, well, never finished. Exactly. The goal is to stop treating the user, to stop treating you as a fixed data point in some old log file, and to start treating you as a dynamic, evolving person. Okay, but let me play devil's advocate for a second because my apps already claim they learn from me, right?

1:46If I click dislike on a song, it skicks it. If I buy a pair of shoes online, it shows me socks. Isn't that learning? It's learning in a very specific, very limited way that's usually based on static data sets. So the model looks at your past logs, you know, snapshots of your behavior from last month or last year, and it builds this really rigid profile. So it's basically looking at a photo of me from 2021 and just assuming I'm still that guy. Precisely. And that static approach, it really fails in three very specific ways. The first is what we call the cold start or the new user problem. Which is when I sign up for a new service and has literally no idea who I am, so it just throws the generic popular stuff at me.

2:27Yes. Without a history log, a static model is essentially useless. It has to wait until you generate enough data to form some kind of pattern. You can't just ask you what you want and immediately get it. I mean, okay, that I can forgive. Everyone has to start somewhere, right? Sure, but what about when you explicitly correct it? Like with your music. You said, stop playing techno. Oh, right. That's the infuriating part. That's the second failure, the stubborn error. Static models just aren't designed for real-time correction. Yeah. When you yelled at your speaker, it probably apologized, right?

2:58Yeah, it said, I'm sorry, I'll remember that for next time. But it didn't update the underlying model. It just flagged that interaction in a temporary little context window. The actual training of the system happens offline maybe weeks or months from now. So tomorrow, it just reverts right back to the weights it was trained on. So it's just gaslighting me. Pretty much. And the third failure, this is the really big one. It's exactly what you're experiencing. Preference drift. Preference drift. That sounds like a complicated relationship status. I mean, it basically just means people change. Your preferences aren't stationary.

3:33You go vegan. You move to a new city. You get a new job. Right. Static models are frozen in time. They struggle desperately to handle the fact that user A today is different from user A yesterday. So this PAHF framework is supposed to fix this. Is it just like retraining the AI's brain constantly? Because that sounds wildly expensive. You can't retrain a massive language model every time I switch my coffee brand. No, no, you can't. And that's a really crucial distinction here. Retraining the parameters, the actual brain of the model, is way too slow and way too costly. PAHF bypasses that entirely by using what we call a read-write memory architecture.

4:13Read-write, like a hard drive. Very similar concept. Most AI today is read-only. It reads its training data. PAHF gives the agent a separate explicit memory module. Think of it like a digital notepad that it can actively write to in real time. Okay, so it has a brain for processing the language and talking to me, and a separate notepad for remembering me. Yes, and it uses a very specific operational loop to keep that notepad accurate. It's a three-step loop. Walk me through it. What's step one? Step one is pre-action interaction. We can just call it the ask. Which I assume is just the AI asking me what I want.

4:45It's a bit more nuanced than that. It's about detecting ambiguity. Before the agent acts, it queries its memory notepad. If it sees a gap, like it thinks, I know he wants music, but I don't know his current vibe, it pauses and asks a clarifying question. Instead of just guessing and hoping for the best. Exactly. It resolves known uncertainty before making a mistake. Then comes step two, action execution. It takes your answer, synthesizes it with its memory, and does the task. Standard stuff. Step three is the game changer, post-action feedback integration. The oops factor. This is where the magic happens.

5:21Let's say the agent bring you a coffee based on its memory. You say, actually, I'm trying to cut down on caffeine. I want a decaf. In the old static model, that feedback just disappears into the void. Right. But in PAHFs, the agent takes that correction and immediately writes a new rule into that explicit memory notepad. It updates its belief system instantly. So it's closing the loop entirely. It asks, it acts, and if it messes up, it rewrites the rulebook on the fly. It's finding that balance, we call it the Pareto frontier, between asking too many questions and making too many mistakes. Because you don't want an agent that asks you about every tiny detail, that's annoying.

5:56Exactly. But you also don't want one that guesses wrong constantly. PAHF uses this memory loop to balance those two extremes. Okay, let's make this concrete for everyone listening. There was a great case study in the material regarding a user named Kate and her drink preferences. It seemed simple on the surface, but the underlying logic was fascinating. The Kate saga. It's the perfect stress test for this. So day one, Kate asks her agent to bring her favorite drink. Now, the agent doesn't know her yet, so it uses step one. It asks, she says, Coke. The agent writes down in its memory, Kate likes Coke.

6:31Simple enough. Day two, Kate says, I'm sleepy. Get me a drink. The agent checks its memory, sees Kate likes Coke, and brings her a Coke. Which is a completely logical deduction based on what it knows. But Kate says, actually, when I'm sleepy, I prefer tea. Ah, so a contextual preference. Right, and this triggers step three. The agent doesn't delete the Coke rule. It adds a condition. If sleepy, then tea. It's essentially building a decision tree for Kate. Okay, but here's where humans get messy, right? What happens on day three when she inevitably changes her mind again? And that's the preference drift test.

7:05So day three, she says, grab a drink to wake me up. The agent remembers the rule, if sleepy, then tea, and brings her tea. But Kate says, actually, I want coffee now when I'm sleepy. The agent should be completely confused now. It literally has a hard rule for tea. A static model would break or just ignore the new instruction. But the PAHF agent realizes the conflict, determines that the new real-time feedback overrides the old rule and overwrites the entry. So now it's, if sleepy, then coffee. It effectively tracks her drift. I want to look under the hood here for a second. You mentioned a notepad, but technically speaking, how does that work?

7:41Is it just a giant text file? Because if I'm chatting with this thing for years, that file is going to be massive. How does it find sleepy equals coffee in a million lines of text? You've hit on the exact engineering challenge. It is not just a text file. They use a decoupled architecture. So you have the LLM, the brain, and then you have a highly structured database for the memory. Structured how? They often use a combination of methods. For really simple rules, it might just be a key value store, like a standard Squalite database. But for more complex, fuzzy concepts, they use a vector database, something like FiloGuy.

8:16Okay, for the listeners who aren't data engineers, break down why a vector database matters so much here. Sure. So a normal database looks for exact word matches. If you search the word sleepy, it scans for the exact letters, S-L-E-P-Y. A vector database stores the actual meaning of the concepts as numbers. Oh, I see. So if you say, I'm feeling drowsy or I need a boost, the vector database knows those phrases are semantically related to sleepy, and it retrieves the relevant coffee preference anyway. So it doesn't need me to use the exact same keyword every single time to trigger the rule? That is huge for making it feel natural.

8:52It is. It makes the retrieval semantic rather than literal, which is how humans actually communicate. But wait, I see a massive risk here. If it's constantly recording everything to this vector memo, what if I'm just venting? What if I've had a bad day and I say, I would literally kill for a massive burger right now, but I'm actually on a strict diet. Does it log? User wants to commit murder for me. That is the hallucination and noise risk. Exactly. And to solve that, the PAHF framework uses what they call a salience detector. A salience detector, like a bouncer for the memory. That's a great way to put it.

9:26It's usually a smaller, specialized LLM that acts as a judge. It sits right between you and the memory database. It analyzes every single interaction and asks, is this piece of information actually useful for a future task? So if I sneeze or if I just complain about the traffic. The salience detector filters it out completely. It only writes down information that constitutes a genuine preference or a reusable rule. It keeps the memory hygienic. That seems critical. Otherwise, the sheer volume of noise would drown out the signal eventually. Exactly. And it handles conflict resolution, too. It detects if a new note contradicts an old one, like with Kate's coffee, and decides whether to update the existing rule or create an entirely new context.

10:09Now, we've been talking a lot about chatbots and coffee and music, but looking at the research, they didn't just test this on conversational AI. They applied it to embodied manipulation and online shopping. Yes, and the embodied manipulation, which is robotics, is fascinating because the stakes are so much higher. How so? Well, in a chat, if the AI makes a mistake, it prints the wrong paragraph of text. In robotics, a mistake means dropping a plate or putting your laptop in the dishwasher. Right. Telling a robot to clean up the table can be interpreted very differently depending on what's on the table.

10:43Exactly. So they used a simulator where a robot arm had to organize objects. But the trick was the human's preferences for where things should go kept changing. Okay. First it was put the bowl on the left, then later put the bowl on the right. The physical version of preference drift. Yes. And they found that without the PAHF loop, specifically without the ability to ask questions before acting, the robots made catastrophic errors. They would just blindly guess. I think he wants the kitchen knife and the toaster. Here we go. Right. And without the post-action feedback, the robot never learned when the user changed their mind.

11:20It would just keep organizing the desk the old way forever. But PAHF allowed the robot to adapt its physical sorting strategy in real time. And what about the shopping benchmark? Because that seems like it would be much easier than robotics. You'd think so, but they introduced something called persona shifts. They simulated a user who starts off shopping for budget items, very price conscious, and then halfway through the session completely shifts to looking for luxury items. Which happens all the time in real life. Maybe I got a bonus at work, or maybe I'm suddenly buying a gift for someone else.

11:51Standard recommendation engines are absolutely terrible at this. They see you bought cheap socks 10 minutes ago, so they keep recommending cheap shirts. Even if you start explicitly searching for silk ties, the momentum of that old data just weighs it down. It essentially anchors you to your past self. Exactly. But the PAHF agents, because they weigh recent feedback and explicit corrections so heavily in that explicit memory, they were able to pivot almost instantly. The personalization error dropped significantly compared to the baselines. It really sounds like the key here is that combination.

12:26You can't just have an agent that asks questions, and you can't just have an agent that listens to corrections. You need the active, continuous loop of both. That is the core takeaway right there. Pre-action interaction prevents immediate disasters with new users. Post-action interaction enables long-term evolution and triff tracking. If you remove either one of those channels, the whole system just falls apart. It's essentially moving from I am executing code to I am building a theory of mind. That's a heavy concept, but it really fits here. In psychology, theory of mind is the ability to attribute beliefs and intents to others.

13:01PAHF is effectively giving AI a structured engineering way to build a theory of user. Which is deeply intimate when you think about it. It is very intimate. It means the AI isn't just serving you anymore. It's actively studying you. And that actually leads to a slightly unsettling thought I had while we were preparing for this. If this system works perfectly over a long enough timeline, if it logs every drift, every tiny change in my taste, every contradiction, eventually that Squealite database is going to know me better than I know myself. It is a very distinct possibility. Human memory is incredibly fallible.

13:37We edit our own pasts constantly. We convince ourselves we've always liked jazz music or that we never really liked that one ex-partner. We totally retcon our own lives. We do. But the explicit memory in a PAHF agent, it doesn't forget unless you explicitly tell it to. It has the receipts. So five years from now, I could be arguing with my agent. I could say I never liked spicy food. And it pulls up a timestamped log from 2024 saying, actually, you ordered the extra hot curry 14 times that year. It becomes an externalized autobiographical memory. And the provocative question is, will we eventually resent that perfectly accurate mirror?

14:13or will we start relying on our AI agents to remind us who we actually are? Hey Siri, who am I today? It might literally be the most accurate answer you ever get. Well, on that slightly existential cliffhanger, I think I need to go have a very serious conversation with my smart speaker. Maybe if I give it some clearer post-action feedback, we can finally get along. Be patient with it. It might just need a vector database upgrade. Good point. Yeah. Well, to everyone listening, thanks for taking this deep dive with us. It's a wild world of adaptive intelligence out there, and it's moving fast. Always a pleasure to unpack this stuff.

14:46We'll catch you on the next Deep Dive.

From the publisher

This research introduces a framework for continually personalizing LLM agents by utilizing a streamlined memory system that learns from two types of human feedback. The system combines pre-action queries, which clarify ambiguous requests before they are executed, with post-action feedback to correct errors when an agent makes an incorrect assumption. This dual approach allows the agent to build a clean database of user preferences and effectively adapt when those preferences change over time, a phenomenon known as preference drift. Evaluated through online shopping and embodied agent scenarios, the method ensures agents do not remain "confidently wrong" but instead refine their behavior through a detect–summarize–integrate pipeline. Ultimately, the study demonstrates that integrating both reactive and proactive feedback channels significantly improves the accuracy and scalability of personalized artificial intelligence.

More from Best AI papers explained

All 475 episodes
Learning Personalized Agents from Human FeedbackBest AI papers explained · 15 min
Listen in VO