PAD: Personalized Alignment of LLMs at Decoding-Time

19 Feb 2026 · 14 min · 9 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Episode topic: PAD (Personalized Alignment of LLMs at Decoding Time) tackles the “alignment tax” where RLHF-trained LLMs drift to a safe, average “committee” voice. It proposes shifting alignment from training to inference by changing token decoding using a user-specific preference vector.

Guest backgrounds

No guests mentioned; this is a solo “Deep Dive” discussion of a research paper.

Key claims

PAD adapts a generic LLM to arbitrary styles (expert, humorous, concise) without retraining or fine-tuning. It uses a decoupling architecture: a base LLM proposes candidate tokens, while a smaller personalized reward model (PRSRM) re-ranks them using successor features (a “nutrition label” vector) and a prompt-derived weight vector (dot product).

Notable examples

Apple internship offer letter becomes precise when “expert” is set; humor vector changes tone with jokes like “AI for the win, baby.” Reported results: 84% win rate on a “Personalized Soups” dataset; best on 7/10 metrics on HelpSteer2. Tradeoffs: 2–3x inference latency and ~17GB extra GPU memory; model-agnostic “plug-and-play” reward model.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Frustration with LLM Outputs

0:45 to 1:40

Discussion on the limitations of current LLM outputs and the concept of alignment tax.

“The average user doesn't actually exist.”

Introducing the Paper: PAD

1:40 to 2:58

Overview of the paper titled 'PAD' and its promise to personalize LLM outputs without retraining.

“Why is the current way of doing things, the status quo, kind of failing us?”

Current LLM Training Issues

2:58 to 4:55

Examining the shortcomings of reinforcement learning from human feedback in LLM training.

“To make a genuinely funny model, you need a huge data set of humor, an expensive fine-tuning run, and then you have to host that model.”

Mechanics of PAD's Approach

4:55 to 6:46

Explanation of how PAD modifies language model behavior in real-time using a decoupling architecture.

“It has to be a likely word, and it has to fit the user's preference.”

Successor Features and Creativity

6:46 to 7:58

Discussion on successor features and how they allow LLMs to understand and generate creative content.

“This is where it gets really surprising.”

Performance and Practical Implications

7:58 to 10:00

Analysis of PAD's performance compared to other models and its practical implications on usage.

“You're composing complex behaviors from these fundamental atomic features of language.”

Challenges and Limitations of PAD

10:00 to 12:34

Exploring the challenges, including inference cost and memory requirements of the PAD model.

“There's no such thing as a free lunch in computer science.”

Ethical Considerations in AI Flexibility

12:34 to 13:34

Discussion on the ethical implications of extreme flexibility in AI models and potential echo chambers.

“Doesn't this kind of extreme flexibility create a paradox?”

Discussion on Decoding-Time Alignment

14:00 to 14:11

Learn about a significant advancement in personalized alignment for LLMs.

“It's a brilliant piece of work that really does solve the average user problem as long as you have the VRAM.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome back to the Deep Dive. so we've all been there right i think everyone listening has felt this frustration especially if you're using llms for you know anything serious you spent all this time crafting the perfect prompt you give it context persona you ask for nuance and then what do you get back corporate beige corporate beige exactly it's i mean it's grammatically perfect sure it's usually factually okay, but it just feels like it was written by a committee. A committee trying to avoid a lawsuit, yeah. And it doesn't matter what I ask for. A stand-up routine, a breakdown of quantum physics, it always seems to drift back to this safe, average voice.

0:40Well, that's the alignment tax we're all paying. We have these incredibly powerful models, but they're trained to satisfy a, quote, general preference. And the average user. The average user doesn't actually exist. They're a mathematical ghost. Which brings us right to the focus of today's deep dive. We're looking at a paper called Payeti, Peripheralized Alignment of LLMs at Decoding Time. It's from researchers at Zhejiang University, National University of Singapore, and the University of Washington. And the promise here is, I mean, it's massive. They're claiming they can take a generic LLM and make it adapt to any personality you want.

1:16Expert, humorous, concise, you name it. And here's the kicker, without a single second of retraining. That is the whole game right there. No retraining, no fine-tuning. This isn't about changing the model's brain, so to speak. It's about changing how it talks in the moment. So you're shifting alignment from the training phase. To the inference phase, in real time. Okay, so before we get into the mechanics of how they pull this off, and the math is pretty clever, let's just set the stage. Why is the current way of doing things, the status quo, kind of failing us? Well, the standard pipeline is RLHF, reinforcement learning from human feedback.

1:54Right. You get the model to generate two answers, a human rater clicks on the one they like better, and you train a reward function based on that preference. But the whole problem is in that human rater part, isn't it? The paper calls it the pluralistic world problem. It is. If you just aggregate feedback from a thousand people, different cultures, different politics, different education, you don't get a perfect answer. You get a smoothie. You get a smoothie where you've blended, I don't know, steak, strawberries, and kale. It's weirdly unappealing to everyone because it tries to be everything to everyone.

2:27So if half your raters want short answers and the other half want these long academic essays. The model learns to just give you a medium-length, semi-formal paragraph. It avoids being bad, but it never gets to be great. It optimizes for boredom. And that's not even getting into the individual user. I mean, my needs change constantly. If I'm debugging code, I want an expert. Terse, technical, no fluff. But if you're writing a birthday card for your grandma... I want creative and warm. And with the current system, to get those reliably, I'd basically have to fine-tune two completely different models, wouldn't I?

3:00Pretty much, yeah. You'd hit the training bottleneck. To make a genuinely funny model, you need a huge data set of humor, an expensive fine-tuning run, and then you have to host that model. And then you start all over again if you want a serious medical bot. It's expensive, it's slow, and it's just not flexible. So pay depersonalized alignment at decoding time says, okay, let's stop trying to bake the personality into the model itself. Right. Let's intervene at the last possible second, right when the model is choosing its next word. This is where we get into what they call the decoupling architecture.

3:34Okay, let's break this down because this is the really cool technical part. So normally an LLM is just doing next token prediction, right? Right. It looks at the sentence so far and gives a probability for every possible next word. Correct. It might say, you know, the word cat has a 20 % chance, dog has a 15 % chance, and so on. And then you usually use a decoding strategial. Like greedy decoding, where you just pick the highest probability word every time. Which is fast, but it can get really repetitive. But PAYE steps in right at that moment before the word is actually chosen. How? It introduces a second, much smaller model.

4:06They call it the personalized reward model, or PRSRM. PRSRMM. them. The best way to think about it is like a sidecar attached to a motorcycle. The base model, that's the motorcycle, it does all the heavy lifting of language. Grammar, facts, logic. It gives you a list of candidate words. So the base model says, okay, my top three choices are apple, banana, or cherry. Exactly. Then the sidecar, the PERSRM, it looks at those options, and it also looks at your instruction. Like maybe your prompt said, be healthy. Oh, okay. And it calculate the reward score for each of those words. So Apple gets a high score, banana is medium, and cherry, you know, maybe that's low if we're thinking about pi.

4:47And then it just combines the two scores, the base model's probability and the reward model's score. It's a weighted combination, yeah. The system picks the word that maximizes both. It has to be a likely word, and it has to fit the user's preference. And it does this over and over for every single token. That seems like it would be slow, but okay, hold on. How does the reward model know what healthy or expert even means without being retrained for every single prompt? This gets to that successor features idea, right? That felt like the magic sauce. It absolutely is. This is the big aha moment of the paper.

5:20See, a normal reward model just spits out a single number, a scalar, like this is a 7 out of 10. It was a black box. A 7 doesn't tell me why it's a 7. Exactly. So PP changes this. Instead of one number, the model learns a vector, a multidimensional list of the features of the text. It's almost like a nutrition label for the text. That's a fantastic analogy. Yeah, that's perfect. Imagine every potential sentence has this label. The model looks at it and says, okay, this sentence is 80 % formal, 10 % humorous, 90 % concise. It calculates these ingredients completely separately from what you, the user, want.

5:56So it understands the attributes of language before it even knows my preference. Yes, those are the successor features. Then your prompt like, be an expert, that gets translated into a really simple weight factor. The diet plan. A diet plan. It just says, I want high formality, high conciseness, and turn the humor way down to zero. And the math is just lining them up. It's just a dot product. You multiply the nutrition label by your diet plan. The word that best fits the profile you asked for gets a huge boost. Okay, now I get why it's so flexible. If I change my mind and say, actually be funny now, I'm not retraining anything.

6:31You're just handing it a new diet plan, a new weight vector. The model already knows which words are ingredients for humor. It just starts picking them instead. And this leads to what I thought was the most amazing result in the paper, the unseen preferences. This is where it gets really surprising. They tested it on things it was never trained on. They did. The reward model was trained on standard stuff, helpfulness, harmlessness, you know, the usual. But then at test time, they prompted it with things like be an expert, be informative, and be creative. And creative is a really hard one. If you haven't trained a model on what creative looks like, it just starts making things up.

7:08But it worked. PayPal significantly outperformed the baselines on these new unseen preferences. But how? I mean, if it's never seen a creative label during its training, how does it know what to do? Because creativity isn't some brand new magical thing. It's a combination of features the model has already seen. Creativity might just be a recipe that calls for, say, high vocabulary diversity, low repetition, and unusual sentence structure. Since the successor features have already mapped out all those basic ingredients, it can navigate to creative just by following the new recipe. So it's like a chef who's never been asked to make a beef wellington before.

7:47Right. But they know how to make puff pastry, they know how to cook beef, and they know how to make a duck cells. You just give them the recipe, the preference vector, and they can assemble it from the skills they already have. That's it exactly. You're composing complex behaviors from these fundamental atomic features of language. So let's look at the actual numbers, because the theory is elegant, but does it win? They put it up against something called personalized soups, which still sounds like a terrible lunch menu. The pea soups method is basically about merging the weights of different models.

8:17and PA, well, it crushed it. On that PSUPS dataset, PAD had an 84 % average win rate. 84%, that's massive. In AI benchmarks, people fight over 2 % or 3%. It's statistically dominant, yeah. And on another dataset, HelpSteer2, it was the best on 7 out of 10 metrics. But honestly, the numbers are less important than the qualitative examples. The Apple internship offer letter. Yes. That case study really shows you what this feels like. I love this part. The prompt was super simple. Run an offer letter for an AI research internship at Apple. So first, they just ran the base model without PA'd, the vanilla version.

8:57And it was fine. We are pleased to offer you the position. You know, boilerplate. It sounded like a template. Then they switched on the expert preference vector. And the whole tone shifts instantly. The language gets precise. It's not just the team, it's the AML team. It starts listing specific responsibilities like computer vision and NLP. It sounds like a real hiring manager from a top tech company. Okay, but the real acid test was the humorous vector. I've got the text right here. It didn't just hack a joke on the end. The whole thing changed. It opens with AI for the win, baby. Which I'm pretty sure would create some HR issues, but it definitely proves the point.

9:32And this line, coffee flows like a never-ending fountain. And my favorite, we promise not to make you work too hard. unless you ask nicely. And the thing to remember, the really crucial part, is that the base model is identical in all three cases. The engine is the same. The only thing that changed was the navigator, the purse RM, telling it which conversational turns to take. So we have a system that's more flexible, it performs better, and it doesn't need retraining. There has to be a catch. There's no such thing as a free lunch in computer science. Oh, there is a lunch bill here. And it's a pretty big one.

10:06It comes in the form of inference cost. I figured. You're basically running two models at once. You are. For every single token the model generates, that little purse RM has to score all the top candidates. And that's a lot of math. The paper is up front about it. Inference time goes up by two to three times. Whoa. So a response that might take five seconds now takes 10 or 15 seconds. You notice that as a user. Yeah. That breaks the flow of a chat. It does. Latency is a real problem. But maybe the bigger issue is memory. In their setup, they needed about 17 gigabytes of extra GPU memory. 17 gigs.

10:40Just to hold the reward model and all the feature data. Wait a second. So the paper talks about this democratizing AI because you don't need to spend a fortune on fine-tuning. But if I need a 24-gig VRAM card, like a top-of-the-line RTX 4090, just to run the thing, how is that accessible? You're right to push back on that. I think when they say accessible, they mean accessible compared to the cost of training. Oh. Fine-tuning a big model like Llama 3 takes a whole cluster of H100s. It can cost tens of thousands of dollars. Compared to that, a single high-end consumer GPU is, you know, relatively accessible.

11:15So it's accessible for a small company or a university lab, maybe? But not for me running it on my phone. Not yet. But the good news is, because the reward model is decoupled, you can optimize it separately. You could probably distill it down to something much smaller. That's a good point. And there's another benefit to the decoupling, right? It's model agnostic. The plug-and-play aspect, yeah. They showed you can train one reward model and then just attach it to Llama 3 or Mistral or Gemma. It doesn't really care. Because the nutrition label for a funny sentence is the same, no matter who wrote it.

11:46Exactly. The features of language are universal. And this opens up a really cool idea. It's like you could have your own personal preference file, a portable vector that's just yours. I really like that idea. It's like in the future, I'll have my profile on a little file and just says likes succinct answers, hates puns, wants technical details, prefers a polite tone. And you just plug that into your car, your phone, your work computer. The base models might all be different, but the alignment is always yours because that preference vector is standardized. It totally solves the cool start problem.

12:19You don't have to spend five minutes prompt engineering every new AI you meet. You just introduce yourself once with math. It's a very clean way to separate the intelligence, which is the base model, from the values, which is your preference vector. But I feel like this leads to a bit of a weird place. Doesn't this kind of extreme flexibility create a paradox? It does. We think of personality as something stable, right? When I talk to you, I know you have a certain style, certain beliefs, a certain way of seeing the world. Right. If I asked you to write a poem celebrating pollution, you'd probably push back.

12:55You'd argue with me. But a paid-aligned model is a total chameleon. It has no core self. If you slide the vector over to agreeable, it agrees with you. Slide it to combative, and it'll fight you. Slide it to expert, it'll drown you in jargon. So we're not really talking to an entity anymore. We're talking to a mirror. A mirror that we can adjust to show us the exact reflection we want to see. And if we can perfectly align an AI to our specific momentary whims, don't we risk creating the ultimate echo chamber? If I only want to hear news delivered in a sarcastic right-wing tone or a sympathetic left-wing tone, Patey could generate that instantly from the same set of facts.

13:34It removes all the friction from encountering a perspective or even just a tone that you don't already agree with. That is both incredibly convenient and kind of terrifying. We started this deep dive trying to escape the boring beige AI, and we found a way. But the price might be that we just end up talking to ourselves. A perfectly customized reality generated for us at decoding time. On that slightly chilling thought, we'll wrap up. The paper is PAD, Personalized Alignment of LLMs at Decoding Time. It's a brilliant piece of work that really does solve the average user problem as long as you have the VRAM.

14:09And as long as you know yourself well enough to set the weights. Thanks for listening. We'll see you on the next Deep Dive. Goodbye.

From the publisher

This paper introduces **Personalized Alignment at Decoding-time (PAD)**, a framework designed to tailor Large Language Model (LLM) outputs to specific user preferences without the need for expensive retraining. Traditional alignment methods often rely on a "one-size-fits-all" approach, but **PAD** uses a unique **personalized reward model (PersRM)** to adjust token-level predictions during the inference phase. By **decoupling text generation from user values**, the system can adapt to diverse cultural, educational, or political leanings in real-time. Experimental results show that **PAD** excels at generalizing to **unseen preferences** and works effectively across various base models. Ultimately, the authors provide a **training-free solution** that balances high-quality personalization with computational efficiency.

More from Best AI papers explained

All 475 episodes
PAD: Personalized Alignment of LLMs at Decoding-TimeBest AI papers explained · 14 min
Listen in VO