In short
Learning LLM personalization from user edits, reframing “frustrated rewriting” as training signal via edit cost and a late ensemble of SFT, DPO, and cost-based RL.
Guests/backgrounds
No guest names or bios in the transcript; it’s a host-style discussion between two speakers.
Key claims
User edits are “free” high-stakes feedback; edit distance (“edit cost”) measures effort and guides training; the same edit can be interpreted as (1) supervised fine-tuning (copy), (2) preference learning/DPO (directional “better than”), or (3) RL minimizing predicted edit cost. SFT is fast but brittle for “weak users” (satisficers); DPO is safer. Late ensemble uses a bandit to pick the best model at test time. Preferences transfer from smaller to larger simulated users.
Notable examples
Fixing a client email tone; SFT overfitting to clunky writing; email writing and summarization experiments; simulated Llama 3 8B-trained preferences tested on Llama 3 70B.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOTurning Frustration into Valuable Data
0:45 to 2:14
Discussing how user edits are a gold mine for AI improvement.
“They are maybe the most valuable resource in AI right now.”
Challenges of User Edits in AI
2:14 to 3:40
Examining how AI interprets user edits to improve its responses.
“I don't write a little note saying, please be 10 % less formal.”
Three Perspectives on User Edits
3:40 to 5:07
Understanding the three methods AI can use to learn from edits.
“I'm guessing that's because, well, not everyone is a great writer.”
Evaluating User Types and Learning Methods
5:07 to 7:27
Distinguishing between strong and weak users in AI editing.
“Even if you don't end up in the perfect spot, DPO just knows that north is better than where you started.”
The Late Ensemble Approach
7:27 to 9:55
Introducing the late ensemble method for AI model training.
“It moves the model in the right direction without getting locked into your specific imperfections.”
Results of the Late Ensemble Method
9:55 to 11:18
Reviewing the effectiveness of the late ensemble in real-world tasks.
“By letting the models compete, it gave every type of user the best experience.”
The Paradox of Perfect AI
11:18 to 13:19
Discussing the implications of achieving zero edit cost in AI.
“I feel a little better about all my backspacing now.”
Transcript
Automatic transcript. May contain errors.0:00You know, there's this very specific modern frustration that I don't think we have a proper name for yet. Oh, I think I know where you're going with this. It's that moment you ask an AI to write an email, let's say a tricky note to a client, and it spits out something that just sounds like a robot. Totally. It's either way too formal or it uses the word delve three times, or it just completely misses the point. Exactly. And then you spend the next five minutes rewriting it. You're backspacing, you're fixing the tone, and by the time you hit send, you're thinking, I basically just wrote this myself.
0:34It feels like such a waste of time. That's the universal feeling. It is. But the research we're diving into today completely flips that on its head. It suggests that those five minutes of, you know, pure annoyance, they're not waste at all. No, they're actually a gold mine. They are maybe the most valuable resource in AI right now. So this is learning from user edits. That's the one. And it's really the frontier of how we get from these sort of general smart models to truly personalized ones. Our mission today is to unpack a framework that takes all your frustrated typing and turns it into a mathematical signal.
1:10Okay, so you're going to vindicate my frustration. I like that. We're going to put it to work. We're going to look at how that data solves the biggest bottleneck in AI development. I'm in. Let's set the stage, though. When we talk about training these huge models, we always hear about RLHF. Right. Reinforcement learning from human feedback. That's the industry standard. And how does that work again? Well, think about the logistics. A company like Google or OpenAI, they have to hire armies of people, human annotators. Okay. And they tell them, hey, pick the better of these two essays or write a perfect response to this prompt.
1:44It sounds so artificial. Those people are just getting paid by the hour. They don't actually care if my email to the client works. Precisely. It's low stakes data. But when you are editing an email to your boss, the stakes are very high. You care. I care a lot. And that makes the data you generate your edits incredibly high quality. And crucially, it's free. It's just a byproduct of you using the tool. Okay, so we've got all this data exhaust from people like me fixing the AI's mistakes. How do we actually use it? Because my edits aren't clean. I don't write a little note saying, please be 10 % less formal.
2:18No, you just delete the slang and type in something else. Yeah. And that's the core challenge. The machine has to guess your intent from your action. The framework we're looking at proposes we measure something called edit cost. Edit cost? Is that just a fancy way of saying how many keys I had to press? Essentially, yeah. It's the edit distance. If you change nothing, the cost is zero. The AI nailed it. And if I rewrite the whole thing. The cost is super high. And the goal of this whole field is to get that cost as close to zero as possible. That seems obvious. But here's where I get a little skeptical.
2:51If I fix the email, why doesn't the AI just, you know, copy my fix, say, okay, do that next time? Why the whole framework? Because just copying is actually kind of dangerous. And this is where it gets really interesting. The research argues that a single edit isn't one signal. It's a medley. A medley, like three different instruments playing the same song. Exactly. And depending on which instrument you listen to, the AI learns in a totally different way. So let's break them down. Okay. The first perspective is what you just said. Supervised fine-tuning, or SFT. The copycat approach. You got it.
3:26The logic is, the user wrote this, so this must be the correct answer. It trains the model to generate exactly what you wrote. Which sounds totally logical. I mean, I wrote it, so that's what I want. It is logical. And it's fast. It's just imitation. The user changed hey to dear sir. So from now on, I'll use dear sir. But you said copying is dangerous. I'm guessing that's because, well, not everyone is a great writer. Exactly. SFT assumes you're what they call a strong user. It assumes you're right. But what if you're not? What if you just made it different, but not better? Or left in a typo? Then the model learns to make my typos.
4:02Ouch! Right. But hang on to that thought, because the strong versus weak user thing is the whole pivot point. Let's get to the second perspective first. Preference learning. Usually done with DPO. DPO. Direct preference optimization. We've touched on this. That's the A is better than method? Correct. So in this case, the model looks at the original AI draft in your edited draft, but the logic isn't copy the new one perfectly. It's much simpler. It's just the new one is better than the old one. That feels like a subtle difference. If it's better, why not just copy it? It's a huge difference mathematically.
4:35It relies on a theoretical proof called the balance equation. Okay. You can't just drop balance equation on us. What does that mean in plain English? Um, it basically proves that your edits, even if they're not perfect, they almost always move things in the right direction towards the optimal version. It justifies using preference learning because it treats your edit as a directional hint, not a final destination. Oh, I like that. So SFT says this is the destination, but preference learning just says you're heading north. That is a brilliant analogy. Yes. Even if you don't end up in the perfect spot, DPO just knows that north is better than where you started.
5:13It's safer. Okay, so we have the copycat, which is SFT, and the navigator, which is preference. What's the third part of this medley? The third is reinforcement learning, but focused on cost. The edit cost from the beginning. Right. So here, the system tries to build a cost model. It's trying to predict, if I give the user this sentence, how much work are they going to have to do to fix it? So it's trying to predict my level of annoyance. In a way, yeah. Yeah. It predicts the negative reward of your editing effort. Then it tries to generate text that it thinks will have the lowest predicted cost.
5:47Okay, so we have the copycat, the navigator, and I don't know, the effort miser trying to save me work. That works. And these are the three ways to interpret the exact same edit. Now for the million dollar question, which one is best? My gut still says the copycat SFT. I mean, if I took the time to fix it, just use my version. It's the fastest path to what I want. And you'd be right. If you were a perfectionist, The research makes this key distinction between two user types, the strong user and the weak user. I have a feeling I know which one I am, but go on. The strong user is meticulous. They edit the text until it perfectly matches their hidden goal.
6:25You know, be professional but warm. For them, SFT is fantastic. The data is perfect, so copying is the right move. And the weak user? The weak user is what they call a satisficer. A satisficer. That's a mix of satisfy and suffice, right? Exactly. A weak user fixes the really obvious mistakes, the wrong facts, the bad grammar. But they might leave a sentence that's just a bit clunky because, well, it's good enough. Yeah, that's me. 100%. I'm not gunning for a Pulitzer with a Tuesday morning email. And that is the trap of SFT. If you train the model on a weak user's edits, it starts to think your response is the gold standard.
7:02It overfits to mediocrity. It learns that clunky is the goal. So it creates an echo chamber of my own just okay writing. It really does. But this is where the navigator, the preference learning, DPO, really shines. It's much more robust for weak users. Because it's only looking at the direction of improvement. Yes. Even if your edit is still a bit clunky, it's almost certainly better than the AI's first draft. DPO just learns that directional gradient. It moves the model in the right direction without getting locked into your specific imperfections. That is such a critical insight. So, SFT is fast but brittle.
7:38DPO is slower but safer. But here's the real problem, right? If I'm deploying this AI, I have no idea who is on the other end. Is it Shakespeare or is it me? And that's the dilemma of the online phase. In a real-world deployment, you are flying completely blind. So, what do you do? Just pick the safe one, DPO, and call it a day? You could, but then you'd be slowing things down for your strong users. You'd miss out on that speed and precision SFT offers them. The solution they propose is really clever. It's called the late ensemble. Late ensemble. Sounds like a jazz band that only plays after midnight.
8:10It's surprisingly close to a sports manager subbing players in and out during a game. Instead of training one master model, you train separate ones. You have your SFT model, your DPO model, your RL model. Okay, so they're all on the bench. They're not blended together. Exactly. They're separate. And then when you're actually using the tool, the system uses a bandit algorithm. A bandit, like a slot machine, a one-armed bandit. That's the one, the multi-armed bandit problem. Imagine you've got five different slot machines. You want to find the one that pays out the most, but you also want to win money while you're figuring it out.
8:44So you have to explore the different machines, but also exploit the one you think is winning. Precisely. And in this case, the slot machines are the different models, SFT, DPO, and the payout is a low edit cost. So the system is watching me. It gives me a draft from the SFT model. If I barely touch it, it's like jackpot. This user works well with SFT. Exactly. It boosts the confidence score for that model. But if you start rewriting everything, the manager sees that edit cost spike and thinks, OK, SFT is failing here. So for my next email, it swaps in the DPO player from the bench. Yes. It adapts in real time at test time.
9:20It doesn't have to stop and retrain the whole giant network. It just picks the best player for the user in that moment. That is so smart. But did it actually work? The results were really compelling. They tested it on email writing and summarization, trying to see if it could learn those hidden preferences like be brief or be respectful. Well, the base model was terrible. No surprise there. High edit costs all around. SFT was great for strong users, but awful for weak users, just like we thought. And the late ensemble. It was the clear winner. It consistently minimized the total amount of editing the user had to do over time.
9:56By letting the models compete, it gave every type of user the best experience. It's basically an admission that there is no one single perfect algorithm. Right, so just keep a whole utility belt of them ready to go. That's amazing. And there was one other detail in the results that I think is huge about transfer learning. They trained the models on edits from a simulated user that was, let's say, moderately smart, like a Llama 3 8B model. But then they tested it on a much smarter user simulated by a Llama 370B. So the student was trained by a junior teacher but had to perform for a tenured professor.
10:31That's a great way to put it. And it worked. The preferences transferred. The model learned the concepts from the smaller model and was still able to satisfy the smarter user. That's massive for scalability. It means you don't need the world's best experts to bootstrap these things. It really validates the whole approach. It's not just a hack. It's a genuine way to pull intelligence out of interaction. Wow. Okay, so to bring it all back home, we started with my frustration about rewriting AI emails. And we ended up with a system that watches you rewrite, decides if you're a perfectionist or just trying to get by, and then swaps out its own brain to match your style.
11:05It really reframes the whole thing. We're moving from hiring people to label data to just watching people work. And treating that work as facts, hints, and scores all at the same time. That's the magic. I feel a little better about all my backspacing now. I'm not just fixing a typo. I'm training an algorithm. You're pretty ensemble. So what's the big takeaway for someone listening? Why should they care about banded algorithms and edit costs? Because this is how AI stops being a generic tool and starts feeling like your tool. Yeah. Right now, most AI has one voice, right? That sort of cheerful, lightly robotic customer service voice.
11:47Yeah. To be useful for real deep work, it has to adapt to you. But it can ask you 100 questions. Do you prefer the Oxford comma? Should I end with best or sincerely? Oh, I'd close the tab immediately. No one has time for that. Exactly. It has to learn by watching you work. This shows a path for an AI to adapt to your unique style without you ever having to explicitly explain yourself. That's the dream, isn't it? An AI that just gets you. Yeah, it is. But it does raise a funny little paradox. Oh, I knew there'd be a catch. There's always a catch. Well, think about the ultimate goal. The goal is to get the edit cost down to zero.
12:21Right. Perfection. The AI writes the email, I look at it, I nod, and I hit send. Zero work for me. But if the edit cost hits zero, the feedback loop closes. Wait, what? If the AI is perfect, you stop editing it. And if you stop editing, it stops getting data. It stops learning. So perfection kills the learning process. In this framework, yes. If we stop correcting the machine, how does it ever evolve with us? How does it learn new slang or new company policies or that my job changed if I never have to touch the text it generates? That is bizarre. We are training it to reach a state where it can no longer be trained.
13:00It's like a self-destruct mechanism for its own education. We might get to a point where we have to make mistakes on purpose just to keep it sharp. That's a terrifying thought. An AI that deliberately messes up your email just to see if you're still paying attention. Hey, I put a typo in the subject line. Just a little test. Hope you caught it. Let's hope they work that part out. Fingers crossed. Well, next time you're furiously deleting a paragraph some AI wrote, just remember, you're not wasting time. You're generating gold standard data. Use your power wisely. Thanks for diving deep with us.
13:32We'll see you on the next one.
From the publisher
This paper introduces a framework for fine-tuning large language models (LLMs) by leveraging user edits found in deployment logs, such as those from writing or coding assistants. Unlike traditional methods that rely on expensive manual labeling, this approach treats user modifications as a rich, multi-dimensional source of preferences, supervision, and cost feedback. The researchers provide a theoretical analysis of how these different feedback types impact model performance and demonstrate that individual algorithms have distinct trade-offs. To address these variations, they propose ensemble procedures that combine multiple learning methods to achieve more robust results. Empirical evaluations on summarization and email writing tasks show that ensembling consistently outperforms single-method approaches. Finally, the study highlights how these techniques can effectively personalize LLMs or adapt them to broader user distributions during real-time interaction.




