In short
The episode argues that today’s LLM coding assistants fail when they generate overly long, low-confidence code, and presents “Empower” (Logit Threshold Empowerment) as a scalable way to train agents to maximize human control rather than guess user goals.
Guest backgrounds
No guests are mentioned; it’s a solo “Deep Dive” episode.
Key claims
Empower trains agents to stop when the model’s confidence drops (entropy/negative log-likelihood above threshold eta), completing predictable boilerplate but handing off before high-impact, uncertain decisions. It avoids costly human accept/reject labeling and avoids constant “are you sure?” interruptions.
Notable examples
Copilot-like 150-line “Kafka pipeline” suggestions from a small typo; empowering completions like imports, helper functions, and loop boilerplate; disempowering giant rewrites. Results: +192% pass@1 vs SFT on LiveCodeBench simulations; better DPR; 18-person double-blind study preferred Empower 78%, with 31% higher acceptance, 38% fewer suggestions, and 26% fewer deletions (12.9 vs 9.56 chars).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOChallenges with AI Coding Assistants
0:45 to 1:40
Discussing common issues users face with AI coding assistants like Copilot.
“It feels like progress for a second because the suggestion is so huge, but then you pay for it, spending ages fixing that one little assumption I got wrong.”
Understanding Empowerment in AI
1:40 to 3:40
Introducing the concept of empowerment in AI agents and its significance.
“really dig into this potential solution called the empower method.”
Empowerment vs. Traditional Alignment Methods
3:40 to 6:40
Contrasting the empowerment approach with standard alignment methods like RLHF.
“those are the branch points that really define your specific program.”
The Empower Algorithm in Action
6:40 to 10:30
Explaining how the Empower algorithm utilizes confidence to enhance user control.
“So a point where the confidence drops, that's the branch point.”
Testing the Empower Method
10:30 to 13:00
Discussing tests and results of the Empower approach versus traditional methods.
“And not just accepted, but requiring less fixing.”
Future Implications of Empowerment
13:00 to 13:38
Exploring the broader implications of the empowerment philosophy in AI development.
“A really interesting question to ponder.”
Transcript
Automatic transcript. May contain errors.0:00Welcome to the Deep Dive, your shortcut to being well informed. Today we're tackling something I think pretty much everyone who uses AI assistance has run into, especially for coding. You know the feeling, right? You're deep in concentration, coding with something like Copilot. And it's great for a bit, completes a dock string, maybe adds some imports, writes some boilerplate. But then maybe you use a variable name it didn't expect or add a comment in the wrong place, just a tiny thing. And suddenly, bam, the Assistant just goes completely off the rails. It spits out this massive, like, 150-line block of code, thinking you need a whole Kafka pipeline when you just wanted a simple file read.
0:38Happened to me recently, one typo in a comment, and it tried to rewrite an entire class. Seriously, I spent 15 minutes debugging the Assistant. Oh, yeah, that's the classic scenario. It feels like progress for a second because the suggestion is so huge, but then you pay for it, spending ages fixing that one little assumption I got wrong. And the core issue is often how these things are trained. They're pushed to be helpful, to complete tasks, even when they're not totally sure about your actual long-term goal. Right. So the way we do it now feels kind of broken. The options seem to be either crazy expensive, like getting people to label why they rejected a huge suggestion.
1:15Yeah. Imagine labeling line three was wrong in this hundred line block. Yeah. Not practical. Or the other option, the agent just constantly butts in asking, are you sure about this? did you mean that? Which totally kills your flow. And that just defeats the whole point of having an assistant in the first place. The real challenge is getting the agent to help as much as possible, but also knowing exactly when to just stop, to hand control back. So our mission today is to really dig into this potential solution called the empower method. It's a pretty clever, scalable approach. It trains the AI agent, not based on guessing what you prefer or what reward it thinks you want, but specifically on maximizing your ability to influence the outcome, your control.
1:57That sounds like a really different way of thinking about alignment. So let's break that down. Maximizing empowerment, what does that actually look like in practice? Well, intuitively, you know, empowerment is about an agent's ability to make changes happen in its environment. But here, with a human and an AI working together, the goal flips. The LLM isn't maximizing its own power. It's trained to maximize the human's power, your ability to shape what happens next. And that's the philosophical shift you mentioned earlier. Think about standard alignment, like RLHF reinforcement, learning from human feedback.
2:33Yeah. That's everywhere now. It basically tries to guess the human's reward function. It tries to figure out the destination, the end goal you're aiming for. Exactly. And power is different. It doesn't actually need to know your ultimate goal. It doesn't care if you're building a game or a data pipeline. It just focuses on making sure you can get to whatever goal you have more effectively by ensuring that your next action, the next thing you type, the next design choice you make has the biggest possible impact on the final result. Okay. Let's make that concrete. In code, what's an example of an empowering action from the assistant versus one that's, well, disempowering, like that giant Kafka suggestion?
3:08Right. So think about the boring, repetitive stuff. An empowering assistant would say, implement standard helper functions, set up common variable declarations, maybe write the usual library imports, finish off particular boilerplate code. The stuff anyone would write basically the same way. Precisely. The things that are obvious in general. By handling that predictable grunt work, the assistant clears the way. It removes the noise so you can focus your attention on the important stuff. The strategic decisions, the creative bits, the points where the code could go in several different directions, those are the branch points that really define your specific program.
3:44And the disempowering action is that huge suggestion. Yeah, that 150-line monster. Even if maybe 90 % of it is okay, you still have to read and check all 150 lines. Then you have to carefully fix the bad 10%. That just eats up your mental energy on verification instead of letting you do the high-impact coding. Okay, so it really boils down to timing. The assistant needs to get out of the way right when a human decision is needed. Like if I type something obvious, say for item in my list, my next few lines are pretty predictable. That's what you'd call a state of low empowerment for me. Exactly.
4:18Yeah. Because you're just following a standard pattern. The assistant should just complete that loop structure for you. And by doing that. By completing the predictable part, the assistant instantly brings you to the point where your next keystroke matters a lot more. Maybe you're about to type a unique function call inside the loop or define some tricky logic. That's the state of high empowerment. Your action has a big impact. Okay, now the techie bit. How does it know when that critical moment arrives? The paper says it's self-supervised, right? It just uses existing code. No explicit accept or reject feedback.
4:51Yes, and that's the really elegant part of the Empower algorithm, specifically something called Logit Threshold Empowerment. Because the assistant is trained on tons of human-written code, it already has a built-in sense of what's predictable and what's not. It uses its own confidence. Basically, yes. It uses its internal likelihood model, how probable it thinks the next sequence of code is, based on everything it's learned. So imagine the LLM looking ahead. If the path is straight, like finishing standard setup code, its confidence is super high. It's almost certain what a human would type next.
5:26Okay. High confidence means low risk for the AI. Yeah. But also low impact for me, the human, because it's just routine stuff. Exactly right. So the LLM calculates this probability that the cumulative likelihood for potential code completions of different lengths. Not just the next word, but a whole chunk. Right. And it picks the longest chunk of code that stays above a certain confidence threshold. But the instant the path forward becomes uncertain, when there are multiple plausible ways the code could go, maybe a complex decision point, the LLN's calculated likelihood for any single path starts to drop sharply.
6:01And that drop is the signal. That's the signal. Technically, it's looking at when its estimate of entropy, or specifically the negative log likelihood, goes above a threshold, which they call eta. Hold on, though. What if it gets this wrong? What if it's too cautious? It stops too early, and I end up typing boilerplate that it should have handled. just because it hits some arbitrary math threshold. That sounds annoying too. That's the balancing act they're aiming for. By using that cumulative likelihood, looking at the whole proposed chunk, it tries to make sure the suggestion is long enough to actually be helpful.
6:34Yeah. But it's designed to stop right when its certainty dips below that preset level. Uh-huh. So a point where the confidence drops, that's the branch point. That's the moment where my next action as the human really changes where the program goes. So it does the obvious stuff, then stops right before the unpredictable high-impact moment. You got it. You make sure the very first thing you type after its suggestion is likely to be a meaningful, important token. Okay, that makes sense. So does it work? They tested this first in simulations, right? They did. They wanted to see if this whole philosophy actually led to better results.
7:07They used simulated human programmers, actually other large language models, working on coding problems from a benchmark called LiveCodeBench. And how did this empower approach stack up against, say, standard training? What is standard training here, just for clarity? Good question. Standard here usually means supervised fine-tuning or SFT. That's where you basically just train the model on lots of correct code examples, teaching it to predict the next token, but without this specific focus on empowerment or careful stopping. Got it. So empower versus SFT. The results were pretty dramatic. For the simulated human programmer, using an Empower trained assistant increased the success rate, like actually solving the coding problem, measured by pass at one by an average of 192 % compared to using the SFT baseline assistant.
7:56Wow. Okay, nearly triple the success rate. That's huge. But like we said at the start, just getting the right answer isn't everything if the assistant drove you crazy getting there. Exactly. Raw success isn't the best measure of assistance. Which is why they introduced that other metric, right? The discounted pass rate, DPR. Yes. DPR is key here. It's designed to measure good assistance. It understands that success is great, but only if it doesn't come with a huge cognitive cost for the human. You want to solve the problem efficiently. So what does DPR actually discount? What's the penalty? It basically starts with the success rate, the pass at one, but then it subtracts points for every extra token the human had to read or verify or maybe delete from a long, confusing suggestion.
8:38It's trying to capture the total effort or burden the assistant imposes. So an assistant making sure perfect suggestions would score way higher than one making long suggestions that are mostly right but need fixing. Precisely. It measures the usefulness relative to the effort. And the empower method came out on top here too. It beat all the baseline models on DPR. Okay, so it wasn't just more successful, it was also less annoying, less burdensome. It seemed to hit that sweet spot, helpful automation, but keeping the human in control. But simulations are one thing. The real acid test is actual humans, right?
9:12Did real programmers like this approach or did they find it, I don't know, less helpful because the suggestions were shorter? No, they definitely preferred it. They ran an 18-person study, double-blinded. They had programmers use either the Empower Assistant built on LAMA 3.18b or a strong baseline that also made suggestions but didn't use the empowerment logic. They called it base 20. And the preference numbers? Pretty clear cut. Participants preferred the Empower Assistant 78 % of the time. And crucially, the question was which one they'd most enjoy using in practice. So that strong preference is a direct thumbs up for the less is more low burden approach.
9:50Okay, that's compelling. Did the actual usage data show that it was being more selective, more judicious, like the theory predicted? Absolutely. Get this. The Empower Assistant actually had a higher acceptance rate for its suggestions, 31 % higher than the baseline. People liked what it suggested more often. Okay. But it generated 38 % fewer suggestions overall, something like 208 suggestions per user compared to 333 for the baseline. Fewer suggestions, but higher quality ones. Exactly. It was clearly picking its moments better. Less noise, more signal. It wasn't just constantly throwing suggestions out there.
10:23So that answers the potential worry. Maybe shorter suggestions mean less help overall. The answer seems to be no, because the ones it did make were more likely to be accepted, more useful. And not just accepted, but requiring less fixing. When people did accept an empower suggestion, they ended up deleting 26 % fewer characters from it compared to the baseline suggestions. Let's quantify that. So with the baseline, if I accept it, I have to fiddle with it, delete almost 13 characters on average. Yeah, about 12.9 characters deleted on average from accepted baseline suggestions. But with Empower, it's closer to 9.5 characters.
10:59So it's 9.56 characters deleted. That's a noticeable reduction in that annoying fix the AIs mistake time we hate. And I assume the suggestions themselves were shorter overall, reflecting that philosophy. Significantly shorter. The baseline suggestions averaged about 82 characters trying to do a lot. Empower suggestions averaged just under 44 characters. High utility, low burden. Complete the obvious, then stop. Okay, so let's try and wrap this up. The core idea here with empower. Yeah, the big takeaway is that empowerment seems to offer this practical, scalable way to get AI agents that are genuinely helpful and aligned.
11:37And crucially, it does this using just offline data existing code through a self-supervised goal. It avoids the huge costs and headaches of methods that need constant human feedback or try to guess complex intentions. It feels like it reframes the whole alignment question, doesn't it? Instead of asking, how does the AI figure out my secret goal? Right. It asks, how does the AI make sure I stay in the driver's seat and can achieve my goal most effectively? And while they showed this works well for code, you can easily see how it might apply elsewhere. Think about, say, AI writing assistants. Yeah, instead of the AI trying to write your whole next paragraph, which you then have to rewrite anyway because it doesn't sound like you.
12:14Yeah, exactly. An empowered writing tool might finish your sentence, maybe handle a standard transition phrase, but then it would start right before you need to make your unique point or add your personal anecdote. The point where your input has the highest value. Precisely. It's about maximizing the human's effectiveness and control. Yeah. Not just the AI's output. And this could have really big implications for how AI gets developed, right? I mean, the big foundation models, the base LLM, they're already trained using a self-supervised goal. Just predict the next token. Yeah. This research suggests that maybe even the alignment part, the fine tuning, could also be done with a self-supervised objective like empowerment.
12:54You potentially bypass a lot of the complexity and, frankly, the debates around human preference labeling. It does open up that possibility, which leads to a pretty interesting final thought for everyone thinking about how we work with AI in the future. If we can tackle this core challenge of helpful assistance, making AI useful without making it annoying or taking over by thinking about human control instead of trying to guess human reward, what other tricky AI alignment problems might we be able to solve or at least approach differently by focusing on maximizing human agency and capacity? A really interesting question to ponder.
13:29Well, thank you for joining us for this deep dive into self-supervised empowerment. Hopefully this gives you a little more, well, empowerment next time your AI assistant gets a bit too ambitious.
From the publisher
This rsearch paper by Sergey Levine's group introduces a self-supervised method for fine-tuning Large Language Model (LLM) agents to be more effective and aligned assistants, particularly in code generation. The core idea is to train agents to maximize the human's empowerment, defined as the user's ability to effect desired changes in the environment, rather than relying on costly explicit human feedback or inferred rewards. The paper details the mathematical connection between their Logit Threshold Empowerment algorithm and the concept of effective empowerment, focusing on training the assistant to complete predictable, boilerplate text so the human can concentrate on key decisions. Experimental results, including a simulated evaluation and an 18-person double-blind user study, demonstrate that the Empower assistant is preferred by users, achieves a higher acceptance rate, and significantly increases the simulated success rate for programmers compared to strong baselines. The authors conclude that this offline data-only approach provides a scalable framework for training useful AI assistants.




