Demystifying the unreasonable effectiveness of online alignment methods

21 Apr 2026 · 18 min · 9 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode explains why online AI alignment methods (online RLHF and online DPO) work better in practice than theory predicts, arguing the mismatch comes from a flawed evaluation metric rather than the algorithms themselves.

Guests

The transcript names only one researcher, Enoch Hunwuk Kang, who authored the “brand new” research discussed (dated April 18, 2026). No other guest identities are provided.

Key claims

Both RLHF and DPO rely on purely greedy updates plus exploratory randomization. The commonly used KL-regularized regret metric conflates learning cost with intentional exploration, making performance look pessimistic. A new “temperature zero regret” criterion evaluates only the deterministic argmax at inference time, yielding constant cumulative regret O(1) instead of O(log T).

Notable examples

A maze analogy for greedy local decisions; and a piano concerto / chef-in-kitchen analogy showing how grading exploratory “noise” can falsely label mastery as failure.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Introduction to Online Alignment Methods

1:05 to 2:17

Discussion on online alignment methods, including RLHF and DPO, and their impact on AI.

“So welcome, everyone, to today's Deep Dive.”

Greedy Updates and Their Implications

2:17 to 4:25

Explanation of purely greedy updates in AI and their short-sighted nature with a maze analogy.

“RLHS uses a reward model based on human preferences to guide the AI, while DPO streamlines that whole thing by optimizing the language model directly from preference data.”

Understanding Regret in AI Learning

4:25 to 6:33

Insights into the regret criterion used in AI alignment and its mathematical implications.

“You get stuck in what they call a local minimum.”

Critique of KL Regularized Regret

6:33 to 10:48

Discussion on the shortcomings of KL regularized regret in evaluating AI models.

“Let's start with the KL regularization piece.”

Proposed Shift to Temperature Zero Regret

10:48 to 13:06

Introduction of the temperature zero regret criterion as a solution for better AI evaluation.

“The core learning mechanism is successfully identifying the optimal path, but its success is buried under the statistical noise of that forced exploration.”

Implications of New Evaluation Metrics

13:06 to 14:00

How the new metric changes the understanding of greedy alignment methods and their performance.

“If we connect this to the bigger picture, that distinction is absolutely paramount.”

Understanding Constant Cumulative Regret in Algorithms

14:00 to 16:34

Learn about the significance of constant cumulative regret and its implications in AI algorithms.

“Instead, they achieve constant cumulative regret, denoted mathematically as O of 1.”

The Impact of Evaluation Metrics on Learning

16:34 to 17:06

Discover how evaluation metrics can misrepresent the capabilities of AI systems and learning processes.

“The synthesis in Kang's research provides such clarity for the entire field of machine learning.”

Reevaluating Complex Systems and AI Behavior

17:06 to 18:05

Explore the importance of proper measurement in understanding AI behavior and uncovering inefficiencies.

“If you grade the practice phase as if it were the final performance, you literally engineer a false diagnosis of failure.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00So, what if I told you that the smartest, most capable AI models operating in the world right now are, well, built on a mathematical foundation that theoretically proves they should be failing. Right, which sounds completely backwards. It really does. I mean, when we look at engineering or like physics, we expect the underlying math to perfectly reflect the physical reality of the bridge we're building. Exactly, or the rocket we're launching. Yeah, but when you step into the world of artificial intelligence and specifically how we align these massive models to actually follow human intent, we're looking at this massive, just glaring contradiction.

0:38A huge paradox, really. Right, because for quite some time now, the theoretical equations have been predicting inefficiency and failure. But the AI models in the real world are doing the exact opposite. I mean, they're functioning brilliantly. They really are. And, you know, it's one of the most compelling paradoxes in modern machine learning. Yeah. We have practitioners achieving just unprecedented success on one side and then theoreticians pointing to these really pessimistic roofs on the other. Which is wild. Yeah. So welcome, everyone, to today's Deep Dive. Our mission today is to explore exactly that bridging the gap between theory and practice.

1:12Right. And we've got some brand new cutting edge research to look at. We do. Published literally yesterday, April 18, 2026, by researcher Enoch Hunwuk Kang. And this is so cool because it sets up a core mystery for you as you listen. Like, why are these highly popular methods used to train AI so unbelievably effective in practice when the math insists they just shouldn't be? Yeah. And to really set the stage for you, we have to look at the specific focus of this research, which is a specific area of machine learning called online alignment methods. Why online alignment? Exactly. And, you know, alignment is just the process of getting an AI to be helpful, harmless and accurate.

1:51The online part simply means the model is learning, adjusting and iterating step by step based on continuous feedback. Instead of just being trained once in a vacuum and like sealed off forever. Precisely. And the industry relies really heavily on a few standard methods to do this. We are talking about online RLHF, that's reinforcement learning from human feedback and online DPO, or direct preference optimization. Which, I mean, those are basically the engines driving the most advanced models available today. They are. RLHS uses a reward model based on human preferences to guide the AI, while DPO streamlines that whole thing by optimizing the language model directly from preference data.

2:32No separate reward model needed. Right. But what Kang's research highlights is this underlying mechanism that both of these methods actually share, right? They rely on something called purely greedy updates. Yeah, purely greedy updates. And here's where it gets really interesting, because the answer to the mystery of why these models work so well isn't about the engineers suddenly building a fundamentally different type of AI. Right. It's not a hardware change. No. Kang's research reveals that the solution is actually about changing how we measure the AI's behavior. But before we get to the new measurement, I feel like we have to understand the old one.

3:06We have to understand why purely greedy updates were causing such a panic among the math folks in the first place. Well, a purely greedy update, just in algorithmic terms, is a strategy that makes the optimal choice at any given immediate step. It does this without considering the overarching, like, long-term global structure of the problem. It's just optimizing for the immediate reward right in front of it. So it's very short-sighted. Yeah, exactly. In the context of online alignment, the algorithm updates the model's parameters to maximize the immediate preference score it gets on a specific batch of data.

3:40It's not calculating some complex, long-term trajectory for the whole learning process. Okay, let's visualize that for a second. If we think about navigating a massive pitch black maze, right? Sure. A greedy algorithm isn't trying to draw a map of the entire labyrinth. It just like stands at an intersection, looks for the single path that seems the brightest or most rewarding right in front of it and takes it. Right. It immediately turns toward that short term payoff. But wait, I want to push back on this architecture a bit because intuitively, isn't a greedy algorithm highly vulnerable to getting trapped?

4:17It can be, yeah. Like, if you only take the immediate reward, you often end up walking into a corner that looked well lit but has no exit. You get stuck in what they call a local minimum. So why aren't these AI models constantly hitting dead ends during their alignment? Well, what's fascinating here is that the maze analogy, while it's super helpful for visualizing immediate optimization, it kind of falls short of capturing the multidimensional architecture of a language model's training environment. Oh, really? How so? Because in online alignment, the AI is not just walking down a single narrow physical hallway.

4:52The training policy inherently relies on exploratory randomization. Exploratory randomization. Yeah. We force the model to continuously sample different distributions of text to try slightly different words, vary its sentence structures, and test alternative reasoning paths. Oh, I see. So it's kind of jiggling the handle constantly. Exactly. Exactly. So even if a greedy update tries to pull the model's parameters into a rigid dead end, that forced variance constantly pushes it back out, making it test adjacent mathematical possibilities. The exploration acts as a buffer against those local minima.

5:25Wow. Okay. That completely changes the framing. The exploration is literally saving the greedy algorithm from its own short-sightedness. It is. But wait, if empirical results show that this greedy approach works like that the AI navigates this complex landscape of human alignment without getting hopelessly stuck, what exactly is the theoretical math measuring that makes it look so disastrous on paper? Well, the theoreticians have been relying on a metric known as the regret criterion. Regret, right. Yeah, regret is a foundational concept in machine learning. It basically measures the difference between the absolute best possible sequence of decisions the model could have made and the sequence of decisions it actually made during the learning process.

6:05So if the AI deviates from the perfect path, it accumulates regret. Precisely. And specifically, the theoretical guarantees for these greedy alignment methods have been calculated using something called O of log T KL regularized regret. Okay, let's unpack this, because we need to break down that specific mathematical ruler, since O of log T KL regularized regret seems to be the primary culprit in this entire misunderstanding. It absolutely is. Let's start with the KL regularization piece. We know that KL divergence, Kolbeck-Leibler divergence, is essentially a way to measure how much one probability distribution differs from another, right?

6:43Yeah, and in the context of alignment, KL regularization acts as a necessary anchor. When we're training a model to align with human preferences, we want it to adapt, sure, but we do not want it to completely forget its foundational training. Oh, right. Like we don't want it to lose its grasp of general syntax or basic facts just because it's chasing a high reward score on a specific prompt. Exactly. So the mathematical framework applies a KL penalty. It penalizes the model if its new aligned policy deviates too aggressively from its original base reference model. Got it. It's like a gravitational pull.

7:16The model is encouraged to explore outward to find the best possible helpful responses. But the further its probability distribution drifts from the base model, the heavier that mathematical penalty becomes. Yes. It forces the model to maintain a softened training policy. So it can't just collapse into giving one single rigid answer. It has to maintain a spectrum of probabilities. That is the crucial mechanism. The KL regularization forces the model to remain somewhat randomized, to maintain a spread of possible outputs rather than becoming purely deterministic. Okay, so where does the O of log t come in?

7:53Right, the O of log t part of the metric. Big O notation describes how an algorithm's performance scales over time. An O of log t bound means that as the number of training steps t increases, the accumulated regret continues to grow logarithmically. Oh, I see. It suggests that the model will basically never truly start making errors. It will endlessly accumulate a deficit compared to a theoretically perfect model. Man, which paints an incredibly pessimistic picture. You have engineers looking at this beautifully aligned language model that writes eloquent code and answers complex questions flawlessly, while the theoreticians are pointing at this O of log T curve, arguing that the system is fundamentally inefficient and bleeding regret with every single step.

8:37Right. And the contradiction stems directly from what the KL regularized regret metric is actually evaluating. Kang's research points out a fatal flaw in this specific theoretical ruler. It argues that KL regularized regret conflates or inextricably mixes together two entirely different phenomena. What are they? It mixes the actual statistical cost of the model attempting to learn the optimal answer with the intentional exploratory randomization that is being forced upon it by that very KL penalty we just talked about. Oh, oh, this is the absolute crux of the issue. The metric is punishing the AI for doing the exact thing it was programmed to do.

9:15Precisely. Think about this in a real world learning environment, right? For you listening, imagine you're learning to play a complex concerto on the piano. Mastering the piece requires exploratory variance. Oh, absolutely. You have to intentionally test different tempos, experiment with alternate fingerings, push the boundaries of your current technique just to discover the optimal way to play the passage. You are intentionally introducing variance into your practice. And that variance is a prerequisite for mastery. Without it, you stagnate. Exactly. Now imagine an evaluator comes in to assess your final capability as a pianist, But instead of just listening to your final performance, they're grading rubric averages in the statistical errors of every single experimental fingering you tried during your practice sessions.

10:01Oof, that would be brutal. Right. Every time you intentionally played a dissonant chord just to map out the harmonic space, it goes on your permanent record as a failure of skill. Yeah, your overall regret score would look catastrophic. You would look mathematically incompetent, even if your final performance was totally flawless. And that is what KL regularized regret is doing to these AI models. It's grading the exploratory noise as if it were a fundamental error in logic. It really is a profound misalignment of measurement. By evaluating the softened randomized training policy, the policy that is forced by the KL penalty to maintain variance, the metric completely masks the true underlying efficiency of the greedy update.

10:45So it's hiding how good the model actually is. Exactly. The core learning mechanism is successfully identifying the optimal path, but its success is buried under the statistical noise of that forced exploration. The mathematical ruler simply isn't fit for the task of evaluating the final decisive capability of the system. Wow. Okay, so Cain's research has identified the broken ruler. The Kale regularized metric is penalizing the necessary practice noise. So how does the recent research propose we actually fix this diagnostic tool? I mean, we need a way to measure the AI's true ability to find the correct answer, completely isolated from that forced randomization.

11:21Right. And the solution proposed in the findings requires a shift in perspective. We have to move toward what Kang terms a decision-centric notion of performance. Decision-centric. Okay. Yeah. To achieve this, the research just completely abandons the KL regularized metric and introduces the temperature zero regret criterion. Oh, this is a brilliant pivot, since a lot of you listening are probably familiar with how temperature settings work in language models. Like, you know that dropping the temperature to zero strips away the probabilistic variance and forces the model to become entirely deterministic.

11:55Exactly. So this means the new metric fundamentally changes when and how the AI is evaluated. It does. It separates the training policy from the inference policy. Got it. During training, the model must explore, right? It has to operate at a higher implied temperature to learn the landscape. But the temperature zero regret criterion does not evaluate that exploratory phase at all. So what does it evaluate? It strictly evaluates the argmax, the absolute highest probability, top-ranked response the model produces at inference time. Inference time. So test day. Exactly. It's the moment the model is deployed and responding to a prompt devoid of any force exploratory noise.

12:32I love this. It's the equivalent of evaluating a master chef. Yeah. Oh, that's a good way to look at it. Yeah. The old math, the KL regularized regret, was like standing in the kitchen taking deductions while the chef was throwing flour, burning experimental reductions, and making a chaotic mess trying to finalize a new recipe. Right. Grading the mess. Yes. But the temperature zero criterion waits until the chef walks out of the kitchen, enters the dining room, and presents their final perfected signature dish. It only judges the final plate. It completely ignores the kitchen noise. If we connect this to the bigger picture, that distinction is absolutely paramount.

13:10Yeah. By focusing solely on the temperature zero output, the research strips away the conflation. We are no longer mixing the cost of learning with the cost of exploring. We are asking a purely objective question. Which is? When the model is stripped of its forced randomization and asked to make its absolute best decision, does it choose correctly? So what does this all mean for those greedy alignment methods we started with? Like we take online RLHF and online DPO. We stop measuring them with the pessimistic, practice-punishing KL metric. And we apply this new streamlined temperature zero ruler that only looks at their deterministic best choice.

13:48What happens to the math? The mathematical landscape completely transforms. Under the temperature zero regret criterion, Kang's research proves that standard greedy online alignment methods do not suffer from endless logarithmic errors. They don't. No. Instead, they achieve constant cumulative regret, denoted mathematically as O of 1. O of 1. Okay, we really need to linger on the magnitude of that shift from O of log T to O of 1. Because in theoretical computer science, proving constant regret is like the holy grail for an algorithm. It fundamentally redefines the system's efficiency. It truly does.

14:24It changes a guarantee of endless accumulating errors into a guarantee of finite limitation. Meaning? An O of one bound means that the total number of errors the algorithm makes over its entire operational lifespan is capped at a constant value. It does not scale with time. Yeah. Whether the AI executes 10 training steps or 10 million training steps, the cumulative regret hits a ceiling. Once the greedy algorithm identifies the optimal policy, it essentially stops accumulating regret entirely. It achieves an optimal state. That is the mathematical equivalent of staring at an intimidating mountain.

14:59Yeah. You know, that endless logarithmic curve stretching into the sky and realizing, once you put on the correct pair of glasses, that the mountain is actually just a small speed bump. That's exactly what it is. You clear the initial hurdle of learning, and then you are on a perfectly smooth, flat road forever. And that mathematical proof finally provides the theoretical backbone for what engineers have been observing empirically all along. The research explicitly demonstrates this constant regret applies to online RLHF and online DPO. So the practice was right the whole time. Yeah. The unreasonable effectiveness of these greedy alignment methods was never unreasonable at all.

15:35The models were learning highly efficiently the entire time. Our diagnostic tool was simply misinterpreting their necessary variants as persistent failure. It's just a deeply satisfying resolution to the mystery. We started with this stark contradiction, right? Real-world performance, vastly outpacing theoretical predictions. Right. We dissected the mechanics of the greedy update and realized that the exploratory randomization forced upon the model during training was actually preventing it from getting stuck in local minima. Exactly. But we also discovered that our measurement tool, KL regularized regret, was unfairly punishing the model for that exact same necessary exploration.

16:15A completely broken ruler. Yep. And by shifting to a decision-centric temperature zero measurement, judging only the deterministic top-ranked response at inference time, the math finally snapped into alignment with reality, proving that these greedy methods operate with incredibly efficient O of one constant regret. The synthesis in Kang's research provides such clarity for the entire field of machine learning. But, you know, the underlying philosophy extends far beyond artificial intelligence. How so? Well, it serves as a rigorous mathematical warning about how we design metrics in our own organizations and educational systems.

16:51That's a great point for everyone listening. Yeah. If you design an evaluation framework that penalizes the variance in noise inherently required during the learning process, you will mathematically mask the true capability of the system you are trying to measure. That is a critical takeaway. If you grade the practice phase as if it were the final performance, you literally engineer a false diagnosis of failure. Precisely. You have to ensure you are measuring the actual decision-making capability, not the statistical cost of exploration. And this raises an important question for us to consider as we look at the future of complex systems.

17:29Oh, I'm ready for it. If fundamentally altering the mathematical ruler completely resolved the mystery of AI alignment efficiency, transforming an apparent structural flaw into a proof of constant perfect optimization, what other inefficiencies or mysteries currently plaguing large-scale AI behavior are actually completely logical, just waiting for us to view them through the correct diagnostic lens? Man, that is a phenomenal thread to pull on. Like, how many other systems look broken simply because we're holding the blueprint upside down? It's a great reminder that sometimes the breakthrough isn't building a better machine, it's building a better way to measure it.

18:05We will definitely be keeping an eye out for more research that challenges these foundational assumptions. But that wraps up our deep dive into Enoch Hanukeng's recent findings. It has been a fascinating exploration. Remember to rigorously question the metrics governing your own complex systems. Stay curious, keep exploring the mechanisms behind the math, and we will catch you on the next deep dive.

From the publisher

This research paper investigates why online alignment techniques for language models perform significantly better in practice than older mathematical theories suggested. The author argues that previous metrics were flawed because they confused the statistical difficulty of learning with the random noise required for exploration during training. By applying a more precise decision-centric evaluation, the study demonstrates that popular methods like RLHF and DPO actually achieve a much higher level of efficiency. Specifically, the paper proves that these greedy algorithms reach optimal performance levels more consistently than once believed. Ultimately, these findings provide a stronger theoretical foundation for the remarkable success seen in modern artificial intelligence fine-tuning.

More from Best AI papers explained

All 475 episodes
Demystifying the unreasonable effectiveness of online alignment methodsBest AI papers explained · 18 min
Listen in VO