Maximum Likelihood Reinforcement Learning

6 Feb 2026 · 16 min · 10 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Maximum Likelihood Reinforcement Learning (MaxRL) argues standard reinforcement learning (RL) is only the first term in a Maclaurin-series expansion of the true maximum-likelihood objective. By adding higher-order terms via a truncation dial T and using a special gradient estimator that normalizes by successful samples (not total samples), MaxRL better weights rare hard successes, improves diversity, reduces mode collapse, and boosts test-time efficiency.

Guests

No guest names or backgrounds are provided in the transcript.

Key claims

RL ≠ maximum likelihood; MaxRL can be up to 20x more efficient in test-time scaling; MaxRL-Pareto dominates GRPO and RLOO; GRPO weights hard problems ~1/sqrt(P) while MaxRL weights ~1/P; GRPO can overweight easy problems due to a math quirk.

Notable examples

GSM8K (pass@k degradation avoided late in training); small MM2 in data-scarce regime (GRPO collapses ~10 epochs; MaxRL improves to 30–50); QIN3 1.7B/4B on hard math benchmarks; test-time scaling with 64 samples and verifier selection.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Revelation of MaxRL

0:44 to 1:45

Exploring the significance of the new paper on Maximum Likelihood Reinforcement Learning.

“And, well, it kind of drops a bomb on how we train AI to think.”

Reinforcement Learning vs. Maximum Likelihood

1:45 to 2:55

Discussing the differences between reinforcement learning and maximum likelihood in AI.

“Why do we even use RL for things like math or coding to begin with?”

Understanding the Mathematical Basis

2:55 to 4:03

Explaining how the Maclaurin series relates to the efficiency of MaxRL.

“If my pass rate goes up, isn't the likelihood of me being right also going up?”

The Truncation Mechanism

4:03 to 6:26

Describing the clever truncation mechanism introduced in the MaxRL paper.

“The researchers prove that if you take the maximum likelihood objective and you expand it using a mathematical tool called a Maclaurin series.”

Revolutionizing the Learning Signal

6:26 to 8:34

Investigating how MaxRL amplifies the learning signal from successes.

“if you had infinite compute and could turn that dial all the way up, Max RL becomes mathematically identical to the cross-entropy loss we use in normal supervised learning.”

Overcoming Mode Collapse

8:34 to 10:35

Analyzing how MaxRL maintains diversity in model training to prevent degradation.

“So the math says the failures are just noise for this specific goal, and the successes are the only signal that matters.”

Efficiency in Practice

10:35 to 13:15

Connecting MaxRL's theoretical concepts to practical efficiency gains in AI models.

“The model learns one way to solve it, and it gets a reward.”

A New Perspective on AI Learning

13:15 to 14:00

Reflecting on the implications of MaxRL for the future of AI learning methods.

“And they showed this on big models, too.”

MaxRL and Training Efficiency

14:00 to 14:30

Learn how MaxRL enhances efficiency in finding answers through diverse models.

“Right, and MaxRL uses that McLaurin expansion to let us turn up the dial, trading a bit more training compute for a much, much higher fidelity goal.”

Reinforcement Learning and Unified Theory

14:37 to 15:25

Explore the implications of reinforcement learning as an approximation and its relation to other AI fields.

“It does make you wonder, though, if reinforcement learning, this huge pillar of modern AI, is just a first-order approximation, what else are we approximating?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Imagine for a second that you've been training for a marathon. You've got this plan. Everyone swears by it. It's, you know, the industry standard. That's what all the pros use. Exactly. But then a team of researchers comes along and proves mathematically that your entire training plan is just a low resolution sort of pixelated thumbnail of the real plan. That's a great way to put it. It's like finding out you've been trying to paint a masterpiece, but you've only ever seen a blurry JPEG of it on a tiny screen. You get the basic shape, sure, but you're missing all the important details. Right, and that is more or less what's happening in the world of AI reasoning right now.

0:39We're looking at a pretty fascinating new paper from Carnegie Mellon University, Fahim Tajwar and his team. And, well, it kind of drops a bomb on how we train AI to think. It's a really significant piece of work. The paper's called Maximum Likelihood Reinforcement Learning, or MaxRL for short. And its core thesis is, it's bold. They argued that the standard method we use for reasoning models, reinforcement learning, or RL, is just a first order approximation of a much deeper principle. Maximum likelihood. Exact order approximation. That sounds like a very polite academic way of saying a rough draft.

1:13I mean, in a strict mathematical sense, that's precisely what it is. It's step one of a much longer process that we've all just been sort of ignoring. And what they found is that by bridging that gap, by going from the rough draft to the real deal, they can unlock these huge efficiency gains. Massive gains. We're talking about making AI reasoning tasks up to 20 times more efficient at inference time. 20 times. In this economy of GPU shortages and just massive compute costs, that is not a rounding error. That's a fortune. It really is. Okay, so let's unpack this. This identity crisis between reinforcement learning and maximum likelihood.

1:49Set the scene for us. Why do we even use RL for things like math or coding to begin with? Well, it all comes down to the kind of feedback you get. For standard supervised learning like teaching an AI to spot a cat in a photo, you have a nice smooth signal. The model gets it a little bit wrong. You can just nudge the weights a little bit. It's a smooth curve. But with reasoning tasks, like solving a math proof or writing code, the feedback is very stark. It either works or it doesn't. Pass or fail. Exactly. It's binary. It's what we call non-differentiable. You can't easily grade it on a curve.

2:22So the industry has settled on using reinforcement learning. You just treat the correct answer as a reward. And the goal is to maximize the expected reward. Right. Maximize the expected reward. Which, I have to say, sounds perfectly logical to me. If I want the model to be right more often, I should maximize the reward. What's the problem? The problem is really subtle, but it has these huge downstream effects. Maximizing expected reward, which is basically maximizing the pass rate, is not the same thing as maximizing the likelihood of the correct answer. Okay, you're going to have to walk me through that distinction.

2:56To me, those sound like synonyms. If my pass rate goes up, isn't the likelihood of me being right also going up? They're correlated for sure, but they behave very differently in terms of how the model learns, specifically with how it weights problems. Oh, weaving. Yeah. So imagine you have a big data set of math problems. Some are easy. What's 2 plus 2? Some are incredibly hard. Standard RL just sees a success as a success. A win is a win. It's a flat structure. Right. You get the easy one right. Reward is one. You get the hard one right. Reward is also one. But maximum likelihood is different.

3:31It naturally places a much, much higher weight on the hard problems, the inputs, where the model has a very low probability of success. So maximum likelihood is like that tough professor who ignores the A's you got on easy homework and just obsesses over that one extra credit question you finally managed to solve. That's a perfect analogy. And that obsession is what drives the learning much harder. The paper proves that standard RL just doesn't emphasize these hard cases nearly enough. and this is where the first order approximation thing really comes into play. Okay. The researchers prove that if you take the maximum likelihood objective and you expand it using a mathematical tool called a Maclaurin series.

4:09Whoa, hang on. Maclaurin series. I thought we were talking about AI ingredients. Now we're digging up 18th century calculus for a 2025 paper. Turns out the old math is the key to the new code. A Maclaurin series is just a way to represent a really complex function as an infinite sum of simpler terms. So like approximating a complicated curve with a straight line, then adding a little parabola to it, then a more complex curve. Exactly. It's layers of detail. Adding resolution to that blurry image we talked about. Yeah, get this. This is the bombshell. The researchers found that if you write out that series for the maximum likelihood objective, the very first term in that infinite list is standard reinforcement learning.

4:52So RL is literally just term number one. And we've just been ignoring the rest of the equation this whole time. We've been truncating the list at one. We've been looking at that low-res thumbnail and thinking it was the whole painting. So this Max RL paper is about figuring out how to access the rest of that list to download the full high-res image. But, I mean, you can't compute an infinite series in a training run. That's impossible. You can't. But you can go further than just the first term. And they introduced this really clever mechanism for it. It involves a parameter they just call T. T4.

5:25Truncation. Think of T as a truncation level, yes. Or even better, just think of it as a dial on a control panel. This dial controls how many terms of that infinite series we're going to use. Okay, I'm with you. I'm visualizing the dial. If you set that dial to T equals 1, you're basically telling the algorithm, just give me the first term, and boom, you get standard reinforcement learning. You're optimizing for just one sample being correct. That's the status quo. Okay, so what happens if I crank the dial? As you turn the dial up, t equals 2, t equals 4, and so on, you get closer and closer to the exact maximum likelihood objective.

6:00But there's a catch. Of course. Turning that dial up costs compute. Specifically, it costs samples. To estimate those higher order terms, you need to generate more rollouts, more attempts at solving the problem during the training phase. Ah, so this is the tradeoff. You are literally trading compute for, what, objective fidelity? you're spending compute to upgrade the math itself. Yes, that's the key insight. And what's amazing is that even in these non-differentiable settings, if you had infinite compute and could turn that dial all the way up, Max RL becomes mathematically identical to the cross-entropy loss we use in normal supervised learning.

6:39It actually unifies the two fields. That's fascinating from a theory perspective, but let's get practical. How do they actually run this? They need a real algorithm, not an infinite one. How do they calculate the gradient? Right, and this gets into the real secret sauce of the paper, the estimator. And when you first see it, the change looks almost suspiciously simple. It all boils down to the denominator. Suspiciously simple usually means I'm about to feel dumb for not thinking of it. Huh. Well, think about the standard reinforce algorithm, the grandpa of all this. You generate a bunch of attempts, sum up the rewards, and then you divide by$9 the total number of samples you generated.

7:14Right, and average over everything you tried, sum of rewards over total attempts. Max RL does something different. It sums up the rewards, but it divides by a dollar the number of successful samples only. Hold on. You divide only by the successes. That seems dangerous. What if I try a problem 100 times and I fail every single time? My denominator is zero. Well, technically, if there are zero successes, the gradient is zero, so you don't learn anything on that step. But think about the case where you try 100 times and succeed just once. Okay. In standard RL, you divide that one success by 100. The signal is tiny.

7:50It gets drowned out by all the failures. In max RL, if you succeed once, you divide by... You divide by one. That makes the gradient huge. Exactly. It massively amplifies the signal from rare successes. But wait, isn't that just survivorship bias? In every other part of machine learning, we're taught that throwing away your negative examples is a cardinal sin. You're ignoring 99 data points that tell you what not to do. That is the perfectly intuitive objection, right? But this is where that Maclaurin math comes back to save the day. The researchers prove it's theorem one in the paper, that when you expand that series for maximum likelihood, the math works out so that the failure cases essentially cancel themselves out.

8:28The true gradient for this objective is the average gradient calculated only from the successful tries. So the math says the failures are just noise for this specific goal, and the successes are the only signal that matters. Correct. It creates a success condition update. You stop diluting your learning signal with your failures. Okay, this is wild. I want to see how this plays out against the current heavyweights. Today, if you're training a big reasoning model, you're probably using something like GRPO. Group relative policy optimization, yep, or maybe RLO. Right. How does MaxRL stack up against those when it comes to weighting the hard problems?

9:06This is where it gets really clear. GRPO is effective because it does emphasize hard problems more than standard RL. If you look at the math, GRPO weights a hard problem by roughly 1 divided by the square root of its success probability. So 1 over root P. Okay, 1 over root P. Taxed RL weights it by 1 over P. Let's put some numbers on that. Let's do it. Say a problem is really hard. The model only has a 1 % chance of solving it, so P equals 0.01. standard RL gives that success a weight of one oh when is a win grpo takes the square root of 0.01 which is 0.1 1 divided by 0.1 is 10 so GLP o gives that rare success a weight of 10 a nice boost that's a big but max RL it just takes one divided by 0.01 that's a hundred Wow a hundred so while GRP o is saying hey pay attention to this max RL is screaming drop everything and learn this now exactly it's exponentially more aggressive and they also found this weird flaw in GRPO.

10:05For really easy problems, its weighting function inverts and it starts up weighting the easy stuff again. That makes no sense. Why waste time on something you've already mastered? You wouldn't want to. It's just a quirk of their math. MaxRL doesn't do that. And the results show this works. They found that MaxRL-Paretto dominates both GRPO and RLOO. Pareto dominates, meaning it's just better across the board, no trade-offs. Higher accuracy and better coverage. And that idea of coverage brings us to the next huge problem that Max RL seems to solve. The dreaded mode collapse. Ah, yeah, I've heard this term.

10:39The models get sharp, but in a bad way. What does that actually look like? So think about a simple math problem. The model learns one way to solve it, and it gets a reward. Then, because it got rewarded, it starts to only use that one method. It forgets any other way to think about the problem. It's a one-trick pony. It found a path that worked once, and now it clings to it for dear life. Exactly. And this is called pass at K degradation. As training goes on, the model loses diversity. It just memorizes one solution path. If that path is blocked for some reason or the problem is slightly different, the model just fails.

11:14It has no backup plan. It becomes brittle. Very. But MaxRL, because it's optimizing for the likelihood over the whole distribution of correct answers, preserves that diversity. They tested this on the GSM 8Pay math data set. And even late in training, MaxRL kept its options open. And didn't they do a stress test on this with a smaller model, the small MM2? Yes, the data scarce regime experiment. They trained it on the same small data set over and over. Methods like GRPO started to fall apart after maybe 10 epochs. They overfit. They collapsed into just memorizing answers. And Max RL. It just kept getting better.

11:49Up to 30, even 50 epochs, it didn't collapse. It just kept squeezing more real generalizable logic out of the data without getting tunnel vision. That resilience is incredible. But I want to come back to the headline number, the 20x efficiency gain. Because diversity is a nice academic concept, but efficiency is what pays the bills. How does that connect? This is the critical so what of the whole paper. It all comes down to something called test time scaling. Right. This is how all the top models work now. You don't just ask for one answer. You ask for, say, 64 different answers and then use Verifier to pick the best one.

12:21Exactly. Now, connect that back to mode collapse if your model is a one-trick pony. Then asking for 64 answers just gives me the same wrong answer 64 times, or maybe two answers 32 times each. You got it. If your model lacks diversity, spending more compute at test time gives you diminishing returns fast. But because MaxRL keeps the model diverse, because it maintains that high pass at k each of those 64 attempts, is much more likely to be a unique, valid attempt at a solution. So I don't need to generate 64 answers to find a correct one. You might only need three or four. This is where the number comes from.

12:56MaxRL achieves up to a 20x efficiency gain in test time scaling. To reach the same accuracy as GRPO, you need to generate 20 times fewer samples during inference. That is just massive. For any company running a service on these models, cutting your inference compute by a factor of 20 for your hardest problems is a game changer. It's an enormous economic win. And they showed this on big models, too. On QIN3, both the 1.7B and 4B models running on really hard math benchmarks, MaxRL just consistently found the right answer with fewer tries. It's so rare to see a paper that gives you this deep theoretical breakthrough, the connection to the Maclaurin series, and also says, oh, and by the way, here's a trick to fix your budget.

13:41It's the complete package. It redefines RL not as this totally separate thing, but just as a stepping stone. It says RL is just step one on the path to maximum likelihood, and then it hands you the tool, that simple change to the denominator to walk the rest of that path. So let's just recap this journey. We started with this idea that standard RL, the thing everyone uses, is just a low-resolution first draft of a better objective. Right, and MaxRL uses that McLaurin expansion to let us turn up the dial, trading a bit more training compute for a much, much higher fidelity goal. We saw how that one simple trick normalizing by successful samples instead of total samples is what actually aligns the algorithm with the true maximum likelihood gradient.

14:19And we saw how that fixes mode collapse. Keeping the model diverse and robust, preventing it from overfitting and memorizing. And finally, the big payoff. Because the model is so diverse, it's up to 20 times more efficient at actually finding the right answer when you need it to. It's just a beautiful example of how going back to first principles, really questioning the math of what you're optimizing, can unlock performance that you just can't get otherwise. It does make you wonder, though, if reinforcement learning, this huge pillar of modern AI, is just a first-order approximation, what else are we approximating?

14:56That's the thought I've been chewing on. We treat all these fields, supervised learning, reinforcement learning, unsupervised, as these separate silos, but this paper shows that two of them are really just different points on the same mathematical curve. We're just looking at different sides of the same mountain. Exactly. So what other distinct fields of AI are actually just low resolution versions of some single unified mathematical truth we haven't fully worked out yet? Maybe we're just waiting for the next person to find the right expansion that connects them all. A unified theory of AI training.

15:27Now that is a deep dive for another day. Thank you so much for walking us through this. This Max RL paper really feels like a glimpse into the future. My pleasure. It's always exciting to see the blurry picture get a little bit sharper. Absolutely. And to everyone listening, thanks for diving deep with us. Keep questioning your approximations, and we'll see you next time.

From the publisher

This paper introduces **Maximum Likelihood Reinforcement Learning (MaxRL)**, a novel framework designed to improve the training of models in tasks with binary feedback, such as mathematical reasoning and code generation. The authors argue that traditional **Reinforcement Learning (RL)** only optimizes a first-order approximation of the **maximum likelihood objective**, causing it to ignore harder problems where success is rare. **MaxRL** bridges this gap by using a compute-indexed objective that approaches exact maximum likelihood as more sampling resources are applied. By normalizing gradients based on successful outcomes rather than total samples, the method places greater emphasis on difficult tasks. Empirical results show that **MaxRL** significantly outperforms existing methods like **GRPO**, offering superior scaling with data and up to **20x gains in inference efficiency**. Ultimately, the framework mitigates the "distribution sharpening" and diversity loss often seen in large reasoning models trained with standard RL.

More from Best AI papers explained

All 475 episodes
Maximum Likelihood Reinforcement LearningBest AI papers explained · 16 min
Listen in VO