In short
The episode explains Trajectory Bellman Residual Minimization (TBRM), a value-based training method for LLM reasoning that replaces step-by-step “micromanagement” with an objective that minimizes Bellman residual error over the whole generated trajectory. It claims TBRM is simpler, faster, cheaper (no critic), off-policy (reuses data), and produces emergent behaviors like verification, backtracking, and decomposition.
Guest backgrounds
No guests are named; the host discusses research from UW–Madison and MIT.
Key claims
TBRM uses the LLM’s own logits as value estimates, avoids a separate critic model, stabilizes value-based learning via trajectory-level Bellman residual minimization, and improves efficiency versus GRPO/PPO.
Notable examples
A geometry sphere-radius solution where the model continues to verify by substituting R=3 back into the volume formula; a hyperbola problem where it detects a contradiction (implying b^2 is negative) and backtracks to try the other form; decomposition where it labels step 1/step 2 for an inequality.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOCurrent Challenges in AI Training
2:41 to 4:00
Discuss the complexities of policy-based methods and their limitations.
“To understand why this new method, TBRM, is a big deal, we have to understand the status quo.”
Understanding the Bellman Equation
4:00 to 6:41
Learn about the Bellman equation and its significance in decision-making.
“Imagine you write a 500-word essay and you get an A.”
Introducing Trajectory Bellman Residual Minimization
6:41 to 8:34
Unveil the TBRM approach and its advantages over traditional methods.
“Richard Bellman, an absolute legend, father of dynamic programming back in the 1950s.”
Efficiency Gains of TBRM
8:34 to 10:54
Explore how TBRM reduces computational costs and improves performance.
“Instead of trying to calculate the Bellman error at every tiny token, which is just noisy and chaotic, it squares the error spanning the whole rollout.”
Emergent Behaviors in AI
10:54 to 13:16
Investigate the unexpected human-like behaviors that emerged from TBRM.
“But we know that in AI, simpler often means dumber.”
AI's Self-Verification Skills
13:16 to 14:01
Discover how AI learns to verify its own work autonomously.
“So the model is solving a geometry problem.”
Exploring TBRM and AI Problem Solving
14:01 to 16:43
Learn how the TBRM method enhances AI reasoning by incorporating metacognition and humility.
“That's what a good teacher drills into you.”
Provocative Thoughts on AI Learning
16:43 to 17:02
Consider the potential of AI to learn human-like cognitive behaviors like empathy.
“If we define the value correctly, what other human cognitive behaviors might emerge?”
Transcript
Automatic transcript. May contain errors.0:00I want to take you back to a moment that I think, well, it probably haunts a lot of us. You're back in high school. It's a math exam. Not just a quiz, but, you know, the big one. Final grade stuff. Oh, no. Already sweating. My palms are clammy just thinking about it. Right. And you are staring at this calculus problem that spans like half a page. It looks like hieroglyphics. But you dive in. You work through it line by line, sweating over the logic, the derivatives, the algebra. You fill three pages of notebook paper. The logic is beautiful. You feel like Newton. But on the very last line, you accidentally add 2 plus 2, and you write 5.
0:36The classic tragedy. And you get the paperback. Giant red X, zero points. All that logic, all that process, completely worthless because of a typo. But then picture the kid next to you. He has no idea what he's doing. Didn't study, he guesses. He writes down 42 by sheer accident, and he gets full credit. Which is, first of all, infuriating. But more importantly, if you're the teacher, how do you even grade that? Do you grade the outcome, the 42, or do you grade the process? Because the kid who guessed 42 learned nothing, but the kid who made a typo, they actually understand the math. And that dilemma grading the process versus the outcome is actually the central conflict in the biggest technology race on Earth right now.
1:17We are trying to teach artificial intelligence to reason, to do complex math, to code. And we are stuck in this exact debate. Do we micromanage every single step the AI takes, like a strict teacher watching your pencil move? Or do we just care if it gets the right answer at the end? It's a huge question. And up until now, the industry has mostly decided to be the micromanaging teacher. We have been obsessed with every single step. But today we're looking at research that suggests we might have been doing it the hard way. We're looking at a fascinating collaboration between UW-Madison and MIT. They've revived a concept from the 1950s, something called Trajectory Bellman Residual Minimization, or TBRM.
1:55Which is a mouthful. I can see why they use the acronym. It is a mouthful. But the promise is huge. They claim this method is simpler, faster, and actually makes the AI smarter than the complex machinery everyone else is using. So what is the mission for today's deep dive? We need to understand how stripping away the complexity and going back to some mathematical basics, specifically value-based learning might be the key to the next generation of reasoning models, and the best part. We're going to see how this simpler method caused the AI to teach itself human-like behaviors, like double-checking its work, without anyone telling it to.
2:33That's the part that really blew my mind. The idea that checking your work isn't something we programmed, but something it learned was, well, valuable. So let's get into it. To understand why this new method, TBRM, is a big deal, we have to understand the status quo. When we talk about training models to think, what are we actually doing? Usually I hear acronyms like PPO. Right. PPO or Proximal Policy Optimization or GRPO. These are the heavy hitters. If you talk to anyone at the big labs, this is, you know, their bread and butter. These are what we call policy-based methods. Okay. Let's strip the jargon immediately.
3:05What does policy-based actually mean in plain English? Think of it like coaching a chess player. In a policy-based approach, you are obsessed with the move. You have a policy manual that says, if the board looks like this, move the knight to F3. You are optimizing the specific option to take in every specific scenario. So it's really prescriptive. Do this, then do that. Exactly. And that works great for simple things. But think about applying that to a large language model writing a complex math proof. A proof might be hundreds of tokens or steps long. In a policy-based world, you are trying to optimize the probability of every single word or token it generates.
3:44You're asking, was the word therefore the best possible move right now? Was the number seven the best move? That sounds incredibly granular. You're micromanaging thousands of tiny little decisions. It is. And it leads to the biggest headache in reinforcement learning, the credit assignment problem. Credit assignment. Okay, break that down for us. Imagine you write a 500-word essay and you get an A. Which word got you the A? I mean, the whole thing. It's the vibe, the argument. Right. But the computer needs to know what to reinforce. Was it the intro? Was it that third sentence in the second paragraph?
4:16If you use the word concomitant in paragraph four, was that a good move or a bad move? Policy-based methods really struggle to figure that out. They have to guess how much credit to assign to every single token for that final result. I can see how that gets messy. It's like trying to pay a soccer team based on who touched the ball, not who scored the goal. That's a great analogy. And to manage this messiness, these methods usually employ a critic. This is literally a second AI model whose entire job is to stand over the shoulder of the first model, the actor, and grade its work in real time. So you have the student taking the test and a separate teacher grading every sentence as they write it.
4:54Yes. Like, that was a B-plus sentence. Ooh, that word was a C. And as you can imagine, running two massive models simultaneously, the actor and the critic, is computationally expensive. It eats up memory. It's unstable. It's like trying to balance two spinning plates on top of each other while riding a unicycle. Okay, so that's the current state-of-the-art, micromanagement. Expensive critic models. High compute costs. It works, but it's heavy. Now, enter TBRM. The paper calls as a value-based approach. How is that different from the policy chess player we just talked about? So, if the policy player focuses on the move, the value player focuses on the state of the board.
5:33Okay. A value-based player looks at the board and doesn't necessarily have a rule for move knight to F3. Instead, they look at the board and say, this position, this has a 95 % chance of winning. They assign a value to the situation. And if I know which board positions have high value? Then the move is obvious. You just make whatever move gets you to the highest value state. You don't need to memorize a policy for every action. You just need to be really good at recognizing what winning looks like. That feels much more intuitive. It's like navigating by landmarks instead of memorizing turn left, turn right, go straight.
6:04If I know where the destination is, I can figure out the path. So why haven't we been doing this all along? Well, historically, in the era of deep learning, value-based methods were kind of the black sheep. They were considered too unstable for something as massive as a language model. The math can get wonky when you have billions of parameters. They thought it was too hard to calculate that value accurately without things just spiraling out of control. But this team from UW-Madison and MIT, they found a way to stabilize it. They did. And it hinges on the concept right there in the name, the Bellman residual.
6:39All right. We have to talk about Bellman. I know we try to keep it light, but this is the core of the engine. Who is Bellman? Richard Bellman, an absolute legend, father of dynamic programming back in the 1950s. The Bellman equation is a fundamental reality check for decision making. It's basically a way to keep your expectations in check. Give me the explain like I'm five version of the Bellman equation or maybe explain like I'm a tired commuter. Okay. Imagine you're using a navigation app to drive home. You start the car and the app says estimated time to arrival, 30 minutes. That's your expected value.
7:13You expect the trip to cost you 30 minutes. Got it. 30 minutes to go. You drive for 10 minutes. You've put in the work. Now you look at the app again. In a perfect world, what should it say? It should say 20 minutes remaining. Exactly. Yeah. But let's say there was hidden traffic. You drive for 10 minutes, and the app now says 25 minutes remaining. Okay, so something is off. I drove 10 minutes, but the ETA only dropped by 5. My expectation was wrong. Exactly. There's a gap. The reality didn't match the expectation. That gap, that error, is the Bellman residual. In a perfect world, your current value estimate should be consistent with the reward you just got, the driving you did, plus the value of where you ended up.
7:55If they don't match, you have a high residual. You have an error. So learning happens by looking at that error and saying, oops, my internal map was wrong. Precisely. You minimize that error. You adjust your mental model so next time your prediction is better. Okay, so that's Bellman. Minimizing the error between expectation and reality. But the acronym is Trajectory Bellman Residual Minimization. What does the trajectory part add? This is the special sauce. Remember how PPO tries to grade every single step, every token? The micromanaging teacher. Right. TBRM says, forget the steps. Look at the whole journey.
8:29The whole essay. The whole proof. The whole trajectory. From the first word to the final answer. Instead of trying to calculate the Bellman error at every tiny token, which is just noisy and chaotic, it squares the error spanning the whole rollout. It basically asks, did this entire path lead to the expected value? That seems like it would smooth things out significantly. You aren't freaking out over every single comma. It does. And because it looks at the whole path, it simplifies the architecture dramatically. Remember that expensive critic model we talked about? The teacher hovering over the shoulder taking up all your RAM?
9:04The memory hog. You can fire him. Gone. Gone. TBRM doesn't need a separate network to estimate value. It uses the LLM's own logits, its raw confidence scores, as the value estimate. The student grades their own test based on how confident they are. Wait, wait, isn't that risky? If I graded my own tests in high school, I would have had a 4.0 GPA. I just say I'm 100 % confident on everything. Huh. It works because of that Bellman reality check. If the model is overconfident, if it says I'm 100 % sure this is right, but then gets the math wrong and gets zero reward, the Bellman error spikes. The math forces it to become honest.
9:43It forces the confidence to align with the actual probability of success. That is incredibly elegant. So no critic model, that must save a ton of memory. It's massive. But there is another efficiency gain that I think is even more important for the industry. PBRM is off policy. Off policy. We hear this term thrown around in AI papers all the time. Why does it matter to the person paying the server bills? Because PPO, the standard way, is on policy. Think of on policy, like trying to learn tennis. but you're only allowed to learn from the swing you just took five seconds ago. So I can't learn from my match yesterday.
10:17No. If you use old data, the math breaks. You constantly have to stop, generate new data with your current brain, and then train on that. It's incredibly wasteful. You're throwing away data constantly. That sounds incredibly inefficient. You're burning compute just to make data that you use once and then delete. Exactly. Off policy means you can learn from any experience. You can learn from your practice last week. You can learn from watching a video of someone else playing. You can recycle old data. So you generate a solution one time and you can squeeze juice out of it forever. Exactly. TBRM generates one solution, checks the Bellman residual, and learns.
10:52One and done. It reuses the data. On paper, this sounds like a slam dunk. Simpler, no critic, recycles data. But we know that in AI, simpler often means dumber. Sometimes you need the complexity. Does this actually work on the hard stuff? That's the big question. So they threw it at the heavy hitters. The math benchmarks. AME24 Math 500 Olympiad Bench. Okay, so these aren't what is 2 plus 2? No, no. AA is the American Invitational Mathematics Examination. It's competition-level math for top-tier high schoolers. It requires serious reasoning. It's the kind of stuff that makes smart adults cry. And how did TBRM hold up?
11:29Let's look at the AME24 benchmark using the Quinn 2.5 Math 7b model. This is a solid, mid-sized model, pretty accessible to most researchers, not just the giants. The standard GRPO method, complex, one-hit, 28.9 % accuracy. Okay. And TBRM? 30.5%. Wait, so it didn't just match the conference method, it beat it? It beat it by 1.6%. Yeah. Which in this world where people fight for 0.1 % gains is pretty significant. But the real story isn't just the accuracy. It's the cost of that accuracy. Right. Because we fired the critic and stopped the data waste. Exactly. Compared to PPO, TBRM achieved that better performance with, get this, 22.5 % less training time.
12:09If you're OpenAI or Google and you are spending billions on compute, 22 % is hundreds of millions of dollars. It's massive. And for the smaller labs, they measured 33 % lower GPU memory usage. That means you can train a smarter model on cheaper hardware. It effectively democratizes the ability to train reasoning models. You don't need a massive supercomputer cluster to do this anymore. You might be able to do it on a much smaller rig. That's huge for open source. But I want to pivot because while the efficiency is great for the accountants, there was something else in this research that I think is way cooler for the psychologists or, you know, the philosophers listening.
12:45The emergent behaviors. Yes. This is what I've been waiting to talk about. The researchers did not program the model how to solve these problems, right? Correct. They didn't say, hey, make sure you double check your work. They didn't say break the problem down into steps. They just gave it the Bellman objective. Minimize the error between what you expect and what you get. Maximize the value. And yet the model started acting strangely human. It really did. The paper highlights a few specific behaviors that just appeared. No code forced them. They just emerged. The first one is verification. Give us the example.
13:19Okay. So the model is solving a geometry problem. It's trying to find the radius of a sphere. It does the math, lots of algebra, and it calculates that the radius is 3. Standard stuff. It spits out the answer. Normally, yes. But here, without being prompted, it continues generating text. It writes, let's verify this by substituting R3 back into the volume formula. It checked its own work. It checked its own work. It ran the calculation backwards to see if it held up. And you have to realize, it wasn't told to do this. The model simply learned that trajectories where it stops to verify the answer tend to have a higher value.
13:55They're less likely to result in that Bellman error shock of getting a zero at the end. That is, that's kind of spooky. That's what a good teacher drills into you. Check your work. Yeah. And the AI derived that principle from scratch just by chasing the reward. It gets better. It also learned backtracking. What does that look like? There's a problem about a hyperbola. The model starts down one path, let's call it method A. it writes out five or six lines of math, but then it hits a wall. It derives something like two ballers or equals one to one. Which, unless we're getting into imaginary numbers, is usually a sign you messed up in standard geometry.
14:28Right. Now, a standard model or a hallucinating model might just plow right through. It might make up a number to keep going. Let's pretend Mabes-a-1 is 1 and hope nobody notices. Exactly. But this TBRM-trained model explicitly wrote, this implies b squared is negative, which is a contradiction. This is incorrect. So let's assume the hyperbola is of the other form. It admitted it was wrong. It stopped, admitted the error, turned around, and tried a different path. That is a level of metacognition of thinking about your own thinking that we usually say AI doesn't have. We usually say they just predict the next word.
15:03It mirrors human problem solving perfectly. Yeah. Though I don't go in a straight line. We fumble. We realize we're stuck. We adjust. TBRM encouraged this naturally because fixing a mistake increases the expected value of the trajectory compared to just giving up or, you know, lying. It's really interesting. By focusing on the value of the path, the model discovered that humility is a high-value strategy. That's a beautiful way to put it. Humility pays off. It also learned decomposition. Breaking things down. Yep. Faced with a complex inequality, it started explicitly labeling step one and step two, solving them in isolation, and then combining them.
15:39It structured its own thinking process to be more manageable. So to recap here, we have a method that is mathematically simpler. It fires the critic model, saving massive amounts of memory. It learns from old data, saving time. And the result is a model that double checks its work, admits when it's wrong, and organizes its thoughts. It really challenges the assumption that we need these incredibly complex algorithms to force AI to be smart. The paper essentially argues that we don't need to handhold the model with PPO. We've been micromanaging. We have. And if we return to value-based principles, specifically minimizing that Bellman residual over the full trajectory, we get a system that is sounder, cheaper, and effectively more human in its reasoning.
16:22It's a return to basics that feels like a leap forward. It makes me wonder about the future of this field. We always assume the next big breakthrough will come from a bigger chip or a larger dataset. But this suggests the next breakthrough might come from a smarter objective function. You know, work smarter, not harder. This raises a really provocative thought for me to leave our listeners with. If the model learned to backtrack and verify just to get a higher reward, what else could it learn? That is the question, isn't it? If we define the value correctly, what other human cognitive behaviors might emerge?
16:56Curiosity. Debate. Maybe even empathy. It's entirely possible. When you stop micromanaging the how and just focus on the value of the outcome, you leave room for the model to surprise you with the solution. So we are moving from being the puppet masters pulling every single string to being the coaches setting the rules of the game and watching the players figure out how to win. And watching them figure out winning strategies we didn't even know existed. Huge thanks to the team at UW-Madison and MIT for this research. The concept is trajectory Bellman residual minimization TBRM and honestly keep an eye on this acronym I think we're going to be hearing a lot more about it.
17:34Thanks for listening to the deep dive. See you next time
From the publisher
This paper introduces Trajectory Bellman Residual Minimization (TBRM), a new value-based reinforcement learning algorithm designed to improve the reasoning capabilities of large language models. Unlike traditional policy-based methods like PPO or GRPO, TBRM optimizes a single trajectory-level objective using the model's own raw outputs as Q-values. This streamlined approach removes the need for complex components like critic models, importance sampling, or clipping, significantly reducing computational and memory overhead. The authors provide a theoretical proof of convergence to an optimal policy even when using arbitrary off-policy data in deterministic environments. Empirical tests on mathematical reasoning benchmarks show that TBRM matches or exceeds the performance of established baselines while being faster and more resource-efficient. Ultimately, the research suggests that value-based RL is a principled and powerful alternative for training models to handle complex, multi-step thinking tasks.




