In short
Reflective test-time planning for embodied LLM robots—moving from “static oracle” behavior (repeat the same failed plan) to trial-and-error learning via simulated reflection, hindsight regret, and on-the-fly policy updates.
Guest backgrounds
No guest names or external biographies are provided in the transcript; it’s a two-speaker discussion.
Key claims
Robots can plan by generating multiple candidate actions and using internal simulation to choose; then use an external reality check. The breakthrough is retrospective reflection: when the robot later gets stuck, it re-evaluates earlier successful steps and retroactively lowers their scores (temporal credit assignment). This enables double-loop learning by updating the doer policy and internal critic during operation.
Notable examples
“Closet of doom” packing; teddy bear scenario (bear placed in green box scores 100, later causes car failure); cupboard fitting task (sweet spot ~6 candidates; temperature 1.25–1.5); action budget effects (30 steps ~51.5% success, 50 steps ~60%, 100 steps dips).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe Limitations of Traditional AI
1:16 to 2:36
Discuss the shortcomings of previous AI models in adapting and learning.
“Until very recently, even the smartest AI robots, we're talking about these really advanced embodied LLMs, they were absolutely terrible at this.”
Reflective Test-Time Planning Explained
2:36 to 4:32
Learn about the concept of reflective test-time planning and its benefits.
“Yeah, to realize when a success was actually a strategic mistake.”
Understanding Reflection in Action
4:32 to 7:35
Distinct modes of reflection and their application in AI and human tasks.
“A robot might be able to critique itself in text, almost like writing a diary entry saying, I failed to pack the bag.”
The Problem of Short-Term Thinking in AI
7:35 to 9:32
How short-term success can hinder long-term goals in robotic tasks.
“You're stuck with your first guess, which is often wrong.”
Retrospective Reflection and Learning from Mistakes
9:32 to 11:28
How AI can learn from past mistakes to improve future actions.
“There is a large green box and a smaller box.”
The Mechanics of Double-Loop Learning
11:28 to 13:39
Explaining double-loop learning and its importance for AI development.
“Okay, so the robot realizes it messed up.”
Evaluating Performance Gains in AI
13:39 to 14:01
Discuss performance improvements of reflective robots compared to traditional models.
“OK, so we have a robot that brainstorms, acts, realizes it messed up later, and then performs brain surgery on itself so it doesn't make that mistake again.”
The Impact of Reflective Learning
14:01 to 15:49
Explore how reflective learning leads to significant performance improvements in robots.
“Compared to baselines like reflection, which is that text-only diary method and standard reinforcement learning, this reflective approach showed massive gains.”
Building Resilient Robots
15:50 to 16:32
Understand the shift towards programming robots for resilience and learning from mistakes.
“We aren't just programming robots to be perfect anymore.”
The Future of Machine Conscience
16:43 to 17:46
Discuss the implications of robots developing a form of conscience and self-critique.
“Before we go, I have a final thought to throw at you.”
Transcript
Automatic transcript. May contain errors.0:00Okay, let's unpack this. I want you to picture a scenario that I guarantee you we have all been in. Oh, I'm ready. You are packing the trunk of your car for a big road trip. Or maybe you're finally trying to organize that one closet in the hallway that has just become a total disaster zone. Right, the closet of doom. We all have one. Exactly. So you're in the zone. You pick up the biggest suitcase you have. You shove it into the back of the trunk, and it fits perfectly. Snug as a bug. Snug as a bug. You feel great. You feel like an absolute Tetris master. I'm giving you that completely false sense of victory.
0:34Because 30 seconds later, you turn around and pick up the cooler. And you realize with just absolute sinking dread that by putting the suitcase in first. You've blocked it. You have completely blocked the only spot where the cooler can actually fit. Classic mistake. Yeah. You've optimized for the immediate moment, but you completely failed the long-term goal. So what do you do? You don't just stand there staring at the bumper forever. You swear a little bit. You pull the suitcase out and you rearrange things. You realize your strategy was wrong and you fix it. It is just basic human trial and error.
1:07It is. I mean, it's fundamental to how we navigate the physical world. We try, we fail, we adjust. But here's the hook for today's deep dive. Until very recently, even the smartest AI robots, we're talking about these really advanced embodied LLMs, they were absolutely terrible at this. It's true. We tend to think of these robots as super intelligent because they can write poetry or, you know, code an app in seconds. But in this specific physical context, they acted more like what the research calls static oracles. Static oracles. That sounds impressive, but I have a feeling you're going to tell me it's not a compliment.
1:44No, it's definitely not a compliment. Basically, the robot would make a plan, try it, and if it failed, like blocking the cooler, it would often just try the exact same thing again. Just shoving the cooler into the suitcase. Exactly. Or it would fail blindly without understanding why. It wasn't learning during the job. It had no capacity to say, wait, that didn't work. Let me rethink my entire approach. So it's like trying to jam that suitcase in over and over again, expecting the physics of the car to change just because you want them to. Right. But that is changing. We are currently looking at a massive leap forward in robotics called reflective test time planning.
2:21Reflective test time planning. It essentially gives robots the ability to think like reflective practitioners. It allows them to simulate actions before they do them, judge the results after they do them. And here is where it gets really wild to look back with hindsight. Hindsight. Yeah, to realize when a success was actually a strategic mistake. Hindsight for robots. That is our mission for today. We are going to unpack how we are moving from robots that just follow instructions to robots that reflect on their own behavior in real time. They are effectively learning from trials and errors without needing a human to step in and save them.
2:56And the psychology behind this is fascinating because it is modeled directly on how our own brains work. Let's start there then. The human blueprint, so to speak. You mentioned this concept of the reflective practitioner earlier. Right. So this goes back to psychology research from the early 90s, specifically Donald Schoen. he was studying how professionals, you know, architects, engineers, nurses, how they actually solve complex problems in the real world. Under pressure. Exactly. And he argued that humans operate using two distinct modes of reflection. First, there is reflection and action. Reflection and action.
3:30Got it. That is thinking on your feet. Imagine you are a chef making a sauce. You taste it, and right in that moment you think, okay, too acidic. You add sugar immediately. You aren't stopping the whole cooking process. Right. You aren't turning off the stove. You are critiquing and adjusting the plan while you are executing it. Okay. So that's the pause before the leap or the adjustment mid-flight. Precisely. If I pull this block from the Jenga tower, will it fall? You simulate the physics in your head. Now, the second mode is reflection on action. On action. So this is past tense. Yes. This happens after the fact.
4:04The tower fell. The sauce is ruined. Now you analyze. So you step back. Yeah. You step back and say the tower fell because I pulled the bottom block. You are looking at the outcome to update your understanding of the world so you don't make that same mistake next time. So one is predicting the future. Yeah. And the other is analyzing the past. In a nutshell, yes. And the robot gap, the reason AI has struggled so much with physical tasks is that previous attempts usually only did one of these. Or they didn't connect them effectively at all. They were completely disconnected. Right. A robot might be able to critique itself in text, almost like writing a diary entry saying, I failed to pack the bag.
4:39But that text realization wouldn't actually change its brain or its underlying motor control policy. So they're just complaining about failing, not actually learning from it. That's like me writing, I will not eat donuts in my journal, and then immediately eating a donut because my hand didn't get the memo. Effectively, yes. What makes this new approach so revolutionary is that it bridges that gap. It uses a three-model setup. It's not just one AI brain. It's three distinct components working together inside the robot. Okay, walk me through the trio. Who are the players inside the robot's head?
5:15First you have the doer. This is the action generation model. It's the part that actually drives the robot, moving the arm, grabbing the object. The muscle. The muscle, but also the initial planner. Then you have the internal critic. This is an internal evaluator. Think of it as a simulator. Okay. It guesses if an idea is good before the robot actually tries it. So that's the reflection and action part, the imagination. Exactly. And finally, you have the external critic. This is the reality check. It looks at what actually happened after the robot moved. It provides the ground truth. So doer, internal critic, external critic.
5:48It sounds like a high stakes kitchen environment. You have the line cook, the expediter and the food critic all just yelling at each other. It can be. And the magic really happens in how they interact. Let's look at that first mode again. Reflection and action. The pause. You know, the best analogy for this, and tell me if I'm totally off base here, is writing a risky email. I like this. You're mad at a coworker. You draft one version that's super aggressive. You draft a second one that's passive aggressive. Then you draft a third one that's actually polite. You read them all, and your internal critic says, do not send the first two.
6:23That is a perfect analogy. That is exactly what the robot is doing. The data refers to it as test time scaling. The robot doesn't just grab the first idea that pops into its neural network. It generates n candidate actions. N being a variable number of ideas. Yes, samples. Say the task is put the toy car in the green box. The robot might brainstorm idea one, put it in the green box, idea three, throw it on the floor. And the internal critic steps in to judge. Right. It scores them. It simulates the outcome. For the orange box idea, the internal critic might say, score zero. The orange box is visually too small.
6:58It catches the error before it happens. But here's a practical question. If I'm writing that email or planning how to pack the car, I can't spend all day drafting 50 versions. I'll never get anything done. How many ideas should the robot brainstorm? What's the sweet spot? That is a crucial question, and the data on this is really specific. In experiments involving fitting objects into cupboards, it turns out the magic number is roughly six candidates. Six. That feels surprisingly low. I would have guessed computers need thousands of simulations. Why six? It's a bell curve of utility, really. If you generate fewer than six, say just one or two, you miss out on the really creative or optimal solutions.
7:38You're stuck with your first guess, which is often wrong. Makes sense. But if you generate more than six, say 10 or 20, you start generating junk. Yunk. Noise. The system starts confusing itself with too many low quality options that don't add value. It slows everything down without improving the result at all. It's essentially analysis paralysis. So six is the limit of productive brainstorming for a robot. In this context, yes. And there's another dial they have to tune, which is the sampling temperature. Temperature. Like... Does the robot get hot when it thinks too hard? Not quite. In AI, temperature refers to creativity or randomness.
8:13If the temperature is too low, the robot is boring. All six ideas look exactly the same. It's rigid. If it's too high, the robot becomes completely chaotic. It starts hallucinating physics that don't even exist. I will face through the wall. Exactly. Or I will balance this heavy suitcase on top of a fragile lamp. The research shows the creativity dial needs to be set between a 1.25 and 1.5. That's the Goldilocks zone for robust planning. Okay, so we've got the robot brainstorming six ideas, staying creative but grounded. It picks the best one. Now it has to act. This brings us to the next phase, right?
8:49Reflection on action. Right. The robot executes the move. Now the external critic steps in to look at the result. And usually this is where we just say, did it work, yes or no? Usually. But this framework uncovered a fascinating problem. The problem of short-term thinking. Or the trap of local success. Local success. It sounds like a good thing, though. It sounds good, but it can be fatal for a long-term task. Just because an action succeeded right now doesn't mean it was the right move for the ultimate goal. Give me an example of that. There is a specific scenario used in the testing called the teddy bear scenario.
9:23It's simple, but it is devastating for a standard robot. I'm listening. The setup is this. The robot needs to put a large toy car and a teddy bear away into boxes. There is a large green box and a smaller box. Okay, two items, two boxes. The robot picks up the teddy bear first. It looks at the green box. The teddy bear fits easily, so the doer says, put the bear in the green box. Success. Immediate feedback. The external critic says, great job. The bear is in the box. Score 100. The robot feels great. It thinks it's winning. I sense a massive butt coming. But then the robot picks up the large toy car.
10:00It looks around. The car's huge. It only fits in the green box. But the green box is full of bear. Exactly. The green box is occupied. The robot is stuck. A successful move 30 seconds ago has caused a total failure now. That is frustratingly relatable. I've totally done that packing the trunk. I put the soft bags in first because they fit, and then I have absolutely nowhere for the hard suitcase. Right. And a standard robot, that static oracle would just freeze or keep trying to jam the car into the small box. It doesn't understand that its past success was actually a mistake. It just thinks I got 100 points for the bear, so the bear move was good.
10:35So how does this new system fix that? Because the bear is already in the box. The mistake is in the past. This is the real breakthrough. It's called retrospective reflection. Retrospective reflection. The aha moment. The system doesn't just evaluate the now. It looks at the history. When the robot gets stuck with the car, it triggers a review. It looks back at the tape. It looks back at that first move, putting the bear in the green box. At the time, remember, it scored 100. Right. With hindsight, the system reevaluates that move. It actually changes the score from 100 down to zero. Wow. It retroactively fails itself.
11:09Yes. It generates a new thought. It says, I shouldn't have put the teddy bear in the green box. It was the only spot for the car. That is huge. That is literally learning from regret. It solves what is known as the temporal credit assignment problem. It assigns the blame for the current failure to the correct past action. Okay, so the robot realizes it messed up. But realizing you messed up is one thing. Ensuring you don't do it again is another. I can realize I shouldn't have eaten that second slice of cake, but I might do it again tomorrow. That is the difference between single-loop learning and double-loop learning.
11:45Okay, let's unpack that. This comes from organizational theorist Chris Argyris back in the 70s. Single-loop learning is just changing the action. The car doesn't fit. Push harder. Which doesn't work. Right. Double-loop learning is much deeper. It's changing the underlying belief system. It's asking, why did I think that was a good idea in the first place? So how does a robot do double-loop learning? It doesn't exactly have a therapist to talk through its childhood. No, it has something arguably better, test time training. The system isn't just remembering text like a diary. It is actually updating its own neural network parameters, its weights, while it is working in the house.
12:22Wait, it's rewriting its own brain code on the fly? Effectively, yes. It creates a self-supervised data set from its own reflections. It uses policy gradient updates to train the doer. The doer? It teaches the doer to favor actions that work in hindsight, not just in the moment. But it also trains the internal critic using supervised learning. So the internal critic gets smarter too. It has to. It teaches the internal critic to predict that the green box is a bad idea next time. It essentially aligns the robot's imagination with reality. That sounds incredibly computationally expensive. Updating a massive AI model in real time.
13:01Doesn't that crash the computer or take a week? It would if they updated the whole thing. But they use a clever technique called LoRa, low rank adaptation. LoRa, I've heard that term in image generation. It's the same principle. Instead of retraining the whole brain, which is like rewriting an entire textbook, they just train a tiny specific slice of adaptation layers. It's like adding sticky notes to the textbook margins. Sticky notes? It's highly efficient. They found a specific configuration. Rank 8, alpha 16 was the absolute sweet spot. Rank 8, alpha 16. Sounds like a secret code. It essentially means they are updating enough of the brain to be flexible, but not so much that the robot forgets everything else it knows.
13:39It keeps the learning stable. OK, so we have a robot that brainstorms, acts, realizes it messed up later, and then performs brain surgery on itself so it doesn't make that mistake again. Does it actually work? Or is this just cool theory? Oh, it works. And the performance gains are significant. Give me the numbers. They tested this on long horizon household tasks and that cupboard fitting task we mentioned. Compared to baselines like reflection, which is that text-only diary method and standard reinforcement learning, this reflective approach showed massive gains. Massive. We are talking about the difference between failing most of the time and succeeding fast.
14:16And the most interesting part was the ablation study, basically turning parts of the system off to see what really matters. And what mattered. If you remove the retrospective part, the hindsight performance just tanks. You absolutely need the regret. you need to be able to look back to move forward. That makes perfect sense. If you never realize the bear caused the problem, you'll never fix the problem. But surely there's a cost. All this thinking and reflecting takes time. It does. It takes steps. And this brings us to the action budget. The action budget. Is that like an allowance? Sort of. Think of it as the runway you give the robot to figure things out.
14:54If you give the robot a tight budget, say only 30 steps to clean the room, it struggles. It has a success rate of about 51.5%. It panics. Right. It doesn't have enough time to fail, reflect, and fix. But if you bump that budget up to 50 steps, success jumps to 60%. So giving it just a little more grace period allows the learning to kick in. Exactly. It needs the runway. However, there is a limit. If you give it too much time, say 100 steps, performance actually dips slightly. Really? Why would it go down? It starts over exploring. It starts wandering. It's like if I gave you all day to pack that car trunk, you might start trying to color coordinate the luggage instead of just closing the lid.
15:34Procrastination by perfectionism. Exactly. You start solving problems that don't even exist. Okay, that is fascinating. There is a sweet spot for failure. You need enough room to make mistakes, but not so much room that you stop trying to be efficient. Precisely. So let's zoom out here. What we are seeing is a massive shift in how we build intelligence. We aren't just programming robots to be perfect anymore. We are programming them to be resilient. That is the key word, resilient. We are building robots that simulate the future, critique the present, and rewrite their understanding of the past.
16:07It's taking the move fast and break things mantra and adding, and then fix what you broke. And learn from it. There's a quote by Katherine Schultz that comes to mind here. Error isn't simple darkness. It sheds a light of its own. I really love that. Mistakes are data. For a long time, we tried to build robots that avoided mistakes at all costs. But by building robots that embrace mistakes, that actually use them as the fuel for test time training, we make them robust. We make them capable of handling the messy, unpredictable world we actually live in. It's empowering in a way. Even for us humans, the mistake isn't the failure.
16:42The failure is refusing to look back and learn from it. Absolutely. Before we go, I have a final thought to throw at you. A bit of a provocation. Go for it. We've been talking about regret in robots. I shouldn't have put the bear in the box. If robots can now regret past actions and literally rewrite their own neural pathways based on that regret, are we approaching a form of machine conscience? That is a heavy question. Or at least a consciousness of quality. And what happens when they start reflecting not just on their own actions, but on the instructions we give them? What happens when the robot looks at a command and says, internal critic scores zero?
17:22That's a bad idea, human. That might be the next frontier, honestly. When the external critic becomes a critic of the user, that would certainly change the dynamic in the kitchen. It certainly would. I'm sorry I can't load the dishwasher that way. It's inefficient. I'm afraid I can't do that, Dave. Exactly. Something to think about the next time you're reorganizing the closet and realizing you put the winter coats in the completely wrong spot. Thank you for taking this deep dive with us. My pleasure. And to you listening, reflect on your own test time planning today. Maybe give yourself a little extra action budget to make a mistake and learn from it.
17:53We'll see you next time.
From the publisher
This paper discusses a framework for reflective test-time planning designed to improve the performance of embodied Large Language Models (LLMs) during robotic tasks. This system utilizes double-loop learning, where agents re-evaluate their past decisions through hindsight assessments to correct underlying strategic errors. By incorporating internal reflection for immediate scoring and retrospective reflection for long-term credit assignment, the model adapts its policy at deployment without requiring additional pretraining data. Experimental results in household and cupboard fitting tasks demonstrate that this approach significantly reduces execution waste and improves success rates compared to standard methods. Furthermore, the researchers employ Low-Rank Adaptation (LoRA) to efficiently update the models, ensuring that the robots can learn from their own trials and errors in real-time environments.




