In short
The episode explains “looped flows,” a new AI architecture meant to let models “scale internal thinking time” for hard logic tasks, unlike standard deep nets that answer with a fixed compute budget. It contrasts looped/recurrent models (which update hidden state via repeated loops) with their training failures: backpropagation through time causes vanishing/exploding gradients, so training uses truncated BPTT, leading to unstable recurrences and “spurious attractors” (confident wrong fixed points or endless spinning).
Key claims
looped flows borrow diffusion/probability-flow ideas, training on progressively denoised “interpolins” that share the same noise sample and target across time steps to create temporal glue and avoid cheating (stop gradients). Inference uses probability-flow velocity with finer temporal grids (e.g., 8→128 steps) and optional stochastic integration to explore multiple solutions.
Notable examples
Sudoku Extreme accuracy 74.5% (8 steps) to 97.9% (128). End queens 10x10: 100% coverage of distinct valid solutions in 20 samples. ARCAGI 1: 44.6%→58.8%; ARCAGI 2: 7.8%→12.2%. On 65,000 Sudoku failure cases, looped flows recovered 90.9% of prior recurrent-model failures.
Guests
No guest names or backgrounds are provided; the episode is presented as a two-host discussion.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Looped Flows in AI
1:30 to 2:52
Learn about a new AI architecture called looped flows and its significance in problem-solving.
“We're going to explore how we can unleash an AI to conquer incredibly complex logic tasks simply by giving it the space to deliberate.”
Challenges of Training Looped Models
2:52 to 4:59
Discover the complexities involved in training looped models and the issues faced during backpropagation.
“It's trying to mimic that mental rotation we just talked about.”
The Problem of Unstable Recurrences
4:59 to 6:52
Understand the limitations of traditionally trained looped models that lead to unstable recurrences.
“So if you physically cannot train the model on the full 50 steps because the math either disappears or blows up, what is the workaround?”
Innovations with Looped Flows Framework
6:52 to 9:46
Explore how the looped flows framework offers a new way to train AI for better reasoning.
“We want the AI to think longer, but if we can't train it to understand the long-term goal of its own thoughts, it just loses its mind and spins out.”
Inferences in Looped Flows
9:46 to 11:31
Learn about the inference process in looped flows and how it enhances AI problem-solving.
“Wait, I have to jump in here and push back on this a bit because a thought just occurred to me.”
Exploring Stochastic Integration
11:31 to 14:00
Understand the role of stochastic integration in enhancing AI's ability to explore solutions.
“But what happens when we take away the training wheels and unleash it on a blank slate?”
Stochastic Integration Explained
14:00 to 14:40
Learn how adding turbulence helps AI explore multiple valid solutions.
“It is mathematically very different from just guessing.”
The N-Queens Problem and AI Challenges
14:40 to 15:30
Understand the complexities of the N-Queens problem and AI's struggles.
“A completely rigid AI just gets blown straight down the mountain to one cabin.”
Looped Flows: Achieving Massive Coverage
15:30 to 16:40
Discover how looped flows allowed AI to find all valid solutions.
“When they tried to map out multiple solutions, they would spit out completely invalid rule-breaking configurations roughly a third of the time and miss entire clusters of valid solutions entirely.”
Improving AI with Loop Flows
16:40 to 18:00
Learn how looped flows improved AI accuracy on complex tasks.
“On ARCAGI 1, the prior Stadia art for this architecture was 44.6 % accuracy.”
Show all 12 chapters
Stability and Understanding in AI
18:00 to 19:10
Explore how looped flows enhance stability and cognitive understanding in AI.
“When we look at that 90.9 % recovery rate, the takeaway is clear.”
The Future of AI Thinking
19:10 to 20:00
Ponder the implications of allowing AI extended time for complex problem solving.
“Yeah, the end of the hallucinated reflex.”
Transcript
Automatic transcript. May contain errors.0:00You know, when you or I are faced with a really difficult problem, like something genuinely complex, we don't just blurt out the very first thing that pops into our heads. Well, usually we don't. Right. Ideally, we sit with it. We mull it over. I mean, if someone asks you to imagine a complex three-dimensional object, maybe a weirdly shaped puzzle piece, and then asks you if it fits into a specific slot, you can almost feel your brain doing the work. Oh, absolutely. You are mentally grabbing that object, you know, rotating it in your mind's eye, looking at it from different angles, and just updating your mental model iteratively until you arrive at a solution.
0:37Yeah, it's a progressive, sustained cognitive effort. We instinctively scale our internal computation time to match the difficulty of the task at hand. Harder problem, you know, more thinking time. But historically, that is absolutely not how artificial intelligence operates. Right. Traditional deep learning models are fundamentally designed to give knee-jerk answers. They process an input through a fixed number of layers, do their pattern matching, and just spit out an answer with a fixed computational budget. Exactly. It's a rapid-fire reflex. There's no pondering phase. Right. But today, our mission is to explore a groundbreaking leap in how artificial intelligence handles logic.
1:16We are looking at a framework called looped flows. Which is really exciting. It is. This is an entirely new architecture that fundamentally changes how AI approaches hard problems, allowing it to finally scale its internal thinking time. We are looking at a paradigm shift that moves machines away from that instantaneous feedforward guessing and into the realm of patient iterative deliberation. And for you listening, whether you're trying to catch up on the absolute bleeding edge of artificial intelligence reasoning or you're just insanely curious about the mechanics of how machines are actually learning to think more like we do, this deep dive is going to break down exactly how AI is moving beyond simple pattern matching.
1:56Yeah, moving way beyond it. We're going to explore how we can unleash an AI to conquer incredibly complex logic tasks simply by giving it the space to deliberate. To really appreciate why looped flows are such a revolutionary concept, though, we first have to understand the brick wall that previous models hit when they tried to do this. Because the idea of giving an AI more time to think isn't entirely new, right? Right. It's not. There is a whole class of neural networks known as looped models or recurrent models. Let's unpack this. A looped model is essentially a neural network designed to recurrently update a hidden state.
2:34Yes. Instead of just passing information forward through distinct layers one time and being done with it, a looped model has a sort of internal workspace or memory. It passes that state back through the exact same computational loop over and over again. Exactly. By doing this, it artificially increases its effective depth. Right. It's trying to mimic that mental rotation we just talked about. And on paper, I mean, it sounds like the perfect solution. You just run the loop more times for harder problems. But there is a massive flaw in this approach, and it all comes down to the training phase. The training.
3:05Yeah. Training these looped models is a computational nightmare. Because of the memory requirements, right? Why is the training so uniquely terrible for a looped model compared to a standard network? It has to do with how neural networks learn, which is a process called backpropagation. When an AI makes a prediction during training, you measure how wrong it is, and you send an error signal backward through the network to adjust the internal weights. Makes sense. But with a looped model, because it's running the same loop over and over across time, you have to do what's called backpropagation through time, or BPTT.
3:43Wait, so if your model looped 50 times, you have to send that error signal backward through all 50 distinct steps? You do. Which means the computer has to keep the memory of every single one of those 50 steps active at the same time, just to calculate the math. Precisely. The memory and time costs grow linearly with every single step, but worse than that, it's mathematically unstable. You run into the vanishing or exploding gradient problem? Oh, let's break that down for a second, because vanishing gradient sounds like a magic trick, but it's actually just basic multiplication, right? It is, yeah.
4:15Think about the chain rule in calculus. When you send an error signal backward through 50 steps, you are essentially multiplying numbers together 50 times. Okay. If the value you are multiplying by is slightly less than 1, say, 0.9, and you multiply it by itself 50 times, the final number shrinks to almost zero. The signal just vanishes. It vanishes completely. The early steps in the loop receive zero feedback on what they did wrong. And I'm guessing the exploding part is the exact opposite. Exactly. If the number is slightly larger than 1, say 1.1, and you multiply it by itself 50 times, it compounds to a massive number.
4:52The gradient explodes, completely destroying the learning process by throwing the network's internal weights into chaos. Wow. So if you physically cannot train the model on the full 50 steps because the math either disappears or blows up, what is the workaround? How have people been training these things up until now? Well, out of pure necessity, the standard practice has been truncated backpropagation through time. You basically cut the gradient. Cut the gradient. Yeah. You let the model run its 50 loops, but you only give it feedback on a tiny window of its process, just one or a few recurrent updates at a time.
5:26The learning signal is artificially chopped up into bite-sized pieces. Oh, I see. It's like trying to learn how to play a massive, complex, hour-long symphony, but your music teacher only ever listens to three seconds of the second movement and grades you purely on that tiny window. That captures the problem perfectly, yeah. The early notes you play have absolutely no idea how they're supposed to set up the grand finale because the teacher isn't giving you feedback on the whole piece. The model is forced to try and discover a globally useful sequence of reasoning steps purely from these isolated local grades.
6:00And the data shows us exactly what the consequence of that short-sighted training is. It leads directly to what we call unstable recurrences. Unstable recurrences. Right. When you deploy these traditionally trained looped models, they exhibit some really bizarre and frustrating behaviors. Half the time, they fail to converge at all. They just spin their wheels forever, trapped in a chaotic loop, never settling on a final answer. So if half the time it spins its wheels forever, I'm guessing the other half of the time it just crashes into a completely wrong answer and refuses to let it go. That's spot on.
6:30We call those spurious attractors. Spurious attractors, okay. Yeah, an attractor is basically a state that the network naturally gravitates toward. A spurious attractor means the AI gets stubbornly stuck on a mathematically stable but completely wrong answer. It becomes wildly confident in a failure. Exactly. So short-sighted training absolutely ruins looped models. We want the AI to think longer, but if we can't train it to understand the long-term goal of its own thoughts, it just loses its mind and spins out. Basically, yes. If we can't train the long-term goal, how do we fix it without burning all the world's GPUs?
7:07That brings us to the core innovation we are exploring today, looped flows. This is a masterclass in rethinking the architecture. If we can't use traditional backpropagation over long periods of time, we need a new way to teach the network how to connect its thoughts. And how does it do that? This framework does it by borrowing a brilliant trick from the world of diffusion and probability flow models. Diffusion, so like the image generators, ones that start with pure television static and slowly refine it, step by step until you have a crystal clear picture of a cat riding a skateboard. The underlying math is very similar, yes.
7:43But instead of generating images, we are applying that logic to pure reasoning and problem solving. Okay, that is fascinating. In looped flows, instead of asking the AI to just guess the final logic answer from scratch, the problem is divided into a sequence of progressively easier denoising tasks. So how does this actually train the internal hidden state of the AI? What's fascinating here is the precise mechanics of the interpolins. During training, the model is fed an interpolin, which is essentially a mathematical mixture of pure random noise and the true, correct target solution. So it's looking at a scrambled, static-filled version of the answer.
8:22Right. And as the training steps progress, the noise level is gradually decreased. The task gets progressively easier. But here is the absolute crucial innovation. The model shares the exact same initial noise sample and the exact same true target across all of these different time steps. Ah. Wait, so the random static it starts with isn't changing every step. It's a consistent anchor. Exactly that. This shared noise and shared target act as a temporal glue. So if you're listening to this and picturing a chaotic mess of static, think of the static as a fingerprint. It's uniquely messy, but it's just the exact same mess across the entire timeline.
9:02And because it's the same mess, even though the learning gradients, the back propagation signals, are still truncated and only cover a few local updates at a time, the model is deeply incentivized to learn recurrent hidden states that transfer useful computation forward. Because it's working on the same underlying puzzle. I mean, if I'm trying to unscramble an image and step one clears up a little bit of the static, Like, I really want step two to remember what step one just did because the underlying image hasn't changed. The shared denoising objective across those decreasing noise levels naturally forces the early steps to actually support the later steps.
9:36It creates a global coherence in the reasoning process without requiring the impossible computational cost of global backpropagation. The early steps learn to lay the groundwork for the finale even if the teacher is only guessing them locally. Wait, I have to jump in here and push back on this a bit because a thought just occurred to me. If the AI is looking at the exact same noise and the exact same target across all these different steps, isn't it just going to cheat? Cheat? How so? Well, neural networks are notoriously lazy, right? They take the path of least resistance. Couldn't it just learn a simple algebraic trick to extract the target from the input instead of actually learning how to reason through the logic of the problem?
10:18That is a highly critical question, and it's honestly one of the first things you have to look for in the data. In theory, on paper, yes. It is mathematically possible for the model to exploit the shared noise and learn a trivial linear shortcut that just algebraically cancels out the noise to reveal the target. Right, just bypassing the actual cognitive work entirely. But the data overwhelmingly shows that it rarely, if ever, does this in practice. Really? Why not? The reason lies in the stop gradients. The model still has those barriers between its training steps. Because of those barriers in the code, the model cannot easily coordinate a grand multi-step cheating strategy.
10:56Oh, because the gradient is cut, the AI can't pass secret notes to its future self. It has no choice but to do the actual homework right there in the moment. That is a great way to visualize it. The local optimization forces the network to genuinely learn features that are useful for making the prediction at that specific moment in time. That makes a lot of sense. Trying to set up a fragile mathematical shortcut across broken gradients is actually much harder for the network to optimize than just genuinely learning the underlying logic of the puzzle. The constraints force honesty. That is incredibly elegant.
11:30So we've forced honesty during training. We fixed the broken training lib. But what happens when we take away the training wheels and unleash it on a blank slate? Like when we ask it to solve a brand new problem during inference? During inference, the AI is no longer looking at the true target, obviously. Instead, it integrates the velocity of a probability flow with its recurrent hidden states. Velocity of a probability flow. Okay, that sounds like a physics midterm. Let's break that down. Imagine the space of all possible answers as a vast mountainous landscape. The model has learned a flow, which is basically a set of wind currents that naturally push random noise toward the correct solution in the valley below.
12:11Okay, I'm picturing it. During inference, the AI drops a point of random noise onto this landscape and uses its hidden state to calculate the velocity and direction of the wind at that specific spot, taking a step forward along the current. And here's where it gets really interesting. Because of this architecture, you don't have to run the model at the exact same speed it was trained on. You can run it on a much finer temporal grid during inference. Yes. You can literally tell the AI, hey, instead of taking 16 big steps to solve this, I want you to take 128 tiny careful steps. Yes, that's inference time scaling.
12:46You are dynamically allowing the AI to spend more computation time on a single problem. And the results of doing this are just staggering. Let's look at the data on the Sudoku Extreme benchmark. These are incredibly difficult 9x9 logic puzzles. They're brutal, yeah. When they were on the loop flow model and gave it eight inference steps to think, it achieved an accuracy of 74.5%, which is solid. But when they scaled the temporal grid up, when they gave it 128 steps to think about the exact same puzzle, the accuracy rocketed up to 97.9%. It's amazing. It literally thought longer, deliberated with itself, and got the right answer.
13:24It is a profound validation of the core theory, but there is a second major detail about the inference phase that is equally important. Because this entire framework is built on probability flows, we aren't limited to deterministic straight line thinking. We can use advanced numerical integrators, specifically something called stochastic integration. Stochastic meaning involving a random variable. I assume that means we are just adding random guessing into the mix. Not random guessing, no. Think of it more as controlled turbulence. Using stochastic differential equations, or SDEs, we can intentionally inject a highly controlled amount of random noise into the AI's thought process at each step.
14:05It is mathematically very different from just guessing. Why would we want to make it noisier while it's trying to think? We do explore multiple paths. Yes, exactly. Think back to the landscape analogy. If the AI just follows the exact center of the wind current, it will always end up at the exact same destination. It will find one single valid solution. But what if a problem has multiple valid solutions? If we add a little bit of turbulence, stochastic noise, the AI gets nudged around as it moves. It can explore different paths, different wind currents, and discover entirely different destinations that are equally correct.
14:37If you're listening and thinking stochastic integration sounds like a math nightmare, stick with us. Just picture that wind current. A completely rigid AI just gets blown straight down the mountain to one cabin. An AI with stochastic integration gets buffeted around a bit, wanders through the woods, and realizes there are actually five different cabins in the valley that are perfectly good places to stop. And the data backs this up beautifully on multi-solution tasks. They tested this on the end queens problem. Let's visualize that for a second. The end queens problem involves placing queens on a chessboard so they can't attack each other.
15:14If you have a 10 by 10 board, you have to place 10 queens so that no two share the same row, column, or diagonal. Exactly. And there are many different ways to do that correctly. A failure state is basically two queens staring each other down on a diagonal, violating the rules. And older models, specifically previous state-of-the-art transformer approaches, struggled terribly with this. When they tried to map out multiple solutions, they would spit out completely invalid rule-breaking configurations roughly a third of the time and miss entire clusters of valid solutions entirely. But the looped flows, they didn't just find one rigid answer.
15:51By using that stochastic integration, they achieved massive coverage. It really did. On the N-Queen's 10x10 board, out of 20 samples, they recovered 100 % of the distinct valid solutions. The looped flows mapped the whole solution space without breaking the rules. It's the difference between a model that memorized one path through a maze and a model that actually understands the layout of the entire maze. So the theory is airtight. We've fixed the training loop and given it the ability to brainstorm multiple paths. But, you know, theory is cheap in AI. What happened when they actually forced this thing to solve extreme logic grids against the benchmarks?
16:26The benchmark destruction here is quite absolute. Let's run through the victories. The ARC-HEI benchmarks. These are considered some of the purest tests of abstract fluid intelligence in machines. Very hard tests. You give the AI a few visual examples of a logical rule like when a blue square touches a red square, turn it green, and it has to apply that unseen abstract rule to a completely new grid. On ARCAGI 1, the prior Stadia art for this architecture was 44.6 % accuracy. Which is pretty typical. But loop flows jumped that to 58.8%. Massive leap. On ARCAGI 2, it improved from 7.8 % to 12.2%.
17:07It'd even beat competitors on May's hard pathfinding. And if we synthesize these victories, we have to look back at the spurious attractors we discussed earlier. Remember how older recurrent models would get confidently stuck on the wrong answer or just spin their wheels forever? The unstable recurrences, yeah. It just completely loses its mind. There was an incredibly revealing in-depth analysis done on roughly 65 ,000 specific Sudoku instances. They ran the old standard recurrent model, and it failed 12.6 % of the time. Just completely broke down, got stuck, or spun out. They took those exact same 65 ,000 failure cases and handed them to the looped flows model.
17:44And let me guess, it didn't just spin out in the exact same spots. The looped flows recovered and successfully solved a staggering 90.9 % of those specific failures. Wow, 90.9%. It didn't just perform better overall. It specifically cured the exact cognitive diseases, the spinning out, the spurious attractors that plagued recurrent AI for years. So what does this all mean? When we look at that 90.9 % recovery rate, the takeaway is clear. The recurrent states that are learned via these flow objectives are radically more stable and vastly more capable than anything we have seen before in this space.
18:20Definitely. By breaking the problem down into those shared noise interpolants, the AI is building an internal logic structure that actually holds up under pressure. It proves that local denoising objectives grading the student on small moments can still teach a globally coherent understanding as long as you anchor those moments with shared noise. Which allows us to wrap this whole journey up. By combining the iterative self-updating loop of recurrent models with the step-by-step progressive denoising of flow models, the code has essentially been cracked on letting AI spend more time to think through hard logic problems without falling into endless chaotic cognitive loops.
18:58And connecting this directly to you, the listener, if you have ever interacted with an AI and felt frustrated because it gave you a rapid fire, superficial or wildly hallucinated answer, well, what we are looking at with loop flows represents the architectural shift to fix that. Yeah, the end of the hallucinated reflex. This is the shift toward an AI that can patiently internally deliberate. It functions so much closer to human reasoning, taking the time to turn the puzzle over and over until the pieces actually click. It's the end of the knee-jerk AI. But it leaves me with one final, incredibly provocative thought to mull over, and I want you to think about this too.
19:34Let's hear it. If this framework allows us to successfully decouple training time from inference time, if we can now dynamically scale an AI's internal thinking process without the model breaking down, what happens when we unleash a looped flow model and just let it think about a single, incredibly complex, world-changing problem for days, or weeks, or even a year before it gives an answer? If jumping from eight steps to 128 steps solves extreme Sudoku. What kind of solutions to humanity's greatest problems are waiting at the one millionth step of inference?
From the publisher
This paper introduces looped flows, a novel framework designed to enhance the reasoning capabilities of neural networks by merging recurrent hidden states with probability flow models. Traditional looped models often struggle with training instability because they cannot effectively backpropagate through many iterations, but this approach sidesteps that issue by using local denoising objectives across various noise levels. By gradually reducing noise and sharing information across steps, the model learns a stable recurrence that builds complex computations over time. During inference, the system solves difficult problems by integrating a stateful probability flow, which allows for increased accuracy through more intensive computation. This method significantly outperforms previous benchmarks in abstract reasoning and complex puzzles like Sudoku and Maze-Hard. Furthermore, the framework enables diverse solution generation by transporting different initial noise samples toward valid final outcomes.




