In short
The episode explains a new “Hierarchical Reasoning Model” (HRM) for multi-step AI reasoning, arguing that flat LLM-style chain-of-thought can be brittle and computationally shallow, while brain-inspired hierarchy can enable deeper, more general computation.
Guest backgrounds
No guests are mentioned; it’s a solo “Deep Dive” discussion.
Key claims
HRM uses two interacting recurrent modules (high-level H for slow planning, low-level L for fast execution) with “hierarchical convergence” that resets L between phases to prevent premature fading. It trains from scratch with ~27M parameters and ~1,000 samples, without chain-of-thought pretraining, using one-step gradient approximation, deep supervision, and adaptive computational time (Q-learning).
Notable examples
Near-perfect Sudoku Extreme (9x9), near-optimal MazeHard 30x30 paths, and ARC-AGI results (40.3% vs O3 Mini High ~34.5%, Claude 3.7 ~21.2%). It also reports inference-time scaling (more compute improves hard tasks) and different internal strategies per task (DFS/backtracking for Sudoku; pruning/refinement for mazes; incremental hill-climbing for ARC).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding AI Reasoning Challenges
0:45 to 3:00
Exploration of the limitations of current LLMs in complex problem-solving.
“comparisons maybe, and a good look at why this stuff really matters for the future of AI.”
Brain-Inspired AI: The HRM Model
3:00 to 4:45
Introduction to the Hierarchical Reasoning Model and its brain-inspired design.
“First, there's a high-level module, the H module.”
Innovative Training Mechanisms
4:45 to 8:00
Discussion of the innovative training methods and efficiency of HRM.
“And third, there's adaptive computational time, or ACT.”
Performance Benchmarks of HRM
8:00 to 11:00
Analysis of HRM's performance on complex tasks compared to larger models.
“Think of it as the size of the internal thinking space.”
Implications for AI Development
11:00 to 12:12
Insights on how HRM could impact the future of AI reasoning and architecture.
“The brain clearly relies heavily on hierarchy for complex thought.”
Transcript
Automatic transcript. May contain errors.0:00Welcome, curious minds, to the Deep Dive. Today, we're plunging into something pretty fascinating and frankly quite challenging in AI reasoning, specifically how even these incredibly powerful systems, you know, our LLMs, large language models, often just hit a wall with complex multi-step problem solving. Yeah, it's a major hurdle. And this deep dive, it's actually sparked by some really groundbreaking new research. There's this paper hierarchical reasoning model coming out of a sapient intelligence in Singapore and also Singapore University. And honestly, it offers a genuinely fresh direction, maybe even a surprising one.
0:34It's a fresh direction. Okay, intriguing. So what's the mission here? What makes this paper stand out? We're going to unpack how this new brain-inspired AI model might finally unlock something like true general-purpose reasoning, expect some surprising facts, some mind-bending comparisons maybe, and a good look at why this stuff really matters for the future of AI. Okay, let's unpack this then. When we talk about current AI reasoning, especially these LLMs, there's this interesting paradox, isn't there? We hear deep learning and it sounds, well, deep, but the architecture itself, some argue it's actually kind of shallow computationally.
1:09Can you explain that a bit? That's a great point. You know, LLMs are amazing at generating text, sure, but their reasoning, it often feels a bit like an imitation game. They tend to break problems down into these explicit language steps. It's called chain of thought or KOTI, but it's almost like a crutch. It can be brittle. If one step is off, the whole thing can fall apart and needs tons of data, it's slow because it works at the level of words or tokens. And fundamentally, it's not really doing deep computation like, say, a universal machine could. LLMs aren't inherently Turing complete for these complex algorithmic tasks.
1:43They can't just simulate any process that limits their depth. Right. And that's where the human brain comparison gets really interesting, isn't it? How do we do it differently? What's the brain's trick? Exactly. We don't sit there consciously translating every single thought into language to solve a problem, do we? A lot of our thinking is what you might call latent reasoning. It's computation happening in our internal hidden states. Much more efficient. Think about learning to ride a bike. First, it's all conscious steps. Then it just happens. It becomes latent. The brain gets this incredible computational depth from its hierarchical organization.
2:18Different brain areas, different cortical regions working together but on different timescales. You have slow, high-level planning happening alongside really rapid, low-level execution. And there are these complex feedback loops, refining things. Crucially, it avoids the huge computational costs, the prohibitive credit assignment costs, the plague AI training methods like backpropagation through time, BPTT. That process is just so memory-heavy for AI. The brain really gives us a blueprint for deep reasoning. Okay, so the brain provides this amazing blueprint. How does this new model, HRM, actually take that blueprint and, you know, build something with it?
2:54Well, that's the core innovation. HRM is structured with two recurrent modules that depend on each other. First, there's a high-level module, the H module. Think of it as doing the slow abstract planning, the deliberate thinking, setting the big goals. And second, a low-level module, the L module, doing the rapid detail computations, figuring out the small steps to reach the goal. And they interact through something the researchers call hierarchical convergence. It's pretty neat. The L module does a bunch of quick calculations, finds a sort of local solution for a subproblem. Then the H module steps in, takes that result, and this is key.
3:30It resets the L module to start on a new phase. This reset prevents that common issue in standard recurrent networks where the activity just fizzles out prematurely. It allows for much deeper multi-step processing. And what blows my mind is how efficient this thing is, given that sophistication. You said only 27 million parameters? Yeah, 27 million. It's tiny compared to most big models today, which often have billions. And it only needs a thousand training samples. That's almost nothing in the AI world. That's exactly. Minimal data, minimal size. And crucially, no pre-training needed, no chain of thought data fed into it.
4:04It learns from scratch on those few examples. Wow. So how does it train so effectively then? What are the mechanisms? There are a few innovative things going on. First, this one-step gradient approximation. It basically makes training much more efficient. It avoids that need for the complex, memory-hungry BPTTT calculations, keeps the memory footprint constant O1, technically speaking, and it's arguably more biologically plausible too. Second, they use deep supervision. This is kind of inspired by brain oscillations, actually. The model does multiple forward passes, segments, and calculates the error and updates parameters periodically.
4:40Importantly, it detaches the hidden state sometimes, stopping error signals from going too far back. This gives frequent feedback, improves stability. And third, there's adaptive computational time, or ACT. Think thinking fast and slow. System one, system two in the brain. HRM uses a learning algorithm, Q-learning, to figure out on the fly how many computation stats or segments a task actually needs. So it spends more effort on harder problems, saves compute on easier ones. Really smart resource allocation. Okay, efficiency is great. The design sounds smart. But let's talk results. Because you mentioned stunning performance.
5:11How good is it, really? The results are genuinely impressive, especially on tasks known to be hard for current AI. It's like you said, small model, big bunch. Take Sudoku Extreme, the tricky 9x9 version. HRM gets near-perfect scores. Now compare that to the big code models. They often score literally 0 % on similarly hard Sudoku sets. They just fail completely. Even massive modders struggle. One paper showed a 175 million parameter transformer trained on a million examples still got less than 20 % on complex maze tasks. And speaking of mazes, on MazeHard 30x30, HRM again finds the optimal path almost perfectly.
5:47But maybe the most striking benchmark is the abstraction and reasoning corpus, ARC-AGI. It's seen as a key test for, well, general intelligence potential. HRM trained from scratch 1 ,000 examples, 27M parameters, achieves 40.3 % accuracy on key parts of ARC. Now, leading Cote-T models, something like O3 Mini High gets around 34.5%. Claude 3.7, even with a huge context window, got about 21.2%. And those models are vastly larger, trained on way more data. And didn't they run a baseline, like a standard transformer with the same training setup but without the HRM architecture? They did exactly that, the direct pred baseline.
6:22And it consistently failed on these hard reasoning tasks. So it really isolates the benefit to HRM's specific hierarchical design. Absolutely. It underscores that the architecture itself is providing the advantage. And another cool thing, inference time scaling. If a task needs more thinking time, like a really hard Sudoku, you can just tell HRM to run more computational steps at inference time, no retraining needed, and its performance often improves. It can just think harder. Okay, so it's solving these complex problems, but how? What's going on inside? Can we get a peek at its thought process?
6:54Yeah, the researchers looked into that. It was quite insightful. For the maze task, it seems HRM starts by exploring several paths at once. Then it prunes the bad ones, sketches out a likely solution, and refines it. Inerative Improvement. For Sudoku, the strategy looks more like a depth-first search. It explores possibilities, and if it hits a dead end, it backtracks. Kind of like how a person might tackle it. But interestingly, for the ARC tasks, it seems to do less backtracking and more incremental changes. Like hill-climbing optimization, making small adjustments to gradually improve the state.
7:27What's really fascinating is that it doesn't seem locked into one strategy. It appears to adapt its reasoning approach depending on the task. choosing the best tool for the job, so to speak. Adaptability. Yeah. Okay. Now, you mentioned brain inspiration. Let's dive deeper into that connection. There's this concept of dimensionality in the brain. How does that relate? Right. Dimensionality in neural representations. It's a key idea in neuroscience. Basically, a brain region's ability to handle diverse, complex information is linked to the richness or effective dimensionality of its neural activity patterns.
8:01Think of it as the size of the internal thinking space. Higher order areas, the ones doing complex reasoning, need a bigger, higher dimensional space to represent more possibilities, more flexibility. Okay. A richer thinking space for complex thoughts makes intuitive sense. How do they measure that? There's a standard measure called the participation ratio, or PR. A higher PR means the activity is spread out across more dimensions. It's a higher dimensional representation. And yeah, this is actually observed biologically. In the mouse cortex, for instance, the PR of neural activity demonstrably increases as you go from lower level sensory areas up to higher level associative areas involved in planning and decision making.
8:38The paper even cites a strong statistical correlation for this in mice. It seems fundamental. OK, so the brain has this dimensionality hierarchy, higher areas, higher dimensions. What does this mean for HRM? Didn't it? Learn this. This is where it gets really compelling, the aha moment, perhaps. Yeah. HRM spontaneously reproduces this exact biological principle just through training. After training, its low-level module, the L module, operates in a relatively lower dimensional space. Its PR is about 30. But the high-level H module, its internal state operates in a significantly larger, higher-dimensional subspace with a PR around 90, almost three times higher.
9:18Wow. Yeah. And what's more, the H module's dimensionality actually scales up. If you train it on more diverse tasks, it expands its thinking space while the L module stays stable. And this isn't something they programmed in. It's an emergent property of the training process on these reasoning tasks. An untrained HRM shows no such separation. The PRs are roughly equal. This really sets HRM apart from many standard deep networks, which sometimes suffer from neural collapse, where representations become less dimensional during training. This high dimensional space in the H module is likely crucial for the model's cognitive flexibility.
9:52That parallel is striking, almost three times higher. Does that number mean anything compared to the biology? It's another fascinating point the paper makes. The ratio of the high-level PR to the low-level PR in HRM, around 2.98, is remarkably close to the ratio measured in the mouse cortex, which is about 2.25. That is uncanny. Okay, so pulling this all together, what does HRM signify for the bigger picture of AI development? Where does this take us? One major implication is that HRM seems to move AI closer to being Turing-complete in practice. It can potentially simulate any computational process, overcoming limitations seen in standard transformers for algorithmic tasks.
10:31It sort of realizes the promise of earlier ideas like neural algorithm reasoners, but avoids the pitfalls like premature convergence or the reliance on costly BPTT that held them back. It's also a different path than, say, using reinforcement learning, RL, to improve chain of thought. RL often helps models use existing cokey abilities better, but it can be unstable and needs lots of data. HRM uses this more direct, dense supervision through gradients. So what's the main takeaway the researchers emphasize? What's the core message here? Their main conclusion is pretty direct. The brain clearly relies heavily on hierarchy for complex thought.
11:06Yet mainstream AI has largely stuck with non-hierarchical models like flat transformers. This work challenges that dominance. It presents HRM as a truly viable, potentially superior alternative to chain of thought for complex reasoning, pushing towards a framework for more universal AI computation. Right. And that really wraps up our deep dive for today. We've explored how this brain-inspired hierarchical model, HRM, achieves really powerful, efficient reasoning, tackling tasks that stump much larger models and doing it with remarkably little data. Absolutely. And maybe the most profound thing isn't just that it solves the problems, but how.
11:41It seems to learn an organizational principle, this dimensionality hierarchy that mirrors intelligence in the brain itself. So here's a thought to leave you with. If a relatively small model, just 27 million parameters, can learn to reason this way and even adapt its internal thinking space like a biological brain, what does that suggest about our path towards truly general AI? Are we seeing a shift away from needing explicit step-by-step instructions towards AI developing more implicit, flexible, almost intuitive reasoning? Something to ponder. Thanks for joining us for this deep dive.
From the publisher
The research introduces the **Hierarchical Reasoning Model (HRM)**, a novel recurrent neural network architecture designed to address the limitations of current large language models (LLMs) in complex reasoning tasks. Inspired by the **hierarchical and multi-timescale processing observed in the human brain**, HRM employs two interdependent recurrent modules: a high-level module for **abstract planning** and a low-level module for **rapid, detailed computations**. The paper demonstrates that HRM significantly outperforms larger LLMs and Chain-of-Thought (CoT) methods on challenging problems like Sudoku, maze navigation, and the ARC-AGI benchmark, achieving high accuracy with **substantially less training data and fewer parameters**. This performance is attributed to HRM's **enhanced computational depth** and its ability to avoid premature convergence through a mechanism called "hierarchical convergence." The authors also highlight HRM's **biological plausibility**, particularly its efficient one-step gradient approximation for training and the emergent **dimensionality hierarchy** within its modules, mirroring brain organization.




