Universal Reasoning Model

6 Jan 2026 · 14 min · 10 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Universal Reasoning Model (URM) for hard algorithmic reasoning benchmarks (ARC-AGI1/2, Sudoku), arguing that smaller Universal Transformer (UT) architectures beat much larger vanilla LLMs by using recurrent refinement and strong non-linearity rather than brute scale.

Guest backgrounds

No guest names or bios are provided; the episode is presented as a “Deep Dive” with two unnamed hosts discussing the research.

Key claims

UT’s recurrent inductive bias (shared parameters reused across loops) converts compute into effective depth; URM adds CON4GLU (depth-wise short convolution after MLP expansion) for stronger non-linearity and uses truncated backprop through loops (TBPTL) for training stability. Ablations show removing attention softmax collapses performance; better optimizers (Muon) speed training but don’t improve final accuracy.

Notable examples

ARC-AGI1 pass@1: URM 53.8% vs TRM 40.0% and HRM 34.4%; ARC-AGI2 pass@1: URM 16.0% vs HRM 5.4% and TRM 4.6%; Sudoku accuracy 77.6%.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Algorithmic Reasoning

0:45 to 1:42

Exploration of complex iterative thought in AI and the role of universal transformers.

“models built on the universal transformer architecture or UTs, are consistently outperforming the huge large language models we hear about every day.”

Efficiency of Universal Transformers

1:42 to 2:28

Discussion on why smaller transformer models outperform larger ones in reasoning tasks.

“Because it forces us all to re-examine that belief that performance just comes from stacking more layers or adding billions more parameters.”

The Parameter Paradox

2:28 to 3:40

Surprising insights on how smaller models achieve better results despite fewer parameters.

“Like a human who has one brain but, you know, runs a problem through it eight times using their memory as a scratch pad.”

Introduction to Universal Reasoning Model (URM)

3:40 to 4:50

Explaining the architecture and enhancements of the URM for improved reasoning.

“When a standard transformer gets more FLOPs by just adding more non-shared layers, those resources often result in what's called redundant refinement in the higher layers.”

Innovations in URM: CON4GLU

4:50 to 6:06

Exploration of the CON4GLU module and its impact on model performance.

“The name suggests they've bolted a convolution onto a standard feedforward block.”

Training Stability with TBPTL

6:06 to 8:18

Understanding how truncated backpropagation through loops enhances training stability.

“It suggests that for these heavy reasoning tasks, the MLP, not the attention mechanism, is the model's main engine for expressive non-linearity.”

The Importance of Non-linearity

8:18 to 11:23

Examining how non-linear components affect the reasoning capabilities of models.

“Let's see how that translates to results.”

Optimizers and Training Efficiency

11:23 to 13:12

Discussion on the impact of optimizers on model training and performance outcomes.

“Attention might be context, but nonlinearity is the engine of thought.”

Key Takeaways on AI Reasoning

13:12 to 14:00

Insights on the architectural limits affecting AI reasoning capabilities.

“So to summarize for you, the listener, the key takeaway from this deep dive is that AI success in these hard algorithmic reasoning tasks is absolutely not a matter of brute scale.”

Exploring Model Capabilities and Future Directions

14:00 to 14:15

Learn about the model's complex functions and the potential need for architectural advancements.

“The intrinsic ability of the model to express the necessary complex functions, that might be the real ceiling for achieving human-level abstract reasoning.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to the Deep Dive. Today we are stepping into a genuinely fascinating corner of AI research. we're looking at how machines tackle problems that require complex iterative multi-step thought right the kind of algorithmic reasoning that really challenges even us exactly we're talking about benchmarks like the abstract reasoning corpus ARC AGI and even you know classic logic puzzles like Sudoku and these tasks are just crucial because they don't just require say massive data regurgitation or simple pattern matching right they demand a kind of depth intensive algorithmic reasoning, almost an inner loop of processing.

0:38And here's the biggest surprise that's coming out of the research. For these specific deep problems, much smaller models, models built on the universal transformer architecture or UTs, are consistently outperforming the huge large language models we hear about every day. Okay, let's unpack this immediately because that idea, the idea that the little guy is winning here, it just fundamentally challenges the, you know,$100 billion assumption that bigger is always better. It absolutely does. So our mission today is really twofold. First, we need to understand why these universal transformers are so efficient.

1:11What's the secret ingredient that lets them trump pure scale? And second? And second, we're diving into a new enhancement, the universal reasoning model, or URM, which takes those core capabilities and just amplifies them. We're really trying to uncover the true architectural secrets that drive this kind of complex reasoning. Exactly. This deep dive is all about recognizing that, you know, smart design and efficiency can beat brute force scale when the task demands sophisticated logic. Let's start right there with that core finding about universal transformers. Yeah. Because it forces us all to re-examine that belief that performance just comes from stacking more layers or adding billions more parameters.

1:50Well, the central finding is that complexity and scale are actually a bit of a red herring for these tasks. UT's success comes mainly from their recurrent inductive bias. And their nonlinear components, right. And their strong nonlinear components. You know, a standard transformer is sort of like an assembly line. Data goes through layer one, then layer two, then three, and so on. It never looks back. One and done. One and done. But UT's, they reuse the same parameters over and over again. They simulate this deep iterative thought through what we call recurrent refinement. So it's not just about having a deep network.

2:24It's about having a network that can think deeply using shared parts. Like a human who has one brain but, you know, runs a problem through it eight times using their memory as a scratch pad. That's a great way to put it. Which brings us to this parameter paradox. Yeah. Just how dramatic is the scale difference we're talking about? The comparison is, well, it's pretty striking. The universal transformer that was analyzed achieved a strong first attempt accuracy score on the ARC-AGI1 benchmark. Okay. And it did this with only four times the number of parameters compared to the smallest vanilla transformer baseline they tested.

2:58So you'd think, okay, just scale up the vanilla model. That should close the gap, right? That's the conventional wisdom. That's what you'd think. Just give the standard architecture more room to learn. But it failed. Dramatically. How badly? When researchers ramped up that vanilla transformer to use 32 times the number of parameters. Wow. Which is eight times larger than the UT. It still remained markedly weaker on the reasoning task. The UT got a much better score while staying so much smaller. Wait a second. If they're using that many more parameters and layers, that also means they're using more FLOPs, more computational cycles.

3:35If the computational budget is the same, why does the UT spend those cycles so much better? And that is the critical insight. When a standard transformer gets more FLOPs by just adding more non-shared layers, those resources often result in what's called redundant refinement in the higher layers. It's like having 20 people look at the same problem once and the last 10 just sort of repeat what the first 10 already said. Perfect analogy. The UT, on the other hand, converts that same computational budget into increased effective depth by sharing its weights and forcing that iterative calculation.

4:07It's just a better fit for the job. So the architecture itself is fundamentally more aligned with the nature of the problem. That leads us perfectly into the universal reasoning model, the URM. They identified recurrence and non-linearity as the secret sauce. And the URM is basically the UT framework with two targeted power-ups to enhance those exact drivers. That's it. The URM blueprint keeps that powerful recurrent core, but surgically enhances its ability to process rich, localized information, and, just as importantly, maintain stability during those deep thought loops. And those two innovations are the CON4GLU module and a training technique called truncated backpropagation through loops, TBPTL.

4:50Exactly. Okay, let's start with Tonda SW4GLU. The name suggests they've bolted a convolution onto a standard feedforward block. What does that actually look like? It's an enhancement to this standard SW4GLU feedforward block, which is already a really powerful gating mechanism in the model. They augment this with a depth-wise short convolution. How short are we talking? Specifically a very small kernel size, K2. Now, K2 is tiny. So it's like giving the model immediate peripheral vision, right? It just looks at its immediate neighbor before deciding if a feature is relevant instead of trying to look across the whole sequence like the attention mechanism does.

5:23That's a great way to think about it. It strengthens the non-linearity without adding all that high sequence-level complexity. It keeps the model efficient, but gives it more expressive wiggle room at every single recurrent step. And here's where the research gets really surprising, right? Placement matters. They tried six different places to insert this little convolution. Where did it actually make the biggest difference? This was the aha moment. The dominant performance gain came when they put it after the multilayer perceptron, or MLP, expansion, position F. Okay, so after the MLP. Crucially, inserting it inside the attention mechanism's pathway actually made performance worse.

6:04Whoa, that's a massive finding. It suggests that for these heavy reasoning tasks, the MLP, not the attention mechanism, is the model's main engine for expressive non-linearity. It really does. They found the workhorse and just made it stronger. The MLP is doing the heavy lifting of transformation, and attention is providing the context for it. Do they see any evidence of this inside the model's internal workings? Yes, the visualizations confirmed it. With Conswageol U integrated, the model's attention matrices became way more diverse and structured. The standard UT, by comparison, had these really homogenous, sparse patterns.

6:41So the better MLP helped the attention mechanism do its job better. Exactly. The non-linearity fixed the quality of the internal representation. Okay, that covers capacity. Now for the second innovation, stability. Truncated back propagation through loops, or TBPTL. If these UTs are running, say, eight or more recurrent loops, training must get pretty difficult. Oh, it's a huge problem. The inherent issue in these deep recurrent structures is optimization and stability. When the model runs all these loops during training, the gradients, the signals telling it how to adjust its weights. It can get noisy.

7:15Right. They propagate backward through all those steps, and you can accumulate a lot of noise and unstable signals from the very beginning of the process. It just hinders learning. That sounds like a classic challenge in recurrent neural networks. So how do they apply this truncation? It's an application of a classic idea, TBPTT. The solution is to limit the gradient computation to only the later loops, where the signal is hopefully a little cleaner. Can you give us an example? Sure. In an 8-loop run, the initial two loops are run forward only. The model processes the info, builds a representation, but no gradients are calculated for those early steps.

7:53And then it turns on the learning for the rest. Precisely. Gradients are only computed for the last six loops. And did that kind of moderate truncation turn out to be the sweet spot? It proved to be optimal. Truncating those first two loops gave them the best performance metrics. It's a really delicate balance. You need a long enough horizon to learn complex dependencies, but you have to jettison that early noise to stay stable. TBPTL provided that balance. So two elegant targeted solutions, comms with GLU for more nonlinear power and TBPTL for training stability. Let's see how that translates to results.

8:27How does the URM stack up against its predecessors like TRM and HRM? URM set a new state-of-the-art across all of these complex reasoning benchmarks, and we should really look at the pass at one accuracy. That's the percentage correct on the very first try. No second chances. Exactly. It's the most rigorous test of a model's sort of instantaneous reasoning power. All right, give us the numbers, especially for the hardest challenges. On the foundational ARC AGI-1 benchmark, URM hits 53.8%. That's a really significant leap over TRM at 40.0 % and HRM at 34.4%. That's a big jump. But the most impressive game is on ARCAGI 2, which is generally considered orders of magnitude harder.

9:06And what was the score there? URM scored 16.0 % pass at 1. That nearly triples HRM's score of 5.4 % and more than doubles TRM's 4.6%. That is not an incremental gain. Tripling the score on the hardest bendy mark? That suggests a fundamental increase in its ability to grasp abstract rules. It really does. And for Sudoku, URM achieved 77.6 % accuracy, surpassing the previous models comfortably. There's another nuance in the data about iterative refinement that really seems to confirm this, right? The gains get even wider when you allow for a larger sampling budget, like pass at 1 ,000. That's the crucial detail.

9:44It highlights the quality of the iterative process. If the model were just making, you know, lucky, brittle, one-step predictions, the gains wouldn't widen when you give it more chances. It would just be making more random guesses. Right. But because URM's gains widen so much, it confirms that its recurrent process is generating a richer diversity of high-quality, viable candidate solutions with each step. It's not just guessing, it's effectively exploring the solution space. That depth of refinement is exactly what you need. Oh. Okay, let's switch gears and go back to the fundamentals. The researchers ran ablation studies to really confirm that strong non-linearity is, well, a mandate for success here.

10:23What did they learn by systematically taking features away? They learned that complex abstract reasoning absolutely requires rich non-linear mappings. And the performance degraded monotonically. Meaning it got worse every single time. Every single time they stripped away one of the advanced non-linear components, the score went down. For instance, what happened when they swapped out the advanced gate they were using? When they replaced the powerful swigula U activation with simpler, more traditional ones like silo or re-ALE U, it led to a substantial and consistent performance drop. The less expressive the function, the worse the model was at abstract thought.

11:01But what was the ultimate proof? The unarguable evidence that non-linearity is just non-negotiable here? The ultimate proof was removing the attention softmax. The softmax function is a key source of non-linearity in the attention mechanism, taking it out completely, while it resulted in a dramatic collapse in performance. How bad. Plummeting the accuracy to near zero, it confirms it. Weakening nonlinearity systematically destroys the model's ability to reason. Attention might be context, but nonlinearity is the engine of thought. Okay, we've covered architecture, performance, and the necessity of nonlinearity.

11:35Let's talk optimization. They also investigated the training process itself, comparing their baseline with a specialized optimizer called Muon. The question was, can a better optimizer actually change the final capacity of the model? That's right. And what's fascinating here is the clear tradeoff they found. The Muon optimizer, which approximates second-order curvature, it's generally considered superior for these complex loss landscapes, it delivered a substantial speed-up. How significant was that speed boost in real terms? It was almost a two-fold speed-up. On the harder ARC AGI-2 benchmark, the Muon-optimized model reached a target accuracy in about 600 ,000 steps.

12:15The Atom baseline needed over 1.3 million steps. That is a huge gain in training efficiency, a massive cost saving. Huge gain. But this is where the implications get really deep. Did that speed, that superior optimization, actually lead to a better final model? And the answer was... No, it didn't. This is the crucial distinction. Despite that dramatic speedup, both optimizers converge to basically the same final accuracies, 53.8 % on ARC-AGI-1 and 16.0 % on ARC-AGI-2. So you're saying Muon cut the training time in half, but it failed to improve the final generalization capacity of the model. Precisely.

12:53This suggests a hard separation. The final, ultimate expressive capacity of the model, its ability to solve the hardest problems, seems to be determined purely by the architecture. by that recurrent structure and the enhanced non-linearity from consuid GLU. How efficiently or quickly you train it seems to be secondary to the inherent limits of the blueprint itself. Yeah. That's a perfect synthesis. So to summarize for you, the listener, the key takeaway from this deep dive is that AI success in these hard algorithmic reasoning tasks is absolutely not a matter of brute scale. No. It's a matter of intelligent structure recurrence and shared parameters combined with surgically enhanced non-linearity, like that tiny consuid GLU component, and training stability from TBPTL.

13:32These targeted architectural decisions led to dramatic state-of-the-art gains. And if we connect this to the bigger picture, to the future of AI development, consider this final provocative thought for a moment. Okay. The finding that a superior optimizer only speeds up training without actually improving the model's final generalization. It suggests that architectural limits might be the true unbreakable bottleneck. Not the training time, not the data. Right. The intrinsic ability of the model to express the necessary complex functions, that might be the real ceiling for achieving human-level abstract reasoning.

14:09An idea that suggests we might be waiting for the next architectural revolution, not just the next bigger training run. Thank you for joining us on this Deep Dive.

From the publisher

This paper introduces the Universal Reasoning Model (URM), a new architecture designed to solve highly complex logic puzzles like ARC-AGI and Sudoku. Researchers found that the success of Universal Transformers in reasoning tasks is driven by their recurrent inductive bias and non-linear depth, rather than overly complex designs. To build on this, the URM incorporates a ConvSwiGLU module to improve local token interactions and a truncated backpropagation method to stabilize training. These innovations allow the model to outperform existing systems while maintaining high parameter efficiency. Ultimately, the study demonstrates that iterative refinement through shared weights is more effective for abstract reasoning than simply scaling traditional model depth.

More from Best AI papers explained

All 475 episodes
Universal Reasoning ModelBest AI papers explained · 14 min
Listen in VO