Beyond Bigger Models: Recursion As The Next Scaling Law In AI

1 May 2026 · 38 min · 17 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How recursion at inference/training time can improve AI reasoning without scaling model size, contrasting hierarchical reasoning models (HRM) and tiny recursive models (TRM) as “next scaling laws” beyond bigger transformers.

Guest

Francois Chaubard, YC visiting partner; discusses AI research history (RNNs, backprop through time) and reasoning limits in LLMs.

Key claims

LLMs are largely one-shot feed-forward at training, lacking latent “compression” like RNN hidden states, so they can’t efficiently perform incompressible algorithms (e.g., sorting lower bounds, Sudoku/mazes). HRM uses three recursion levels (low/high latent refinement plus outer refinement) and a trick: truncated backprop through time with stop-grad while reusing updated latent states as “mini-batches.” TRM simplifies architecture (shared net, fewer layers) and backpropagates through one full latent recursion step, improving performance with fewer parameters.

Notable examples

sorting unsorted-to-sorted lists (comparison lower bounds), Sudoku and mazes as incompressible problems, ARCPrize 1/2 results (HRM ~27M params, TRM ~7M params; HRM ~70% ARCPrize 1, TRM ~87%).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Introduction to Recursion in AI

0:45 to 3:49

Discussion on recursion's role in improving model reasoning beyond just scaling size.

“you already did an amazing lecture on RNNs and LLMs in one of the previous videos.”

Limitations of LLMs in Reasoning

3:49 to 8:02

Exploration of why LLMs struggle with certain reasoning tasks and their limitations compared to recursive models.

“You know, many people think about LLMs as doing reasoning, and we're going to talk about that a little bit later.”

Deep Dive into Hierarchical Reasoning Models

8:02 to 11:39

In-depth look at HRMs, their design, and how they achieve effective results with minimal parameters.

“And I guess that's one interpretation of it, of the way that they're talking about classifying these hierarchies of frequencies.”

The Mechanics Behind Recursive Models

11:39 to 14:00

Understanding the backpropagation and training techniques that enhance performance in recursive models.

“If I take a batch of like ImageNet or CIFR10, and I forward pass through the model, and I get some loss, and I back prop, and I update the weights, I would go get a different batch for the next one.”

Understanding Deep Recursion in AI

14:00 to 14:40

Exploration of why certain machine learning methods work, particularly deep recursion.

“That's not sufficient support for why it's working.”

The Bioplausibility Argument in Machine Learning

14:40 to 16:00

Discussion on the relevance of biological theories in machine learning and their implications.

“And like, it didn't work and you didn't need that.”

Dismissing Bio-Plausible Constructs

16:00 to 18:10

Critique of the reliance on bio-plausibility in machine learning models and the divergence from biological accuracy.

“Exactly, it runs better on a GPU, it's more efficient in some capacity that is relevant to how we actually encode it in a computational system.”

Memory Caches and Algorithm Efficiency

18:10 to 21:00

Insights on the role of memory caches in enhancing algorithm efficiency and performance.

“And like that's the easiest thing to do and you don't need backprop at all.”

Key Takeaways from the HRM Paper

21:00 to 23:10

Summarization of critical insights from the HRM paper and its implications for machine learning.

“Basically, the Sapien authors, which is huge kudos for this paper, because there's so many innovations in this paper, didn't really do like scaling ablations on every single one of the inputs.”

Comparing TRM and HRM: Key Differences

23:10 to 25:40

Analysis of the differences between TRM and HRM papers, focusing on their architectural innovations.

“It seemed like another thing that changed was having this sort of double layer of higher order thinking and lower order thinking.”
Show all 17 chapters

Optimizing Recursive Methods in TRM

25:40 to 28:04

Examination of recursive methods in TRM and their advantage over traditional transformer approaches.

“that is a candidate answer, a proposed latent answer that is just an embedding space away, a one MLP lookup away from the true answer.”

Understanding Recursion Levels in AI

28:04 to 29:18

Learn about the structure and significance of recursion in AI models.

“We go from XRAW to X, which is the maze state or whatever it is, initial maze state.”

Training and Testing Loops Explained

29:18 to 30:52

Discover how training and testing loops function in AI recursion.

“So that's two of the three recursions you said.”

Deep Recursion and Optimization Techniques

30:52 to 32:52

Explore optimization techniques in deep recursion for AI training.

“So it's in a different part of the latent space.”

The Future of AI Models with Recursion

32:52 to 34:19

Discuss the implications of combining recursion with larger AI models.

“And then it outputs, and then you're good to go, and you train it exactly as the same way before.”

General vs. Task-Specific Models in AI

34:19 to 36:19

Understand the differences between task-specific models and general-purpose models.

“Even Alexia is limited by backprop through time.”

Efficient Architectures and Reasoning in AI

36:19 to 37:44

Learn how efficient architectures can enhance reasoning in AI models.

“The model trained to do Sudoku cannot do ArchPrize inherently, it has to be trained on the price set to do so.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Francois Chaubard:Welcome back to another episode of Decoded. Today I'm back with YC visiting partner Francois Chaubard to talk about one of the most interesting recent trends in AI research, recursion. Specifically, we're going to talk about how we can improve a model's reasoning performance by using recursion at inference time rather than by just making the model bigger and bigger. There were two papers that made the power of this approach really clear in 2025. One on hierarchical reasoning models, or HRM, and another on tiny recursive models, TRM.

0:35Francois Chaubard:Francois, thanks for joining us. Can you tell us a little bit about these two models and what was so interesting about them?

0:42Ankit Gupta:Sure, I guess to set up a little bit of a foundation, you already did an amazing lecture on RNNs and LLMs in one of the previous videos. I won't overdo it, but just to give the cliff notes. An RNN is just a model that you recursively call again and again and again on itself. And we were very much in the belief that this was required to get to AGI peak RNN use was probably until 2016 with Alex Graves NeurIPS keynote, which is just fantastic, and all his adaptive compute time work.

1:15Francois Chaubard:So this is about 10 years ago, people who were working on these models. This was in the era of LSDMs and LSDMs with attention.

1:21Ankit Gupta:Yeah, and depending which professors you talk to, before attention was invented.

1:26Francois Chaubard:Yes, yes, totally.

1:26Ankit Gupta:Yeah. And I think what really was the limiting step on RNNs in general was this thing called backprop through time, where you have to roll out the model. And then to update the weights, you need to approximate the gradient, and you step back, back, back, and you keep rolling out. And as the model gets bigger and bigger, and as you roll out for more and more steps, then you have all these accumulation of errors, and the gradient gets noisier and noisier, and then it just kind of stops to work.

1:54Francois Chaubard:Yeah, so you have these vanishing or exploding gradient problems. And it's because if you have an input with 20 steps, you're like multiplying these matrices 20 times and that causes training.

2:01Ankit Gupta:And we're talking about doing context length of like a million or like a billion. And so it's not even just 20, it's like a billion. And even worse, you have to retain the activations at every single step. And so like if this were happening in your brain, you would need like a million copies of your brain at every single activation so that I can back prop through it. There's tricks around this that you can do and you can do a gradient checkpointing and things like that to reduce that issue. But then you just like trading off memory for wall clock time and compute.

2:29Francois Chaubard:Right, so now if you could pass that with LLMs, the ones that people are widely using. These, while at face value they appear to be similar, at training time, they're doing basically this one shot feed forward process for every input. Right, the LLM, the transformer block can take all of the inputs in parallel. It's not actually iteratively going over them one at a time at train time. So you don't have this needing to store tons of activations problem or this giant vanishing gradients problem with them.

2:55Ankit Gupta:Yeah, exactly. Like it's actually like all happening in time in one shot magically. And that was like the trill or lower triangle trick that kind of happens, this causal mask that occurs. And so you actually do all time steps in one shot and you forward pass a feed forward model on all time steps in one shot and you backwards in one shot. And it's amazing for train time in terms of like wall clock. It requires a lot of flops and and it still requires a lot of the memory, you still need it there, but you don't have the vanishing gradient issue. And what you actually paid for that you have to give up is this latent reasoning thing and this compression in the time direction.

3:34Ankit Gupta:There is no compression in LMs. Every single decode that I do, I still have to retain the entire Shakespeare novel just to decode a little bit. And in RNNs, you don't have to do that. It's all compressed in this hidden state that you kind of roll out.

3:46Francois Chaubard:Okay, so let's talk about that in a little bit more detail. Like, you refer to this inherent reasoning ability. You know, many people think about LLMs as doing reasoning, and we're going to talk about that a little bit later. But help me understand where you see the biggest limitations in LLMs reasoning ability is in terms of what the model does in an actual forward pass.

4:08Ankit Gupta:Yeah, and so I guess we go back to chat GPT-2. GPT-2 was this landmark architecture and paper that basically was just get next token, next token, next token, and it kind of worked. And we just watched val loss go down, perplexity goes down, like the model is more performant, looks better, starts to make some Shakespeare that actually sounds somewhat plausible. Right. And then we have to get these things to reason and to actually solve some really hard problems. And I've done extensive experiments on this, but if you take, for example, You have infinite amounts of unsorted lists and you give it sorted lists.

4:46Ankit Gupta:You keep feeding it to the model, it should work, right? It's actually impossible for the model to map from unsorted list to sorted list. If I have - In a one shot, basically. In a one shot basis. It's like literally that we know a theoretical lower bound that for comparison sort you can't do better than n log n steps. And if I have a list that's 31 characters or elements long and my transformer is 30, I run out of steps to do comparisons. It's not possible for me to do all the steps that is needed to be done. In HRM and TRM, they use Sudoku as an incompressible problem. Similarly, and so are mazes, those are incompressible problems.

5:27Ankit Gupta:Rolling sum, incompressible problem.

5:29Francois Chaubard:So when you mentioned the sorting algorithm, when I think back to my algorithms class from college, the one way you could get faster than n log n in a sorting algorithm is if you had some access to an external memory cache. If you had some tape you could write to, then you can actually do faster than n log n by basically selectively putting things onto this memory. And I suspect that's a key limitation of these LLMs in that because there's no external memory tape inbuilt into the model, you lose certain performance possibilities in terms of how fast you can go.

5:57Ankit Gupta:That's right. And so I guess like a rate extort would be like the most common depending on the number of buckets that you have, you can kind of get from n log n to order n. You can't get less than n, you have to touch all the elements. Sorry, you have to do that. And if you run out of layers and transform layers in your neural network, then you ran out of chances to do that.

6:20Francois Chaubard:So this is just like going back to Alan Turing now and like a Turing machine, right? So what's the analogy there exactly that we should think about in terms of LLMs, I guess not quite satisfying how you think about a Turing machine?

6:30Ankit Gupta:Yeah, so if we let's just talk about like chat GPT to GPT to the original like no bells and whistles It's just a feed-forward model. And so just forward passing one step and

6:41Francois Chaubard:Taking an input creating a bunch of outputs and the Sudoku case

6:45Ankit Gupta:If I have 50 different Squares it's provable that I can only do one given this information then and I have this many layers then that's all I can do. And the cheat is the chain of thought. And so it's completely true that at test time, they are turn complete. And you can simulate all turn computable functions at test time, but how do you get it to learn it? You need to train it. And that's where, unless you're training it on human labeled traces, for which there's a lot of problems like the millennial prize problem, we don't have the trace for it. So we'd love to have the trace for it. It doesn't exist.

7:23Francois Chaubard:Totally, makes sense. Okay, so with that context in mind now, let's talk about these two papers, because I think that sets up a lot of the contrast we're going to draw between these papers and the models that people are maybe more used to. So let's talk about HRMs first. Walk me through a little bit about how this model works and some of the intuition behind it.

7:43Ankit Gupta:Sure. So this is directly in the lineage of RNNs. There's not that much novel from the RNN standpoint, at least in my opinion. they do have this idea of, you know, inspired by the brain, where I have, like, there's different parts of the brain that operate on different frequencies. There's some that operate at a really high frequency, which is on the low level of the hierarchy, some that operate in a really low frequency, which is the higher level of the hierarchy, and the interplay between those things is really interesting.

8:13Francois Chaubard:So this is, like, literally in the human brain, there's some, like, bio-inspiration here, which is that, like, you have, like, different waves running at different frequencies at different parts of the brain or something like that. Okay, cool.

8:22Ankit Gupta:And I guess that's one interpretation of it, of the way that they're talking about classifying these hierarchies of frequencies. And the most interesting part, at least for me, is the way that they train the neural network. You take in some X, some input, whether it's an incomplete Sudoku puzzle, a maze, or an art prize challenge. You do TL steps with the lower level module. Then you do, to go to H, you do that TH times, and then you have N sub outer refinement steps.

9:03Francois Chaubard:So you basically are like running through the input with a given matrix, with a given transformation repeatedly on it, and you're doing that through two levels of refinement, and then basically running that process several times.

9:15Ankit Gupta:Yes, so there's exactly three levels of recursion occurring here. There's the low level, there's the high level, and then there's the outer refinement steps.

9:22Francois Chaubard:And we're calling it recursion because it's the same weights that are being applied repeatedly. We're not changing the weights in between these steps.

9:28Ankit Gupta:Exactly right. You get to recurse on the Lnet, TL times. You've recurse on the TH and the TL, this looped recursion, TH times. And then you do N-sub, you do this whole outer refinement step N-sub times. Cool.

9:42Francois Chaubard:And so what's the basic intuition for why that works? Like, why does that produce an effective paper result? And what even were the results that this paper showed?

9:50Ankit Gupta:Yeah, and so, I mean, this got state of the art on ArcPrize 1 and 2. This was only a 27 million parameter model that was only trained on ArcPrize.

10:04Francois Chaubard:So it's like 1 ,000 inputs or something like that, like puzzles basically.

10:06Ankit Gupta:Yeah, literally 1 ,000 tasks, which is extremely small. There is no pre-training at all. This starts from like literally Tagula Rasa weights and it can outperform at that time if we go back You know we had oh three if you remember back way back when And it drew me and it oh three gets zero Literally zero and this got like something like 70 % on our prize one at least At the time which was just a huge breakthrough And so kind of the the way you can kind of think this is like variable scoping and so like if I have like three nested functions. I guess the lowest level function has scoped variables, which they'll call ZL, which is the carry that inits the zero.

10:48Francois Chaubard:A latent variable. Latent variable.

10:50Ankit Gupta:In traditional RNN literature, they would call this the hidden state, the low level hidden state. I get to recurse, recurse, recurse, and then I pass back that ZL back to the outer scoped function, the higher level one. I let that one do one iter. it goes back and calls the lower level again, it does this whole thing in a third Adder loop which is called Adder Refinance Step.

11:12Francois Chaubard:But when you describe it like that, it seems like it would have the same backprop through time problem that you would have at one-on-ends, and I think they came up with a clever trick to basically get around that. So like what was that trick that they figured out?

11:22Ankit Gupta:And this is really the crux of the paper that like differentiates it in my opinion in the literature is they instead of doing what Alex Graves did in all of his papers from neural turning machines to adaptive compute time to differential neural computers is he always back propped through all of the recursion steps and he was limited by back propped through time so you could only make the model so big you have all these issues vanishing gradients etc etc and what they do is they they kind of have this uh deq uh of method of doing fixed point iteration sounds like deep deep equilibrium learning where if I take a batch, and this is completely counterintuitive as a computer vision person, because you'd never do this, but it actually does make sense.

12:11Ankit Gupta:And I'll explain why. If I take a batch of like ImageNet or CIFR10, and I forward pass through the model, and I get some loss, and I back prop, and I update the weights, I would go get a different batch for the next one. But what they do instead is they actually do that 16 times. And so, and as you do that, you actually can see the change in your residuals get less and less and less. And why it actually makes sense is because when in the RNN case, the ZL and the ZH, which are the carry, the task carry start out - The hidden states. The hidden states start out at zeros. Those are zeros. Then we go through this whole loopy recursion, at least the two loops, the two lower loops, the TL and TH steps.

12:51Ankit Gupta:And then I back prop just through the two modules, just once. And I don't recurse all the way back. I do a stop grad, I stop right there. And then there's a huge residual, and then I don't reset ZL and ZH. I do it again at a different point in the carry or hidden variable space. And so one can actually look at it as a different batch every time, even though it's the same exact x's.

Read the full transcript

13:19Francois Chaubard:Yeah, like the way I kind of think about it is like the 16 or whatever that you're recursing over, it's like constructing a mini-batch, not from different inputs, but from like different memory states basically. It's like across this hidden or carry memory access basically.

13:38Ankit Gupta:And that math holds and it works. It follows DEQ directly in the event that the delta in ZL and the delta in ZH go to zero, which it actually doesn't do. And so we'll get to TRM, but Alexia basically shows that it's just not the case, and you can't actually apply this math. And that's why it's working. That's not sufficient support for why it's working. We actually don't know why it's really working. and she figures out that you actually can back prop through all the way to the deep recursion, which we're going to get into TRM in a second, and that actually improves performance much, much more.

14:17Francois Chaubard:Interesting. Okay, so before we get into TRM, yeah, on this paper, I think there's a bunch of different ways people have looked at this, right, in terms of how they came up with it and then why this may or may not be working. One, it's a sort of bioplausibility argument. As you know, I'm usually not super keen on these. I think machine learning tends to have a long history of people starting with bio plausible arguments and then Realizing that there's some variant of them that seems highly bio implausible that actually works better I think you have example

14:44Ankit Gupta:Yeah, classic the first deep learning paper that started this whole craziness is AlexNet and in AlexNet there's actually this funny little thing called like local receptive activation or depression or something like that where like once this activation fires, then I have this like, you know, a refactory region or something like that. It actually doesn't work at all. And like, it didn't work and you didn't need that. And then VGG came out and said, get rid of all that, just go deeper and three by three conv. And it actually just like outperforms dramatically. And so like, this is like, always the, maybe you need to do it to get accepted into NeurEPS and have some biological possibility.

15:18Ankit Gupta:Yeah, sure, totally, totally, yeah. You're definitely the expert here, but what do you consider to be bio-plausible and what's not?

15:24Francois Chaubard:Well, I think that a lot of machine learning literature has overlapped a lot with people working in neuroscience. I think it is very natural for us to ask questions about how does our brain work, because our brain is like an incredible instrument that does a ton of computing, obviously, and does it in a very shockingly efficient manner, it seems like. And so a lot of machine learning research has for a long time sought analog from how we think to understand our brain to work and try to encode that in various machine learning systems. So from the very basic concept of what a neural network is, it's called a neural network, because we think it's some basic model for what a neuron is.

15:59Francois Chaubard:How certain activation functions work are meant to be inspired by certain biological premises. Do you think that's a misnomer? The thing about them is that often we use bioplausibility to inspire us to come up with ideas, but we end up veering away from the bioplausible to something adjacent to them that is likely bioimplausible, but that seems to work better.

16:18Ankit Gupta:Something that runs better on a GPU.

16:20Francois Chaubard:Exactly, it runs better on a GPU, it's more efficient in some capacity that is relevant to how we actually encode it in a computational system. So I find thinking about bio-plausibility fun and interesting, and it's definitely a great way to inspire us to think about new things, but I tend to not be bounded by bio-plausibility when I think about what machine learning systems we should prioritize working on or think as particularly exciting, other than as an interesting scientific launching point for a deeper exploration. I think the version of this that I find more compelling is actually that original discussion we were having around automata theory, basically.

16:52Francois Chaubard:And honestly, just actually like fundamental data structures and algorithms theory, which is that if you're running a complex algorithm, having access to sort of a memory cache is actually very useful for being able to run that algorithm efficiently. And I kind of think of this set of hidden states or carry as akin to a Turing machine tape or akin to the radix sort memory bank, where you can basically train a model to use this memory cache in an intelligent way in a single forward pass so that you can get a more efficient time operation that would otherwise require some sort of more complicated reasoning.

17:27Ankit Gupta:Yeah, I think that a point I wanted to make earlier is that like we did this COT stuff and this tool use thing as ways to get beyond the limitations of GPT-2. And so the way that we get, I've done this experiment. You can actually, if you give me infinite amounts of un-sorted list and un-sorted list, if I can do chain of thought, and I can do every single step and teach it to do every single step, then I can actually get it to do sort and become a Turing machine at test time. And similarly, an even cheaper one that is much easier to do is you teach it and you say, hey, there's this Python function called sort.

18:10Ankit Gupta:Just call the function. Just call the function. And like that's the easiest thing to do and you don't need backprop at all. And so those are the two hacks. Now, well Francois, this is solved. Like we're done, right? No, because I needed to know what sort was. What happens if we didn't know what merge sort?

18:25Francois Chaubard:The chain of thought is not going to inherently discover sorting from first principles. It's finding it from our historical knowledge of everything it's trained on.

18:33Ankit Gupta:Yeah, I mean this is like the, the Demis had this whole thing about like the ultimate test is the Einstein test. like go back to 1911 and then like have it rebuild all the physics up until now. Similarly, let's just pretend that we only had bubble sort. We knew other no other sort system. If you chain of thought it on all the bubble sort input and output it will only do bubble sort. In fact it won't even do bubble sort that well. So this is the best situation and then the tool use of course it can only know bubble sort. I want to get to merge sort. How do I discover merge sort?

19:00Francois Chaubard:And I think the interesting thing just to emphasize here because it may not have been extremely clear is there already exists some type of recursion that people are used to in LLMs, which is chain of thought we mentioned earlier. But that is a recursion that's happening in the token space of the models outputs, not inherent to the model itself. That's sort of the fundamental limitation is that the model can only do a feed forward one shot output. And then we basically just have this hack that if you keep letting it output of things, then it can read its outputs and do somewhat intelligent seeming things with it.

19:37Francois Chaubard:But it seems to sort of be upper bounded by the data that we feed it that the labs are very hungrily buying right now and not the sort of inherent underlying recursive reasoning.

19:46Ankit Gupta:Yeah. So in both cases, both hacks to solve this in COT and tool use, you're bounded by the bounds of human knowledge. In the event it's outside the set of human knowledge, then you're kind of SOL. And so that's one. The other, you make a great point about discrete versus latent space. Reasoning in a discrete, it can only output the carry in the case of LLens, has to be snapped back to some discrete token space. And in the case of RNNs in general, they remain in this continuous latent space, which is much higher dimensional. If you give me like a tape that's this long and you cut it up into 10 buckets, like versus all the possible values.

20:30Ankit Gupta:Right, exactly. It's much more expressive to be in continuous space. But we can't train it that way because we actually, you know, because you're inhibited by backdrop through time, largely. And this is why this paper is so exciting.

20:40Francois Chaubard:OK, so before we then go over to the TRM paper, let's just summarize here. What matters most from the HRM paper that we should take away before we transition and contrast it with the TRM paper?

20:52Ankit Gupta:Yeah, I think that the number one piece to take away is this outer refinement loop. The outer refinement loop scales. And there's a great breakdown. Basically, the Sapien authors, which is huge kudos for this paper, because there's so many innovations in this paper, didn't really do like scaling ablations on every single one of the inputs. But this guy, Constantine, at François Chalet's company, India, actually did. And it's this amazing breakdown that he posted on YouTube that you can go check out. But basically the main takeaway is that the outer refinement loops is the main beneficiary is the main reason why these things work so well which Alexia basically takes the she found I think in parallel and And scales up and and shows that you can get rid of a lot of all this other stuff

21:46Francois Chaubard:Okay So like a lot of machine learning the follow-on paper is basically delete 75 % of the first paper as we've often done in videos here and keep the magic basically. Yeah. So okay, so what's the magic then? Like what's the part that actually matters in terms of what stays in the TORM paper? And let's not contrast the core architectural differences between these two papers.

22:05Ankit Gupta:Yeah, so I think that I guess if I break it down into two major things. This outer refinement loop thing is really great and works really well. And that this like truncated backprop through time, which is backprop through time except I truncate at some time. Some earlier earlier point. called t, t back, t equals one is actually completely sufficient. And so truncated back row up to time t equals one, completely sufficient. And that's very counterintuitive. Which is what HRM found. Which HRM found and TRM does a little bit further, rather than going through just one call to the H net and the L net, it actually goes through one full recursion loop.

22:45Ankit Gupta:So if I do it 16 times, I just go back through one time and that is kind of sufficient. And if you do it with this fixed point iteration thing, pseudo fixed point iteration thing, where you keep hitting it with gradient at every single step, it weirdly works. And this batch size across the carry space actually works.

23:07Francois Chaubard:So that part is also kept between these two models. It seemed like another thing that changed was having this sort of double layer of higher order thinking and lower order thinking. It seems like it collapsed it down into just a single one. What's the intuition there and how does that actually work in the TRM paper?

23:25Ankit Gupta:Yeah, so it's interesting. She actually ablates having two separate networks versus just having one. I guess the more important space is the variable scope, is that you should have low level features and high level features, but the same network. And so the best performing model - The same network can extract both, basically. Yeah, you weight share between the Lnet and the Hnet, and it's just called net. And you do just one transformer layer versus the four like they do in Sapient, and just whittle it down to one and do more of a cursion. But you keep ZL and ZH to be distinct and separate. And she calls it X and Y, which I found very confusing.

24:00Ankit Gupta:X, Y, Z, which is very confusing. And it's just like ZH and ZL is just cleaner.

24:04Francois Chaubard:So if you read the paper, Y is actually like latent space. It's like Z, basically. And it is not a label.

24:10Ankit Gupta:Yeah, okay. Which really threw me through. But anyway, so we'll go through some code here, and I'll walk you through it. So I replaced all of her nomenclature and used the sapient notation, which is much cleaner and more straightforward to me at least.

24:23Francois Chaubard:Okay, cool. Now, before we dive into the code for a sec, like in terms of how these TRMs actually work, it's pretty interesting because this recursion advantage now gives you a bunch of advantages over transformers. Rather than having 500 or 1 ,000 or a million or whatever transformer layers and having tons and tons of parameters, you get compute depth basically without this parameter depth. Mm-hm, right. And the optimization process looks like more of like an iterative, kind of like expectation maximization algorithm. You want to talk about how that worked in the TRM paper? Because I thought that was also pretty interesting.

24:57Ankit Gupta:So both of them kind of have the same kind of EME feeling thing, where like we update ZL condition upon the input x and ZH, the last ZH, ZH T minus one, let's say. And then we keep updating ZL, ZL, ZL, ZL, ZL, and we keep updating it. And then we go holding, we update ZH, condition upon ZL, and actually it's just ZL, it's not even X. And then we just update ZH. And the way to think about ZL and ZH is ZL is like your local scoped variables that are just being overwritten and updating, updating, updating. and then ZH and Alexia makes this point, I'm sorry, Alexia makes this point, that is a candidate answer, a proposed latent answer that is just an embedding space away, a one MLP lookup away from the true answer.

25:54Francois Chaubard:So you're kind of like, yeah, I mean, just to like zoom out a little bit, you're kind of maximizing the probability of the correct information stored in your memory, conditioned on a given output and maximizing the right output conditioned on the information stored in your memory quote-unquote in parallel. And like that optimization algorithm leads to you ultimately learning a recursive method that stores the right information to this local memory basically and then outputs the right thing.

26:25Ankit Gupta:It really like if we actually think of Sudoku, it's actually a really natural way to think about what's actually happening on the hood where Sudoku is an incomplete puzzle. You can't guess every cell at any one time. Actually, it's designed where you can only guess one or two cells based on the available information. So it's an incompressible problem. You actually can't do it unless you're just randomly guessing and guessing and guessing, which is a very high combinatorial space. And so what the ZL is doing is some type of, let me try this, try that, do some computation, think about local things, and then it proposes.

26:57Ankit Gupta:And then we go to condition upon something that it may have found. It sends it to ZH, ZH fills it in and now we have a little bit more of a filled in Sudoku puzzle. And the training process is training the algorithm to know to do that, right?

27:12Francois Chaubard:It's like it's maximizing that it's like oh this strategy for what you save tends to lead to correct outputs.

27:18Ankit Gupta:Without chain of thought.

27:19Francois Chaubard:Without chain of thought. That's the most important part.

27:21Ankit Gupta:If we had Sudoku and we know how to solve Sudoku because like we were just you know dumb homo sapiens that didn't know how to solve Sudoku, like it would just have solved it. And that's why it's cool because it actually is able to discover things without being teacher forced via chain of thought. Right, interesting, yeah.

27:37Francois Chaubard:Should we look at some code? Let's do it. Okay, let's dive in. And I would love to see what these papers or models look like, just distilled down to their core essence. I know there's lots of details on how you train them, but kind of the core training algorithm, and it'd be great to contrast the two methods. Yeah.

27:52Ankit Gupta:So, I mean, they're remarkably similar. And so, largely one and learning one is learning the other. But basically, you start out with some ZH and ZL that are just zeros. Yep. You have some input embedding space. We go from XRAW to X, which is the maze state or whatever it is, initial maze state. And then with no grad, you don't pass any gradients back through this. You - So, this is the trick, basically.

28:18Francois Chaubard:This is the trick. To not backprop through time.

28:20Ankit Gupta:Here are two of the three recursion levels. So yeah, this is like the, they do this just for simplicity, but I hit ZL, T low times, and then once for modulo, T low, then I hit the ZH and I do it again and again. And like you said, I'm updating ZL condition upon ZH and X. Right. And then I update ZH condition upon ZL. Right.

28:45Francois Chaubard:So this is the expectation maximization style. Exactly.

28:48Ankit Gupta:And then you don't really need this, this is just for cleanliness to show clearly that there's no gradients occurring above this line.

28:56Francois Chaubard:Just freezing the weights past that.

28:58Ankit Gupta:Exactly, and then I hit Lnet and Hnet one more time.

29:01Francois Chaubard:Which is the same thing as up above. So this is just, it's literally just the no grad thing running one more time. Exactly. Cool.

29:06Ankit Gupta:Yeah, and just make it really clear. And then there you go. And that's your HRM model.

29:11Francois Chaubard:Cool, that's quite simple.

29:12Ankit Gupta:Two and two is completely sufficient. If you actually go much higher, Konstantin showed very clearly that it doesn't actually help.

29:23Francois Chaubard:So that's two of the three recursions you said. The third happens in the actual train loop.

29:26Ankit Gupta:The third is in the train loop and at the test loop. They both have this M test or N supervision, which Alexia calls deep supervision. They call it adder refinement steps. It's just whatever you want to call it. Call it N sup.

29:39Francois Chaubard:And so you do this n sub times during training and then during test time there's a different hyper parameter for how many times it recurses over each model which is m test. They're actually the same. Okay.

29:51Ankit Gupta:And so this and this we can probably just call this the same. Yeah. And but it's the same. And if you actually, Constantine does a good job of this. If you actually train on 16 and you test on only one, you get like 7 8ths of the performance, or like almost all the performance. So it's actually quite interesting that this is just too much compute, and it doesn't actually help you all that much. So setting this to one is actually like pretty much -

30:21Francois Chaubard:But presumably for like more complicated problems, having more test time compute is still useful. It's like the reason you would set it up this way. Yeah, for sure.

30:28Ankit Gupta:And so we call our HRM, we get some loss, We backprop through just those two little parts here, and then we step. We zero out the gradient, but we do not update ZH and ZL. These are still the same in it. So that's the really important detail there. And then as we go back, we pass in the ZH and the ZL from the previous one. So now this is actually not the same batch. Because we have updated ZH and ZL. So it's in a different part of the latent space.

30:59Francois Chaubard:Cool. That's the key like mini batch construction through memory space concept.

31:05Ankit Gupta:Yeah, exactly. And then at test time, it's simply the three loops. So there's your outer refinement loop, which turns out just at train time. Mostly doesn't matter. Train time recursion was important, but test time recursion was actually not that important, which is kind of counterintuitive. And then the HRM inside that has your two other loops. Makes sense. And that's it. So pretty simple. Now the TRM. The only two changes, the main two changes here, is that they collapse Lnet and Hnet into just net. Great. And it's important detail, these are four transformer layers, this is four transformer layers, and this is just one transformer layer.

31:41Ankit Gupta:And Alexei actually shows that going deeper actually didn't help.

31:44Francois Chaubard:Yeah, and actually on some tasks, it was just the feedforward net actually worked just as well as a transformer there, right? It was like on Sudoku, I think.

31:50Ankit Gupta:Yeah, on Sudoku, MLP actually outperformed the tension. It scored zero on the maze. The MLP scored zero on the maze. And so it's not clear, it's not obvious that the transformer is always better. So there's the weight sharing. And then instead of going back just the one, two, this back propping through just these two, you actually back prop through one latent recursion step, all the way through one latent recursion step. So let me just walk through this a little bit. So we have the same thing here. Same starting point, yeah. It's mainly the same thing here. We're doing this six times. And then we go one more time here.

32:30Ankit Gupta:And then we do our deep recursion. This is the outer loop, n sub times. And so again, we have the no grad, we have the detach. And then this is where it's different. So I am calling this latent recursion after the detach.

32:45Francois Chaubard:Yeah. So it's one full recursive loop is happening. Versus here.

32:49Ankit Gupta:And so that's the main difference in the optimization. Otherwise, it's effectively the same. And then it outputs, and then you're good to go, and you train it exactly as the same way before. And then at test time, it's the same thing again. And so largely the same. Cool.

33:06Francois Chaubard:And so in many ways, it's sort of a simplification, right? You're collapsing certain parts of it. You're simplifying this net architecture. It's slightly more complicated along this backprop through time part because you're actually back propping through more than you did before. But it's like taking a bunch of lessons from the first one and basically simplifying most of it.

33:25Ankit Gupta:Which is actually why she needs, I think, is why she needs to make the model smaller. And so it's a 28 million parameter model for HRM. Now she brings it down to a 7 million parameter model. It actually gets from 70 % to 87 % on ArtPrize 1 and does actually quite well on ArtPrize 2 as well. And so, yeah, so she makes the model, you know, three, four times smaller, but because it has that recursion, it actually outperforms. And there is one, there is this researcher named Melanie Mitchell that writes this book talking about this very phenomenon, which is like, it is sufficient, not necessary to go bigger and get better performance.

34:08Ankit Gupta:And it is sufficient and not necessary to add more recursion. And so where I'm really excited is what happens if you do both. Right. And you're still limited by backprop through time. Even Alexia is limited by backprop through time. That last step from a memory perspective for sure. And so if you can make the model really big and you have lots of recursion and we do something else other than backprop through time, then we can get all the benefits of this and all the benefits of the giant LLMs and then you can get some crazy stuff.

34:39Francois Chaubard:So now to wrap up, why don't we talk a little bit about the bigger picture. What does this mean for the field of AI research? How should people think about where these models fit into the current span of research happening, especially given that it seems like a bit of a departure from a lot of the methods that people are used to hearing about and increasingly seeing products that people use?

34:57Ankit Gupta:Well, I think for one, from the arguments that Schmidhuber makes and that we've talked about today, recursion is important, and it's not going away. And clearly the benefit is here of adding recursion into models. And you've seen things like the recursion language models out of Google that are pretty powerful and cool. And so that's definitely one piece that's, I don't think, going away anytime soon. The next one is this adder refinement loop, like tbtt, t equals 1, truncated back wrap through time t equals 1. I think that that is a really powerful idea. And the fact that that works so well, we have yet to really explore that extremely really understand what's happening there and then the third is that idea of like okay we know that recursion works we have these tiny recursive models that are 7 million parameters it can solve really small a hundred million hundred billion hundred probably a trillion trillion parameter model can't solve trained on the entire internet and a 7 million parameter wins like the right answer is to like take the amazingness here and take the amazingness here, which probably is already in Gemini already or some of these, it might be at least in some part.

36:08Ankit Gupta:But when you take the benefit of both these TRMs and these giant models and you actually slam them together, I think that it's just going to take off and it's going to be really huge.

36:18Francois Chaubard:Yeah, one of the things that's really interesting about these TRMs and HRMs is they're not general purpose models, right? These were task specific models, right? The model trained to do Sudoku cannot do ArchPrize inherently, it has to be trained on the price set to do so. Versus the LLMs that are used on these tasks are general purpose models that maybe get some additional fine-tuning data or in-context learning data on those tasks. And so I think that's where the interesting overlap might come is if you can make these more general purpose agents that can somehow be general purpose in the way that the sort of next token prediction algorithm has given us and do more complex reasoning to achieve that.

36:52Francois Chaubard:It seems like you can have really efficient architectures to do scale up reasoning.

36:57Ankit Gupta:Right, a lot of the view of what these LMs are doing is finding really amazing embedding representation spaces, but reasoning inside that space is actually not done all that much.

37:09Francois Chaubard:Yeah, it's always through the token space.

37:10Ankit Gupta:It's always through the token space, and so like what you can imagine is we found mapping from token space or from vision, from pixels, some really cool latent space where like things are just nicely semantically separated and we can, you makes it really easy for downstream tasks to do. But now in that space, use this, like tiny reasoning models, use some type of recursion inside that and train that model on that, a little small model on that reasoning space. I think that's really gonna work.

37:40Francois Chaubard:Francois, thanks so much for breaking it all down for us. See you all in the next episode of Decoded. Thank you.

From the publisher

A 7-million parameter model outperforming models a thousand times its size on tasks like ARC Prize. That's what recursive reasoning unlocks.In this episode of Decoded, YC's Ankit Gupta and Francois Chaubard break down two recent papers on recursive AI models, HRMs and TRMs, that are achieving state-of-the-art results with a fraction of the parameters of today's largest models.They explain why standard LLMs hit a fundamental ceiling on certain reasoning tasks, how recursion at inference time gives small models the compute depth to break through it, and what happens when you combine these ideas with the power of large-scale foundation models.

More from Y Combinator Startup Podcast

All 148 episodes
Beyond Bigger Models: Recursion As The Next Scaling Law In AIY Combinator Startup Podcast · 38 min
Listen in VO