Learning to Reason in 13 Parameters

11 Feb 2026 · 19 min · 9 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The episode explains “Learning to Reason in 13 Parameters,” arguing that a 7B-scale language model can dramatically improve math reasoning by updating only 13 parameters (TinyLoRA), using reinforcement learning rather than supervised fine-tuning.

Guest backgrounds

No guests are named; the host leads the discussion.

Key claims

Full fine-tuning or even LoRA typically trains millions of parameters, but TinyLoRA steers reasoning with ~26 bytes of FP32 parameter data. Weight tying shares a “master key” across the model, with tiling by layer depth outperforming module-based sharing. The method works because GRPO reinforcement learning rewards only correct final answers, removing stylistic noise that supervised fine-tuning must imitate. Larger base models may require smaller steering updates; Quen is more steerable than Llama.

Notable examples

GSM8K accuracy rises to ~91% after training 13 parameters; harder benchmarks (AIME, Math500) retain ~87% of the gain with 196 parameters. The episode contrasts Quen 2.5 vs Llama 3, noting Llama “flatlines” under ultra-low-parameter updates. It also warns about security/alignment risks if tiny steering files can be swapped on edge devices.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The 13-Parameter Breakthrough

1:42 to 3:58

Discussion on how a tiny update of 13 parameters can drastically improve AI reasoning.

“We have the data, we have the benchmarks, and we need to unpack how this works, why bigger brains might actually need smaller steering wheels, and what this all means for the future of AI on our devices.”

Revolutionizing Fine-Tuning Methods

3:58 to 6:40

Explaining how traditional fine-tuning methods are inefficient and how new techniques improve efficiency.

“Right, because you aren't squeezing the rules of algebra into the numbers.”

Understanding Weight Tying in Neural Networks

6:40 to 9:33

A deep dive into the technique of weight tying and its implications for AI model training.

“SFT supervised fine-tuning is the standard way, right?”

The Impact of Reinforcement Learning

9:33 to 11:15

How reinforcement learning provides a cleaner signal for model training, stripping away noise.

“The 13 parameters are just the switch that turns the light on.”

Scaling Model Size and Efficiency

11:15 to 14:00

Discussion on how larger AI models can be controlled with smaller updates, challenging traditional scaling laws.

“But, and there's always a but, not all brains are created equal.”

Understanding Tiny Laura's Parameters

14:00 to 14:48

Learn about the parameters behind Tiny Laura and its impact on reasoning.

“The American Invitational Mathematics Examination, or AIM, is another.”

Limitations of Tiny Laura in Creative Tasks

14:49 to 15:14

Explore the restrictions of using Tiny Laura for creative writing and subjective tasks.

“This works for math, coding, logic, domains with a verifiable truth.”

Implications of AI on Edge Devices

15:15 to 16:16

Discuss the potential of running AI on edge devices and its impact on personalization.

“What does this actually mean for us, the people holding the phones?”

Safety Concerns with Fluid AI Personalities

16:17 to 17:36

Examine the safety implications of easily alterable AI models and their personalities.

“You don't need a massive GPU cluster to share your custom model.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Okay, I want you to visualize the most complex digital structure we have right now. A large language model. We're talking hundreds of billions of connections. A neural network so vast requires a warehouse of GPUs and, I mean, enough electricity to power a small city just to exist. It's a digital leviathan, a web of parameters so dense we can barely even map it. Right. Now, imagine you want to teach this leviathan a new skill, let's say. You want to turn it into a math genius. Standard logic, and honestly just common sense, says you need a massive update. You need to rewire millions, maybe billions of those connections to get the information in there.

0:39That's the standard logic. Usually that's exactly what it takes. You want a big change in behavior. You make a big change to the brain. But today we're looking at something that just breaks that logic completely. We're looking at a method where you can take a 7 billion parameter model and drastically improve its math reasoning by changing 13 numbers. Yes, 13. Not 13 million. Not 13 ,000. 13. Which, if we're talking file size, is roughly 26 bytes. 26 bytes. That is smaller than a single text message. It's smaller than the empty space in a text file. It's smaller than the file icon sitting on your desktop.

1:15And yet, this microscopic update drove a model's math accuracy from mediocre to over 90%. It sounds impossible. Effectively, like magic, but it's real. And what's fascinating isn't just the efficiency. it's what this 26-byte file tells us about how these models actually think. We're diving into TinyLaura and how reinforcement learning allows us to steer these massive intelligences with a rudder the size of a grain of sand. So welcome to the Deep Dives. We have the data, we have the benchmarks, and we need to unpack how this works, why bigger brains might actually need smaller steering wheels, and what this all means for the future of AI on our devices.

1:52It's a complete inversion of how we usually think about training. Let's get into it. I want to start with the paradox here, because usually when we talk about fine-tuning AI, it's a heavy lift. Give us the baseline. If I'm a developer and I want to teach a model like Llama 3 to be, I don't know, a doctor or a lawyer, what do we usually have to do? The old school way. You do full fine-tuning, you take the base model, that pre-trained giant that knows a little bit about everything, and you update every single parameter based on your new data. Every single one. Billions of numbers changing. It's slow, it's expensive, and it requires massive hardware.

2:26It's like rewriting an entire encyclopedia just to add a new entry on cardiology. Which is why everyone started using Loray, right? Low-rank adaptation. That's been the buzzword for a while. Exactly. Loray was the first massive efficiency breakthrough. Instead of rewiring the whole brain, you freeze the model, you don't touch the original encyclopedia, and you just add a few small adapter layers on top. Ah, okay. So think of it like adding sticky notes to a textbook page instead of reprinting the whole book. Perfect analogy. But even those sticky notes add up, don't they? They do. Even with standard L 'Oreal, which we consider efficient, for a model like LAMA-38B, you're typically training about 3 million parameters to get good results.

3:06Okay, so the industry standard for efficient is 3 million numbers. And this new research comes out and says, we can do it in 13. That is just a staggering difference. It is orders of magnitude smaller. The research is titled Learning to Reason in 13 Parameters. They use a method called TinyLaura on the Quyn 2.57B model. And just to be clear on the results, this wasn't some minor improvement where they saved space but lost performance. No, and that's the kicker. They tested it on GSM 8K, which is the standard benchmark for grade school math word problems. The base model is decent, you know, scores in the mid-range.

3:43But after training just those 13 parameters, the accuracy shot up to 91%. 91%. with 26 bytes of information. Mechanically, how is that even possible? You can't fit the rules of algebra into 13 numbers. I mean, I can't even fit my grocery list into 26 bytes. Right, because you aren't squeezing the rules of algebra into the numbers. You're using those numbers to unlock what's already there. But the how, it requires us to look under the hood a bit. They use a technique called weight tying. Weight tying. Okay, walk us through that. In a standard neural network, you have layers, lots of them. And inside those layers, you have module-specific parts of the brain that handle things like memory, attention, processing inputs.

4:23Usually, every single one of those modules gets its own unique set of parameters to update. So it's like a company where every single employee gets a unique 100-page handbook on how to do the new job. Exactly. And that's why you end up with millions of parameters. You're printing a lot of unique handbooks. Tiny Laura says, forget that. Let's write one single instruction, a post-it note, and give a photocopy of it to everyone. So they train one tiny vector, a string of numbers, and they just reuse it across the entire model. Essentially, yes. It's a master key. Now, they do some math tricks, projecting it through fixed random matrices to make it fit the different locks mechanically.

5:02But the core information is just that one tiny set of numbers repeated everywhere. So instead of teaching every neuron a new trick, you're broadcasting one signal to the whole brain. Precisely. But here's where it gets nuanced. If you have this one master key, you have to decide who gets it. Do you give it to all the attention heads, or do you give it to the layers based on how deep they are in the brain? I assume the location matters. It matters a lot. They compared different strategies. One was sharing based on the type of module, like giving the key to all the memory managers. The other was tiling, which means sharing parameters based on the depth of the layer.

5:40Depth meaning how far into the processing chain the information is. Right. Layer 1 is the surface. Layer 30 is deep abstraction. Surprisingly, they found that tiling sharing based on depth works significantly better. Huh. Why would depth matter more than function? It implies that the processing stage of the network is the critical factor. It suggests that the brain of the AI is organized in stages of reasoning. By targeting the depth, you're essentially tuning the level of abstraction rather than just tweaking specific mechanical knobs. Okay, so we have a master key, we're tiling it across the deep layers.

6:14But I'm still stuck on the why. I can't write 13 numbers on a piece of paper and teach a human calculus. Why does this work for an AI? Why does changing 13 numbers suddenly make it good at math? This is the most critical insight of the entire discussion, this 13-parameter trick. It only works because they switched how they teach the model. If you try this with standard supervised fine-tuning, it fails completely. Okay, let's unpack that. SFT supervised fine-tuning is the standard way, right? Here's the question, here's the answer, now copy it. Why does that fail here? Think about what SFT actually asks the model to do.

6:50We call it behavior cloning. You show the model a math problem and you show it the solution. The model's job is to predict the next token word by word to match that solution exactly. So it's not just trying to get the right answer. It's trying to match the style. It's trying to match the therefore and hence the specific phrasing. Exactly. It's trying to replicate the exact path the human took. That is incredibly information dense, but it's also incredibly noisy. The model is worrying about whether to say x equals 5 or the value of x is 5. That's noise. It has nothing to do with whether the math is correct.

7:24Precisely. And to memorize all that stylistic noise, you need memory. You need capacity. You can't store all that phrasing nuance. The difference between therefore and thus in 13 parameters. It's physically impossible. That's why SFT needs millions of parameters to work. It's using them to memorize the script. So SFT is like trying to memorize a monologue. You need a lot of storage for all the words. How does reinforcement learning fix this? They used a technique called GRPO group relative policy optimization. But forget the acronym. The philosophy is what matters. In this RL setup, the model generates an answer, and the system just checks the final number.

8:01Is it right? Or is it wrong? It doesn't grade the grammar. It doesn't care about the therefore. No. It creates a very sparse but very clean signal. The model tries a path. If the answer is correct, it gets a reward. It doesn't matter if the model sounded like a pirate, a professor, or a toddler. As long as the logic held up and the answer was 42, it gets the points. So it just strips away all the stylistic noise. Completely. It separates the signal, the reasoning, from the noise of the phrasing. Because the update is only driven by, did I get the math right? The model doesn't need to memorize new words.

8:37It just needs to adjust its internal logic processing. I love the analogy you used in our prep call about language learning. I think that really solidifies this. Yeah, think of SFT like trying to learn a speech in a language you don't speak, just by mimicking the sounds phonetically. You'd have to memorize every syllable, every intonation, every pause. It takes a ton of brain power and storage. Exactly. That's SFT. You're memorizing the sound, not the meaning. Now, imagine you already speak the language fluently, which these models do, and someone just tells you, hey, be more polite. That's a tiny instruction.

9:07Be polite. That fits on a sticky note. Right. You don't need to learn new words. You don't need to memorize a dictionary. You just need a tiny nudge to shift your existing vocabulary into a polite mode. That is what Tiny Laura does. It's a 26-byte nudge that says, hey, use your math circuits more. That completely reframes it. We aren't teaching it math. We are steering it toward the math it already knows. Correct. We're amplifying a capability that is already dormant in the weights. The math is there. The 13 parameters are just the switch that turns the light on. Which leads us to the next mind bender in this.

9:42You mentioned that this nudge theory actually changes how we think about model size. Because usually in tech, bigger engine means bigger fuel tank. Bigger model means bigger update file. That is the scaling law we're used to. If you have a more complex machine, you usually need more complex tools to fix it. But for steering behavior, it turns out to be the opposite. As the base model gets larger, the update size needed to steer it gets smaller. Wait, so a massive trillion-parameter AI, a god-tier model, might be controllable with fewer numbers than a small, dumber chatbot? That's exactly what the trend lines show.

10:17They found that as they scaled up the model size they could achieve peak performance with fewer and fewer parameters. Why? That feels so backward. A bigger brain should be harder to change. It goes back to the nudge. Think of the model's knowledge like a library. A small model is like a messy, disorganized room with piles of books everywhere. To find a specific piece of information, you need a detailed map, a big update file to tell you exactly where to go step by step. And a large model. A large model is a perfectly organized, indexed library. It understands the structure of the information, what mathematicians call a manifold of low, intrinsic dimension.

10:56Dancy term. It just means the solution space is already structured. The model knows where the math reasoning lives. It knows where the poetry lives. Because it's so well organized, it just needs a tiny pointer, a 13-parameter signpost that It says, look over here. So the smarter the AI, the easier it is to boss around. Ideally, yes. The bigger the brain, the smaller the steering wheel. A genius just needs a hint. A novice needs a manual. But, and there's always a but, not all brains are created equal. This really stood out in the findings regarding the difference between Quen and Llama. Oh, it was a massive difference.

11:30They compared the Quen 2.5 family against the Llama 3 family. Yeah. Quen is incredibly steerable. As we said, it hit that 91 % on math with just 13 parameters. Llama 3 basically flatlined at that scale. Flatlined. It just didn't learn. When trained with fewer than five parameters, Llama barely improved above baseline. It just generally struggled in this ultra-low parameter regime. It needed much larger updates to get the same result. Why the gap? Is Quinn just smarter at math to begin with? It's hard to say definitively without seeing the training data, but the speculation is that Quinn is primed.

12:05Its pre-training likely included more exposure to similar math examples or a structure that makes math easier to access. So Quinn was like a pot of water already at 99 degrees. It just needed a tiny bit of heat to boil. Llama was cold. That's a good way to put it. Llama might have the capability. It's a very powerful model. But that capability is buried deeper, requiring a louder shout, a larger update to get it to pay attention. It shows that steerability is a hidden trait that we haven't really been measuring until now. I want to touch on the technical constraints for a second. We keep saying 13 parameters, but when we talk about computer files, usually we try to save space by lowering quality, right, lower precision.

12:43We compress the MP3, we lower the resolution of the image. Right. In AI, we typically use something called BF16 Brain Floating Point 16. It cuts the data size in half compared to standard numbers to save memory. It's usually good enough for training massive models. But Tiny Laura flipped this script too. It did. They asked a really practical question. If I'm strictly limited by file size, say I only have a budget of 100 bytes to transmit a skill, how should I spend those bytes? Logic says use the smaller number format so you can fit more parameters in. Quantity over quality. That's what I would have guessed.

13:18But the data says no. It turns out, storing the parameters in FP32 full 32-bit precision, which takes up twice the space results in better performance per byte. So it's better to have 10 very precise numbers than 20 fuzzy ones. Exactly. When you only have 13 numbers to define an entire personality shift, every decimal point matters. You need that precision to get the trajectory exactly right. If your rudder is tiny, you better aim it perfectly. Quality over quantity. That is a great lesson. Now, I have to play the skeptic here. We keep talking about GSM 8K, which is grade school math. Is this just a trick for easy problems?

13:54Yeah. If I throw advanced calculus or competition math at this 26-byte update, does it just crumble? It's a fair question. Grade school math is one thing. The American Invitational Mathematics Examination, or AIM, is another. They tested this on the hard-stuffed benchmarks like math, AIM, and math 500. These are competition-level problems that stomp most humans. And did the tiny update hold up? They had to scale it up slightly, but not by much. To handle the really hard tasks, they trained 196 parameters. 196. That's still, what, a kilobyte? Basically, it's still microscopic. And with just those 196 parameters, the model retained 87 % of the performance improvement you'd get from full, massive training.

14:35Wow. So yes, it genuinely unlocks deep reasoning, not just simple arithmetic. It's not just a parlor trick. However, there is a limitation we need to flag. This relies on reinforcement learning, which means you need a right or wrong answer to train it. Exactly. This works for math, coding, logic, domains with a verifiable truth. The system needs to be able to automatically check the answer to give that reward signal. So I can't use this to teach the AI to write better poetry or be a funnier comedian? Not easily. If you tell an AI, write a funny poem, there's no objective true false signal to feed into the algorithm who decides if it's funny.

15:09Without that clean signal, the noise reduction of RL doesn't work. So Tiny Laura isn't the silver bullet for creative writing or subjective tasks yet. But for reasoning, it's a game changer. So let's zoom out. What does this actually mean for us, the people holding the phones? Because my phone doesn't have a server farm inside it. We talk about running AI on edge devices, but storage is always the bottleneck. That is the so what. Think about personalization. Right now, if I want a custom AI one that knows my coding style, or one that specializes in contract law, I basically need to load a separate huge model a shutter.

15:46It takes up memory. It's slow to switch. It's like changing the operating system every time I want to do a different task. But if a skill is just 26 bytes. If a skill is 26 bytes or even a kilobyte, you could store millions of them on your smartwatch. Library of skills. Exactly. You could have a Spotify of AI personalities. You simply stream the 26 bytes for physicist when you ask a science question, and then instantly in the blink of an eye, swap it for the 26 bytes for chef when you ask about dinner. The switching cost is effectively zero. It effectively makes the AI fluid. It can morph instantly based on context because the instruction set for the new personality is smaller than a tweet.

16:21It democratizes fine-tuning. You don't need a massive GPU cluster to share your custom model. You can literally tweet the code to turn a generic AI into a math genius. A text message can carry a doctorate-level capability. That is incredible. It opens up so many possibilities for sharing knowledge and customization. But, and here's where my brain goes to the darker side. It also feels a little unnerving. In what way? Well, we spend a lot of time worrying about AI alignment, making sure these massive intelligences are safe, aligned with human values. We treat alignment like this heavy concrete foundation that we build into the base of the model.

16:59I see where you're going. You're thinking about stability. Right. If I can radically alter the behavior of a model, make a genius, or potentially make it act in ways we didn't intend by flipping 13 numbers, doesn't that imply that safety is incredibly fragile? It's a profound thought. If these massive intelligences are just giant, distinct personalities separated by a few bytes of code, then what does aligned even mean? It suggests that good AI and bad AI aren't different physical structures. They are the same brain, just steered by a microscopic rudder. And anyone with 26 bytes of data can grab the wheel.

17:35It definitely raises the stakes for how we secure that last mile of steering. The base model might be safe, but the personality is fluid. We used to think you needed a supercomputer to create a dangerous or powerful AI. Now it seems you might just need a base model and a very precise, very small set of coordinates. A cruise ship steered by a rudder the size of a playing card. And then that rudder is very easy to replace. It means we need to think about security not just at the level of the massive model, but at the level of these tiny, swappable instruction sets. On that slightly terrifying but awe-inspiring note, we're going to wrap up.

18:10It's amazing to think that the future of AI efficiency lies in getting smaller, not bigger. We went from billions to 13. Who knows what we can do with 10? Sometimes the biggest changes come in the smallest packages. It's going to be a fascinating year for efficiency. Thanks for joining us on this deep dive into Tiny L 'Ori. We'll catch you on the next one.

From the publisher

This research introduces TinyLoRA, a breakthrough method for fine-tuning large language models that scales down to as few as one trainable parameter. While traditional techniques like LoRA require millions of updates, the authors demonstrate that models can achieve over 90% accuracy on complex math benchmarks using just 13 parameters. The study reveals that Reinforcement Learning (RL) is far more effective than Supervised Fine-Tuning (SFT) in this ultra-low parameter regime because RL provides a cleaner, more task-relevant signal. Experiments on Qwen2.5 and Llama-3 show that larger models are increasingly "programmable," requiring fewer absolute updates to reach peak performance. Ultimately, the paper suggests that the knowledge for reasoning already exists within pre-trained models, needing only a minimal stylistic shift to be unlocked.

More from Best AI papers explained

All 475 episodes
Learning to Reason in 13 ParametersBest AI papers explained · 19 min
Listen in VO