LoRA Without Regret

1 Oct 2025 · 22 min · 9 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

LoRA (low-rank adaptation) for adapting large language models efficiently, focusing on when it matches full fine-tuning (“low regret”), plus operational benefits, compute/FLOPs savings, and behavior in RL.

Guests

No named guests; the episode is a host-led “Deep Dive” discussion.

Guest backgrounds

Not applicable (no guests mentioned).

Key claims

LoRA can match full fine-tuning if applied to all layers (especially MLP/and MoE), avoids capacity constraints (rank too small causes a cliff), and uses a LoRA learning rate about 10x higher than full fine-tuning. Attention-only LoRA underperforms MLP/MoE LoRA. Large batch sizes hurt LoRA more than full FT. In policy-gradient RL, LoRA can match full FT even at rank 1 because reward signals are information-poor.

Notable examples

TULU3 supervised instruction tuning; OpenThoughts3 subset (~10,000 examples) showing widening LoRA-vs-FT gap with larger batches; RL math dataset (~10,000 problems, 32 samples each) where rank-1 LoRA has ~3M parameters vs ~320k bits estimated needed. FLOPs comparison: full FT ~3N^2D vs LoRA ~2N^2D per step.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding the Inefficiency in Fine-Tuning

0:45 to 2:26

Exploring the mismatch in data sizes for fine-tuning large AI models.

“We could be talking gigabits, sometimes even megabits of new information.”

Benefits of LoRA in Deployment

2:26 to 4:21

Discussing the operational advantages of using LoRA for model adaptation.

“I mean, why did it become the standard way to deploy even when that performance parity thing was still kind of up in the air?”

Key Conditions for Low Regret with LoRA

4:21 to 6:10

Identifying essential factors for achieving performance parity with LoRA.

“The third one is just pure quality of life, really.”

Learning Rate Adjustments for LoRA

6:10 to 9:41

Discussing the critical learning rate adjustments necessary for effective LoRA training.

“Yep, performance falls off, and your training efficiency tanks your capacity constraint.”

Batch Size Sensitivity in LoRA

9:41 to 13:21

Examining how LoRA's performance is affected by batch sizes during training.

“Find your best full FTLR, multiply by 10, and that's your starting point for LoRa.”

LoRA in Reinforcement Learning

13:21 to 14:00

Exploring the efficiency and performance of LoRA in reinforcement learning contexts.

“So for best quality, stick to smaller batches and LoRa should be fine, but be aware of the tradeoff if you push batch sizes way up.”

Understanding LoRa's Impact on RL

14:00 to 15:58

Explore how LoRa performs near full fine-tuning levels with minimal data.

“Training models to be agents that can do complex stuff like mathematical reasoning or coding.”

Evaluating Computational Efficiency of LoRa

15:58 to 19:52

Learn how LoRa reduces computational costs and improves processing speeds.

“The total amount of information the model actually needs to absorb from the RL process itself is minimal.”

Key Takeaways on LoRa Application

19:52 to 21:57

Summarize the advantages and best practices for using LoRa in model training.

“Okay, so it's faster per step and uses less memory.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome back to the Deep Dive. Today, our mission is really to cut through the complexity around these massive language models, you know the trillion parameter ones. We want to figure out how we can adapt them efficiently, really digging for those key operational insights that, you know, separate a research concept from something you can actually use. So we're diving into this core contradiction, really, in modern AI engineering. You've got these enormous base models, right, pre-trained on just vast data sets. Yeah. Tens of trillions of tokens. I mean, terabits of storage just for the weights.

0:32Exactly. And that's where all this sort of general knowledge is baked in. Right. That's the foundation. But then when you, the user, want to specialize it, maybe for a specific instruction set or a particular reasoning task, your data set is often tiny. Tiny? How tiny? We could be talking gigabits, sometimes even megabits of new information. It's a huge mismatch. It just feels, well, wildly inefficient, doesn't it? Telling this giant model to learn, say, three new things, and the default method is to tweak a whole terabit of weights. Yeah, it's a massive computational waste. And that feeling, that inefficiency is exactly what we're trying to tackle.

1:07And that's really the driving force behind perimeter efficient fine tuning or PFT, as most people call it. And the method that's just completely taken over since 2021 is Loray, low rank adaptation. Loray. Right. And the idea there technically is instead of messing with the huge main weight matrix, the W, you only train two much smaller ones, A and B. That's it. The update gets applied as this low-ranked decomposition of whatever change is needed. The actual math is like W prime equals W plus alpha over R times B times A. Where R is the rank of that adapter? Exactly. And, you know, LORI was adopted super quickly just because it was convenient.

1:46But for years, this big research question kind of lingered. Which was. Can LORI actually match the performance of doing the full fine-tuning, you know, full FFT, without any compromise? Across sample efficiency, final accuracy, the whole deal. Right. That's the crucial uncertainty. And that's what we're focusing on today. The research we're dodging into specifically these studies looking for what they call the low regret regime. It basically confirms, yeah, LoRa can match full FTT performance, but and it's a bit, but you have to get certain details exactly right. Spot on. There are specific rules you got to follow.

2:19Okay. Before we get into those performance secrets, let's just pause on why industry jumped on Laura so fast. I mean, why did it become the standard way to deploy even when that performance parity thing was still kind of up in the air? Well, it really boils down to three huge operational wins that just make real world deployment so much easier. Okay. What's number one? The first, and this is probably the biggest one for cost, is multi-tenant serving. Ah, so serving lots of different customers or different customized models from the same hardware. Exactly. Yeah. Because LORIC only trains those little A and B adapter matrices, right?

2:53Yeah. The huge original base weights, the W, they stay completely untouched, frozen. So a single inference server can load that massive base model just once. And then it can keep dozens, maybe hundreds, of these tiny adapters in memory at the same time. Wow. Yeah. And when a request comes in for, like, customer A's fine-tuned version, the inference engine just, poof, swaps in that lightweight adapter. That's a game changer for costs. Instead of needing, I don't know, 100 GPU clusters for 100 models, you might need just a handful and just swap adapters. Completely changes the economics, yeah. Yeah.

3:29Makes specialized models way more feasible. Okay, second benefit is about accessibility. Training layout size. Hmm, okay. If you do full FT, the old way, you need to store the original weights, obviously, but also the gradients for everything and the optimizer state, atom states, for example. And often that needs higher precision, like float 32. Right, which balloons the memory requirement. Massively. It often takes an order of magnitude more GPU memory than you need just for inference, just to run the model. So just because my GPU can run the giant model doesn't mean I can actually fine-tune it myself.

4:01Correct. Full FT puts a huge barrier up. It severely limits who can even do the training. But LoRa, because it trains such a tiny fraction of the parameter. And here's just way less memory. Way less. Yeah. You can often train a LoRa adapter using a memory setup that's only slightly bigger than what you need just for inference for sampling. Yeah. So suddenly advanced training is accessible on way more hardware. Makes sense. And the third benefit, you said three. Oh, yeah. The third one is just pure quality of life, really. Ease of loading and transfer. Meaning. Your update file, the thing you need to move around to deploy, is now this tiny adapter.

4:36It might be megabytes, maybe even kilobytes sometimes. Instead of a whole copy of a 70 billion parameter model, which is huge. Exactly. Gigabytes upon gigabytes. So the speed of transferring that adapter, setting it up, deploying it, it's a massive win for things like CI, CD, rapid iteration. You know, simplicity just always wins in production environments. Okay, that makes total sense why it took off even before the performance proof. So operationally, LoRa is fantastic. Got it. Now let's hit that performance question head on. To get into this low regret regime, basically matching full of T's sample efficiency, what are the absolute non-negotiable conditions?

5:13Okay. There are two main ones the research points to. The first edition is about layer application, where you put the LoRa adapters. Right. And the finding is LoRa must be applied to all layers of the network. And critically, this really includes those big feed-forward blocks, the MLP layers, or if it's a mixture of experts model, the MoE layers too. Wait a sec. I thought some of the early thinking was like you should focus, Laura, just on the attention layers, you know, the query and key matrices, because that's where context is processed. Yeah, that was exactly the sort of early conventional wisdom that this research pretty forcefully overturned.

5:49Okay. So apply everywhere, especially MLPs. What's the second condition? The second one is avoiding what they call the capacity constraint. Basically, you need to make sure the number of trainable parameters in your LORI adapter is actually enough to learn the new information. How do you estimate that? Usually it's estimated based on the size and complexity of your fine-tuning data set. If you're rank, the RORI value is too small for the task. Performance drops off a cliff. Yep, performance falls off, and your training efficiency tanks your capacity constraint. Let's go back to that layer thing because that sounds like a major practical correction.

6:24You're saying applying LORI just to attention layers is not good. It's significantly worse. The research suggests that putting LoRa on the attention layers gives you basically no extra benefit if you've already properly applied it to the MLP blocks. Really? No benefit? The difference was pretty stark in the results. They showed that an attention-only LoRa, even with a really high rank, like rank 256, it actually performed worse than an MLP-only LoRa with half the rank, say, rank 128. Wow. So the learning is happening more in the MLPs for these tasks? It strongly suggests that, yeah, for fine-tuning on instruction datasets or reasoning tasks, it seems the bulk of the knowledge update, the stuff the model needs to learn, gets stored in those MLP components.

7:07They handle more of the factual recall and general knowledge processing. Okay, so practical takeaway. If I'm doing instruction tuning and I'm tight on memory, I should prioritize putting a decent rank, like 64, 128, on the MLP layers before I even think about adding LoRa to attention. Absolutely. Focus your LARAC capacity where the learning seems to be happening. And for most of these post-training tasks, that appears to be the MLP and MOE blocks. Attention matrices are vital for processing context during inference, sure, but may be less critical for storing the new specialized knowledge during fine-tuning.

7:42Fascinating. Okay, this changes the whole optimization picture then. If you're going from optimizing potentially a trillion parameters down to just these two little matrices, A and B, the hyperparameters must need adjusting, right? Specifically, what about the learning rate? Ah, yes. This is probably the single most crucial hyperparameter finding from the research. It's what they call the 10x rule. The 10x rule. Yeah. Across basically all the tests they ran, supervised instruction tuning on data sets like TULU3, even reinforcement learning experiments we'll talk about later, the optimal learning rate, the LR for LoRa, was consistently 10 times higher than the optimal LR they found for full fine tuning in the exact same setup.

8:2010 times higher. Whoa. So if someone just took their usual full FFT learning rate and used it for LoRa. Their LoRa run would likely perform very poorly, maybe even fail to converge properly. It'd be way too low. Is that factor of 10 just something they observed like empirically or is there a mathematical reason behind it based on that LoRa formula? It's actually both. It's consistently observed, but there's a structural reason. Remember the formula. What are our W plus frac alpha BA? Right. With the alpha and the R. Okay, so that alpha term is a scaling factor, and typically people set alpha equal to the rank.

8:53Okay, so if alpha equals R, then alpha divided by R just becomes 1. The scaling factor seems to disappear. It does in that final update equation, but the scaling happens implicitly during training. Here's the key insight. The research suggests that the gradient magnitude, the size of the update signal, tends to be much smaller for these low-rank$1 out of updates compared to the update for the full matrix dollars. Ah, okay. So to get the same effective step size in the overall parameter space, to make the model learn at a similar effective rate, you have to compensate for that smaller gradient magnitude by using a much higher learning rate.

9:31And that compensation factor turns out to be consistently around 10x? Empirically, yes. About 10x seems to be the sweet spot needed to equate the actual update step sizes. Okay, that makes tuning Laura much less scary if you know that rule going in. Find your best full FTLR, multiply by 10, and that's your starting point for LoRa. Pretty much. And what's really helpful is that this structure, especially the Alphur scaling, when Alphur makes the optimal LoRa Ray learning rate approximately independent of the rank to where you choose. As long as you're not capacity constrained, you know, as long as R is big enough.

10:02So I don't need to retune the LoRa much if I change the rank from 64 to 128? Generally not significantly, which is great. Now, there is one small nuance. For very, very short training runs, like maybe the first 100 steps or so, the optimal multiplier might actually be closer to 15x, not 10x. Huh. Why would it be higher just at the very beginning? It likely comes down to the standard initialization. Usually you initialize matrix B to all zeros while A gets random values. Right. In those very first few steps, the model's kind of finding its footing. That zero initialization on B acts like a temporary implicit learning rate schedule.

10:41It dampens the update initially, so you might need a slightly stronger push, a higher LR, like 15x, just to get things moving effectively before the dynamics settle down into that stable 10x multiplier regime for the rest of the training. Interesting. A little detail, but good to know for those initial steps. Okay, so we've figured out where to put L 'Oreal layers, focus on MLPs. We've got the 10x learning rate rule. We seem to be getting performance parity. So are we basically done? Are there any, like, hidden drawbacks or subtle differences even when you use LoRa correctly? Unfortunately, yes.

11:15There appears to be one key difference in the training dynamics that still persists, even with optimized LoRa. Okay. LoRa seems to show greater sensitivity to large batch sizes compared to full fine-tuning. Sensitivity to large batches. Okay, walk us through that. How does that actually show up in practice, and maybe why is it happening? Well, in some of the tests they ran, particularly on maybe small to mid-sized data sets, they mentioned one using a 10 ,000 example subset of OpenThoughts 3. They found the performance gap between LoRa and FullFT consistently widened as the batch size got bigger.

11:47Widened. So FullFT handled large batches better. Exactly. FullFT seemed much more robust to scaling up the batch size than LoRa did. So if I try to speed up my training by cranking up my batch size to, say, 1024 tokens or samples, LoRa might suffer a noticeable hit in final model quality that full FT wouldn't experience quite as badly. That's precisely the finding, yes. And importantly, this penalty seemed to happen regardless of the LoRa rank are they used. Huh. Any idea why? What's the mechanism? The likely technical reason, or at least the hypothesis, is that the more constrained parameterization of LOR, that product of matrices, B times A, just, has less favorable optimization dynamics compared to updating the full unconstrained matrix W.

12:31Less favorable how? It seems to struggle more with smoothing out the inherent noise you get in gradients when using very large batches. A full update has more degrees of freedom to kind of average out that noise effectively, whereas the low-rank structure is perhaps a bit more rigid, more susceptible to getting pulled around by noisy large batch gradients. Hmm, okay. But hang on, isn't there also a general understanding in deep learning now that really, really large batch sizes can sometimes hurt the ultimate performance anyway for both full FT and LoRa, leading to these sort of sharp minimizers that don't generalize as well?

13:05That's a really crucial point, yes. You're absolutely right. Often, both LoRa and full FT tend to achieve their best final quality using smaller, perhaps more noise-rich batch sizes. So maybe this LORA penalty at large batch sizes isn't such a disaster in practice? It might not be a catastrophic limitation, no, because if you're already aiming for optimal quality, you're probably using smaller batches anyway, where this difference between LORA and full FT seems to largely disappear. Okay. It just means that if you find yourself in a situation where you must use large batches, maybe for throughput reasons or hardware constraints, you should expect to take a proportionally larger performance hit with LoRa compared to what you'd get with full FT under those same large batch conditions.

13:47Got it. So for best quality, stick to smaller batches and LoRa should be fine, but be aware of the tradeoff if you push batch sizes way up. All right, let's pivot a bit. Let's talk about reinforcement learning, RL. That's a huge area now, right? Training models to be agents that can do complex stuff like mathematical reasoning or coding. How does LoRa stack up there when it's a reward signal driving the learning, not, you know, a big label data set? This is where the whole LoRa efficiency story gets genuinely, well, astonishing. Oh, how are we? For the common policy gradient RL algorithms, the research shows LoRa achieving performance basically equivalent to full fine-tuning, even when the LoRa rank is pushed down to almost impossibly low levels.

14:29We're talking rank one. Rank one. Seriously. Yeah. Just one dimension of update. I find it hard to picture how a single number, effectively per layer, can capture a complex policy shift for solving math problems. What does that tell us about the learning process in RL? It tells us something really fundamental about policy greedy in RL. It is incredibly information poor. Information poor. Yeah. To understand why, you have to contrast it with supervised learning, like the instruction tuning we discussed. Supervised learning, SL, is informationally very rich. The amount of information the model gets is sort of proportional to the number of tokens it processes.

15:06Millions of tokens means a huge potential information input. Okay. Makes sense. Policy gradient RL, though, it primarily learns from a scalar reward signal, or more specifically, the advantage function. That signal just says if an action or a sequence of actions led to a better or worse outcome than expected. Right. Just a thumbs up or thumbs down, basically. Essentially, yeah. Mathematically, it works out that policy gradient methods provide only O1. That's order of one bits of information per learning episode, regardless of how many steps or tokens were in that episode. So wait, whether the agent took 10 steps or 10 ,000 steps to solve some complex problem, the ultimate learning signal from that whole trajectory is just one chunk of information, like getting a single good job or bad job at the end.

15:51That's a great analogy. It's like trying to learn physics by just getting a single pass-fail grade after taking a whole exam versus getting detailed feedback on every question. The total amount of information the model actually needs to absorb from the RL process itself is minimal. Wow, okay. And the source research actually tried to quantify this. For their math data set experiments, which involved around 10 ,000 math problems, and they used 32 attempts or samples for each problem, they estimated the total information the model needed to absorb throughout the entire training process was only about 320 ,000 bits.

16:24320 ,000 bits for the whole training run. Okay. And how much capacity did their Rank 1 LoRa adapter actually have, even at Rank 1? Well, even a rank 1 LORA applied to a model like LAMA 3.18b contains roughly 3 million trainable parameters. 3 million parameters to store 320 ,000 bits of information. Exactly. It's almost 10 times the estimated required capacity. So this confirms that for standard policy gradient RL, even the absolute lowest practical LORA ranks are still drastically overcapacity. The learning signal is just that sparse. That's incredible. So for RL, you barely need any LORA capacity at all.

17:01Okay, let's pull this together. LoRa can match full FT performance if you do it right. It's way more memory efficient. And for RL tasks, it's ridiculously overcapacity, even at rank one. But what about the actual compute time? The raw processing power, the FLOPs, does it save time per step? Oh, absolutely. Yeah. When you measure progress, not just by training steps, but by the fundamental currency of machine learning floating point operations, or FLOPs, LoRa, has a very clear and substantial advantage. every single training step. Okay, let's try and unpack the math there a bit. You mentioned 3 and 2, 2 for full FT earlier.

17:37For listeners, maybe not deep in matrix math daily, can you break down why a full fine-tuning pass costs roughly 3 and 2 operations? Sure. So$9 here represents the size like the dimension of the big weight matrix dollars. In full FT, every single training step involves three main computational chunks that scale with 2, not less 2. Okay, what are they? First, the forward pass. You multiply the input by the weights. That's roughly 2 and 2 multiply add operations. Then in the backward pass, you have two major steps that also scale similarly, calculating the gradients with respect to the weights, and then actually updating the weights using those gradients.

18:11Those two together are about 2N22 arch. Okay, so 2N2 for forward, 2N2 for backward, that adds up to 3N2N22 total per step for full FFT. Precisely. Now contrast that with LoRa. Right. What happens with LoRa? With LoRa, you still need to do the full forward pass through the original weights, N2 hours. And you still need to do the first part of the backward pass to calculate the gradients needed for the adapter weights in 2NN2. So that's 2N2, 2N right there, same as the first two parts of full FT. Okay, so far it's 2N2 hours worth, the saving. The saving comes because you completely skip that third Ndola 2 operation, the one where full FT calculates and applies updates to the entire massive W matrix gradient.

18:50Instead, you do the much, much cheaper update for just the small A and B adapter matrices. Ah, right. And how expensive is updating just A and B? Does it scale with N? It does, but linearly, not quadratically. The updates to A and B together require about 6 in Rola multiply hands in total, where R is that small rank. Okay, 6 in Rola, and since R, the rank, is usually tiny compared to N, the dimension, like maybe R is 128 and N is 8 ,000. Exactly. That 6 in Rola term becomes almost negligible compared to the Tendrotti 2 term you just saved by not updating W directly. So the total FLOPs for LoRa step ends up being roughly 2N202 plus that tiny 6N roll a bit, which is basically just 2N2D all year.

19:33Pretty much, yeah. It approximates to 2N2 and others. So comparing 2N202 for LoRa to 3N2 for full FDA. LoRa takes slightly more than two-thirds of the FLOPs per training step. That's right. So if you're measuring efficiency in terms of actual computation performed or wall clock time on the same hardware, LoRa has a clear speed advantage per step. It's faster, less compute intensive. It's a win for throughput. Okay, so it's faster per step and uses less memory. Seems like a pretty compelling package. Hashtag tag outro. Wow, okay. That's a fantastic set of practical results. Let's try and summarize the key takeaways for you, the listener, tuning in.

20:07Laura really looks confirmed now as a, well, a low regret choice for almost all post-training situations. Provided you stick to the blueprint we've discussed, the one derived from this research. Yeah, you've got to get two main things right. First, apply Laura everywhere. to all the layers, but really make sure you're prioritizing those MLP and, if you have them, MOE blocks. Right. Don't just focus on attention. Exactly. And second, use the 10x rule for your learning rate. Find the optimal LR for full FT in your scenario, then multiply it by 10 for your lower run. And if you follow those rules, you get these huge benefits, less memory, faster compute per step, much easier deployment, essentially, without giving up the final performance you need.

20:45You match full FT. Yep. And, you know, stepping back a bit, this whole line of research into PKFed methods, like LoRa, it gives us this really interesting lens on the fundamentals of deep learning. It forces us to think harder about model capacity, about data set complexity, about sample efficiency. Like that RL finding. Exactly. Seeing that policy gradient RL needs almost zero LoRa capacity because the learning signal itself is so incredibly weak. That tells you something profound about the algorithm itself. Which leads us right to a final thought to leave you with. The source material we looked at hints at this.

21:21While policy gradient RL might only give O1 bits of information per episode, other RL approaches are emerging. Things like model-based RL, where the agent tries to learn a richer model of how the world works from the same interactions. Right. Those methods might potentially extract more information from the same trajectory data. So for you, the curious learner, here's the question to ponder. If these more informationally rich RL techniques like model-based RL become the standard way we fine-tune foundation models in the future, how much would the necessary LoRa Ray rank need to increase to actually capture that richer signal without becoming capacity-constrained again?

21:55Something to think about.

From the publisher

This research provides a detailed analysis of Low-Rank Adaptation (LoRA), a parameter-efficient fine-tuning (PEFT) method for large language models, comparing its performance against full fine-tuning (FullFT). The authors establish a "low-regret regime" where LoRA matches the performance and sample efficiency of FullFT, particularly for small-to-medium-sized datasets, provided key implementation details are correct. Operational benefits of LoRA, such as improved multi-tenant serving, reduced training memory footprint, and easier transferability, are highlighted as reasons for its growing popularity. The research emphasizes that for optimal performance, LoRA must be applied to all model layers, especially the MLP/MoE layers, and that its optimal learning rate is consistently about ten times higher than for FullFT. Finally, the analysis shows LoRA's significant advantage in reinforcement learning scenarios due to the inherently low information capacity required for such tasks, and discusses its computational efficiency advantage, requiring slightly more than two-thirds of the FLOPs of FullFT per training pass.


More from Best AI papers explained

All 475 episodes
LoRA Without RegretBest AI papers explained · 22 min
Listen in VO