In short
How to scale the reinforcement learning (RL) stage of LLM training with a principled compute-scaling framework (“Scalar RL”), replacing guesswork with predictable A-B scaling curves.
Guests
No named guests; the episode is a host-led deep dive summarizing a study.
Key claims
RL performance saturates (bounded reward/pass rate), so power laws fail; use a sigmoidal scaling model with ceiling A (max reward), efficiency B (compute needed to climb), and midpoint compute C. Predict ceilings from small runs to avoid wasted GPU hours.
Notable examples
8B dense model, 100k GPU hours: power-law fit predicted A=1.0, but sigmoid fit found A=0.645. Ceiling increases via CISPO (truncated importance sampling loss; stability) and FP32 precision at the final logits (A from ~0.52 to ~0.61). Scaling knobs: MoE 17B reached same A with 1/6 RL compute; longer generation (“patience paradox”) raised A (0.61 to 0.65) but lowered initial B. Larger batch increased A (0.605 to 0.645).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding RL Scaling Challenges
0:46 to 2:15
Discussion on the complexities and costs associated with scaling reinforcement learning.
“This deep dive is all about a huge new study they clopped over 400 ,000 GPU hours to finally bring some science to RL scaling.”
Introduction to Scalar RL Framework
2:16 to 5:18
Exploring the Scalar RL framework and its parameters for improving RL performance.
“This is your absolute performance ceiling.”
Key Parameters of the Sigmoidal Model
5:19 to 8:16
In-depth analysis of the three key parameters defining the S-curve model in RL.
“It's much more stable and reliable over extremely long training runs.”
Ceiling Raisers in RL Training
8:17 to 11:43
Identifying crucial factors like loss types and precision that can raise performance ceilings.
“If the model consistently gets a prompt right, say, pass rate over 90%, they stop showing it that prompt in later training epochs.”
Optimizing RL Training Strategies
11:44 to 13:18
Review of how different strategies affect scaling, including model size and generation length.
“So bigger batches, more context, both seem to trade initial speed for a higher ultimate potential instability.”
Transcript
Automatic transcript. May contain errors.0:00Welcome to the Deep Dive. We're here to get you up to speed fast, turning complex research into, well, expertise you can use. Anyone tracking large language models knows about the training bottleneck. Free training. We've pretty much nailed down how to scale that predictable power law, it's right. Right, that part's become almost routine if you have the compute budget. But then you hit the second stage, reinforcement learning. Yeah, the RL phase. That's where models learn these really advanced skills, like step-by-step thinking or acting like agents. Exactly, and that part. It's historically been, let's say, more art than science.
0:34a bit of guesswork, a lot of GPU hours. We're talking hundreds of thousands of GPU hours spent, often just trying things out. Super expensive, super unpredictable. Until now, hopefully. This deep dive is all about a huge new study they clopped over 400 ,000 GPU hours to finally bring some science to RL scaling. Okay, so they're proposing a principled framework. What's the goal here? The mission is to unpack this framework, look at their best practice recipe called Scalar RL, and figure out, you know, which decisions actually lift the maximum performance versus just making things faster. Ah, separating the ceiling from the climb rate.
1:10That sounds useful. It's potentially a game changer, moving from just, like, intuition to actual engineering principles for RL, and it means predicting outcomes from smaller, cheaper runs. Okay, so how do we do that? You mentioned throwing out old assumptions. Yeah, the first big shift is the scaling model itself. Pre-training follows power laws, more compute, more capability, seemingly forever. But for RL, especially on tasks with a clear pass-fail or accuracy metric. Like getting a math problem right or wrong. Exactly. That expected reward or pass rate, actually$100, it's bounded. It can't go above 100 % or 1.0.
1:46A power law just doesn't fit that reality. It predicts endless improvement. So what do they use instead? A sigmoidal curve. Think of an S shape. It models things that start slow, accelerate, and then naturally level off as they approach a maximum limit. That makes intuitive sense for something like accuracy. You hit a ceiling eventually. Precisely. And the beauty is in the parameters that define this S-curve. There are three key ones we need to understand. All right, let's dig into those parameters. What's the first one? First up is A, the asymptotic reward game. This is your absolute performance ceiling.
2:18It's the highest possible pass rate the model can achieve with this specific RL setup, no matter how much compute you pour in later. So the theoretical maximum for that particular recipe, like the peak of the mountain. You got it. That's parameter A. Then there's B, the scaling exponent. This tells you about compute efficiency. How steep is that climb up the S curve? Ah, so a higher B means you reach that ceiling A fafter using less compute. Exactly. B is your speed, your efficiency. A is your destination height. Okay, A is the ceiling. B is the speed. What's the third parameter? That's CIA, the midpoint compute.
2:54It's simply the amount of compute needed to reach 50 % of the total possible gain defined by A. It gives you a concrete point on the curve. Got it. A, B, see me, dollar. And you mentioned validation. How do we know this sigmoid approach is better? Right, the crucial test case. They ran an 8 billion parameter dense model for a massive 100 ,000 GPU hours. If they'd tried to fit a power law to the early data from that run. It would have predicted continued improvement. Yeah, it wrongly predicted the ceiling A would be 1.0, like perfect performance was achievable. But the sigmoidal fit, using the same early data, correctly predicted the actual saturation point the model hit later on.
3:35Right. A was 0.645. Wow. Okay, so the sigmoid saw the plateau coming while the power law was way off. Exactly. That predicted power from smaller runs is the key. It saves potentially hundreds of thousands of dollars evaluating a method that's doomed to hit a low ceiling. Huge cost saving. All right, now that we have the measuring tools A and B, let's talk about what actually changes A, what raises that ceiling, because the study says different RL methods hit different limits, right? Absolutely. And this brings up the bitter lesson again. But for RL, you can't just look at performance early on at low compute.
4:07Because a method might look great initially, maybe it has a high B value, really efficient. But if its fundamental ceiling A is low, it'll get overtaken by a slower starting method that ultimately reaches a higher peak. So judging too early is dangerous. You need to know where the curve is headed. Precisely. You need to identify the design choices that genuinely push that asymptote A higher. Okay, what did they find? Which choices were the real ceiling razors? The first big one was the loss type. They compared several, including common ones like DAPIO and GSPO, and a variant they ended up recommending called CISSPO.
4:43CISSPO, what did that stand for? Truncated importance sampling RL loss. And the difference was significant. DAPIO plateaued around A-crete 520, but CSPO and GSPO pushed much higher, up to A around 0.590 to 0.595. It's a substantial jump in the ceiling. Why is CISPO over GSPO if they were close? Well, CSPO was maybe marginally better right at the end, but the real decider was robustness. DAPIO and similar methods calculate these importance weights, and sometimes those weights can just explode, especially in long, complex RL training. Leading to instability. Catastrophic instability sometimes. times.
5:17CISPO uses truncation basically capping those weights, which prevents those blowups. It's much more stable and reliable over extremely long training runs. That robustness is critical at scale. Makes sense. Stability equals scalability here. Okay, what was the second big ceiling shifter? You mentioned something surprising. Yeah, it's about numerical precision, specifically the FP-102 precision fix. It sounds basic, almost like just good coding hygiene. FP32, full 32-bit floating point, as opposed to mixed precision. Exactly. They found that using standard mixed precision, especially at the very final layer of the language model head where it calculates token probabilities, caused tiny numerical mismatches between the model generating answers and the model learning the policy.
6:03And those small mismatches added up. They destabilized the entire RL training process. By simply forcing that final layer calculation to use full FT32 precision, they dramatically boosted the achievable ceiling, A, from 0.52 all the way up to 0.61. That's incredible in numerical detail setting the performance limit. Was there a tradeoff, though? FP32 is slower, isn't it? It is slightly slower, yes. But the study suggests the cost of instability from mixed precision at that critical point was far, far higher. The stability gained by using FP32 allowed the model to learn more effectively and reach that higher ceiling, making the small compute overhead easily worth it.
6:41It's a net win for scaling. So the big ceiling razors are the CISPO loss for algorithmic stability and the FP32 fix for numerical stability. Got it. Now, how do these fit into the full Scalar RL recipe they recommend? Okay, Scalar RL integrates these ceiling razors with other choices optimized for efficiency for boosting that B parameter. We can sort of group the components. All right, lay out the full recipe for us. First, there's what you might call the efficiency engine. This starts with their asynchronous setup, Pipeliner RL8. It's not standard PPO. It lets data collection and training happen in parallel using eight steps of off-policy data.
7:18Much better GPU utilization. So faster training cycles. Right, and paired with that is batch-level advantage normalization to keep the updates stable and smooth. Okay, that's the hardware and update efficiency. What about the core stability parts we just discussed? That's the stability guard. This includes the CI's PL loss function we talked about and that crucial FP32 precision fix at the logits. They also found that prompt-level loss aggregation averaging losses across all generations for a single prompt hit the highest day. And how do they handle controlling the length of the model's output?
7:50Sometimes RL models ramble or stop too soon. Instead of unstable length penalties, they used forced interruption. Basically, they just append a special end-of-thinking phrase after a certain point. Simple, stable. Clever. Avoids messing with a reward signal directly. What's the last piece? The last part is the data optimizer, their beta curriculum strategy. Two techniques here. First, zero variance filtering. If multiple answers to a prompt get the exact same reward score, that prompt isn't offering any useful gradient information. So they drop it. Gets rid of useless training data. Smart. And second, no positive resampling.
8:25If the model consistently gets a prompt right, say, pass rate over 90%, they stop showing it that prompt in later training epochs. Focus the training compute on the problems the model still struggles with. Exactly. Don't waste time on solved problems. And when they put this whole Scalar RL package together and benchmarked it against other methods like DeepSeq, Quen 2.5, using Dampio, Minimax, Scalar RL consistently reached the highest ceiling, that Ailsor 0.61. Validating the whole integrated approach, what about those leave one out tests? Did they show any single component was non-essential? That's interesting.
9:01The LOO ablations, where they removed one component at a time, showed that removing most individual pieces didn't actually lower the ceiling A very much. Oh, so maybe not all parts are strictly needed for the max performance. For the ceiling, perhaps not individually catastrophic, but removing almost any component consistently hurt the efficiency of the B parameter and made the climb slower. So the full recipe is really about maximizing both A and B, getting to that high ceiling as efficiently as possible. Okay, that makes sense. The whole system works together for optimal scaling. Let's shift gears now to applying this.
9:33How does this A-B framework help us evaluate the big scaling knobs we can turn, like model size, context length, batch size? Right. This is super practical. If you have a limited budget, where do you invest? Let's start with model scale. They tested a bigger model. Yeah, they trained a much larger 17 billion parameter mixture of experts, 16 experts, they called it Scout. It followed the same predictable sigmoid scaling. But here's the kicker. What happened? The big Moe model reached the same performance ceiling, A, as the smaller 8 billion dense model, but it did so using only one-sixth of the RL training compute.
10:09One-sixth. That's a massive efficiency game. So scaling up the model size itself is a huge lever for RL efficiency. Absolutely. It suggests the bigger model architecture is just inherently better suited to learn from the RL phase, reaching the same limit much, much faster. That's a really strong argument for using MoE architectures for RL, assuming inference costs are manageable. Okay, what about generation length, letting the model think for longer? This showed a fascinating tradeoff, what they called the patience paradox. When they increased the allowed generation length from 14 ,000 tokens up to 32 ,000.
10:43Did it learn faster? No, actually the opposite. the initial progress was slower, the B value went down, the curve was flatter at the start. It took more compute to see gains. So more patience required. But did it pay off? It did. Consistently, the longer generation length resulted in a higher fitted asymptote. A, it went from 0.610 up to 0.650. So giving the model more room to think genuinely raises the ultimate performance ceiling. But you have to be willing to invest more compute overall to get there. Trade speed for height. Exactly that tradeoff. A crucial insight for planning long RL runs, and global batch size showed something similar.
11:21How did batch size scale? Moving to larger batch sizes, like 2048 prompts per update instead of smaller ones, was also initially slower lower B, higher seam aurora, took more compute to get going. But it reliably led to a higher final ceiling, A. It climbed from about 0.605 up to 0.645. Plus, the larger batches seemed more stable. They avoided the performance stagnation sometimes seen in smaller batch runs. So bigger batches, more context, both seem to trade initial speed for a higher ultimate potential instability. That seems to be the pattern. And just quickly, they also checked if this holds up for multi-pasc RL, training on math and code together.
12:00Did the sigmoid framework still apply? Perfectly. Even with the complexity of joint training, the performance curve for each domain, math, and code followed its own clean, predictable sigmoid trend. The methodology generalizes well. Okay, let's pull all this together then. This study seems to really shift RL tuning from guesswork to something predictable. Absolutely. The big takeaway is that RL scaling is predictable now, using this AB sigmoid framework. We know the ceiling A is largely set by fundamental stability choices like the loss function CSPO and numerical details like FP32 precision. And most other factors.
12:34The asynchronous setup, data filtering, normalization primarily influence the efficiency. B, determining how fast you reach that ceiling. Right. And if you want to push beyond a given ceiling, the key knobs are model scale, especially MOE and generation length, though you need to budget for the increased compute those require. This provides a much clearer roadmap for anyone doing serious RL fine-tuning. But, as always, research opens new questions. What's the final provocative thought here? Well, the paper itself points to the next logical step. They've given us predictable scaling laws within the RL phase.
13:09The frontier now is developing a unified theory. Unifying what? Unifying everything. Deriving scaling laws that connect pre-training compute, model size, the amount of RL training data, and the RL compute budget all into one grand optimization problem. How do you optimally allocate resources across the entire LLM lifecycle? Ah, the scaling laws for the whole pipeline end to end. That's the next really big challenge. That's the next deep dive we're waiting for. Indeed. We'll definitely keep watching that space. For now, that wraps up this deep dive. Thanks for joining us.
From the publisher
This paper studies scaling reinforcement learning (RL) compute for large language models (LLMs), introducing a principled framework to predict performance. The authors develop ScaleRL, a best-practice recipe derived from ablating various algorithmic choices, and demonstrate its predictable scaling trajectory using a sigmoidal function to fit compute-performance curves. Accompanying figures illustrate validation performance over increasing GPU hours (log scale) for different RL configurations, showing that ScaleRL achieves higher asymptotic performance and efficiency than prevalent methods while maintaining stability across various scaling axes, including model size and batch size. The work establishes that predictable scaling laws, similar to those in LLM pre-training, can be applied to the RL fine-tuning stage.




