In short
The episode argues that “power sampling” can make base LLMs outperform or match RL post-training methods (e.g., GRPO) for reasoning tasks using only inference-time sampling—no new training, data, external rewards, or verifiers.
Guest backgrounds
No guests are mentioned; it’s a solo “Deep Dive” discussion of a research paper.
Key claims
RL mainly performs “distribution sharpening” (concentrating probability mass on one best reasoning path), which boosts single-shot accuracy but collapses diversity and hurts generalization. Power sampling instead reshapes the model’s probability distribution by raising token/path probabilities to an exponent alpha, preserving diversity while improving accuracy.
Notable examples
On HumanEVOL, power sampling improved PHY 3.5 mini-instruct by +51.9% vs GRPO. A HumanEVOL string-filtering task: power sampling produced a correct Python list comprehension; GRPO failed with redundant f-string concatenation. Pass@K: GRPO’s curve flattened quickly, while power sampling kept improving with more samples. Compute: ~8.84x inference cost vs greedy, estimated comparable to one GRPO training epoch.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOChallenging the Reinforcement Learning Paradigm
0:45 to 3:37
Discussion on the conventional view of reinforcement learning in training AI models and the potential of power sampling.
“We're looking at research on something called power sampling.”
Understanding Distribution Sharpening
3:37 to 4:54
Explaining how reinforcement learning sharpens probability distributions and its implications for AI reasoning.
“The amazing thing is the sources show they can get that high single-shot accuracy without the RL just by cleverly using the model's own internal likelihoods when generating the answer.”
The Mechanism of Power Sampling
4:54 to 7:01
Detailed exploration of how power sampling works, including its mathematical foundation and its difference from low temperature sampling.
“You're dramatically upweighting the sequences the model already thought were pretty good, and you're heavily penalizing the ones it was unsure about.”
Implementing Power Sampling with MCMC
7:01 to 8:28
Description of how the Metropolis-Hastings algorithm is applied in power sampling for AI reasoning.
“If you're considering the likelihood of the entire future sequence for power filler, how do you actually compute that?”
Comparing Power Sampling to RL Models
8:28 to 11:15
Assessment of the performance of power sampling against traditional RL models like GRPO, highlighting significant results.
“In each step, they propose a change to the token sequence within that block, maybe resampling a subsequence.”
Diversity in AI Reasoning Outputs
11:15 to 12:26
Exploring how power sampling maintains diversity in generated outputs compared to RL models, enhancing multi-shot accuracy.
“They showed a concrete example from human evil.”
Cost Efficiency of Power Sampling
12:26 to 14:00
Discussion on the computational costs of power sampling versus traditional RL training, emphasizing practical benefits.
“It maintained that high single-shot accuracy without sacrificing the ability to explore multiple valid reasoning paths.”
Analyzing Computation Costs in AI Training
14:00 to 14:48
Explore the comparison between inference costs and training costs in AI models.
“Okay, nearly 9x the compute per token generated sounds like a lot.”
Reevaluating AI Model Development Strategies
14:48 to 15:24
Discuss the potential shift from complex training to intelligent inference techniques.
“Maybe the focus should shift away from just more and more complex training towards more intelligent ways of using the models we have at inference time.”
Unlocking Latent Knowledge in AI Models
15:24 to 16:04
Learn about the importance of leveraging existing model knowledge for better outcomes.
“well, it's quite a humbling finding for the field, I think.”
Transcript
Automatic transcript. May contain errors.0:00Welcome back to the Deep Dive. Okay, so if you've been following the world of, you know, large language models, the standard story for getting really high-level reasoning. Right, like the heavy stuff, math proofs, complex coding, science problems. Exactly. That story always involved massive, expensive training, usually reinforcement learning, right? Things like RLVR, GRPO. Those became the sort of industry standard. Yeah, the gold standard. The assumption was you had to do this post-training reinforcement learning step to make the model smart enough for these tasks. But what if that whole consensus was, well, maybe not wrong, but missing something fundamental?
0:39What if all that complex retraining wasn't actually necessary? And that's exactly what we're digging into today. It's pretty disruptive stuff. We're looking at research on something called power sampling. Power sampling. And the kicker is zero training, zero new data, zero external rewards needed. It's purely an inference time technique. Wow. Okay. So no retraining cycles, no math of reward data sets. None of that. And the claim is bold. Just by being smarter about how the base model generates text, how it thinks, essentially during inference, this method can match and sometimes even beat those state-of-the-art RL methods.
1:15So the researchers are basically saying, hey, the base model, it's way smarter than you've been giving it credit for. That's the headline, yeah. The base model is smarter than you think. Okay, let's unpack that assumption first because this feels like the core shift. The belief was that RL post-training was actually, you know, creating new skills, teaching the model to reason better. Right, that it was adding fundamentally new capabilities. But if power sampling works just by changing how we sample from the existing model, then RL wasn't really creating new knowledge, was it? What was it actually doing?
1:49Well, the evidence here suggests RL was mostly doing something called distribution sharpening. Distribution sharpening, okay. Think of it like this. The base model, from its huge pre-training, could already generate correct high-quality reasoning paths. It knew how, somewhere in its probability distribution. RL post-training just came along and found those existing high-likelihood answers and made them way, way more probable, hyper-concentrated, you could say. Ah, okay. So if the model initially knew, say, 20 different ways to solve a math problem, some good, some okay, RL comes in, finds the best one according to its reward, and basically cranks up the volume on that one path while silencing the other 19.
2:32That's a really good way to put it. That's the core idea. It sharpens the probability landscape. Yeah. Which is great if you just want the single best answer on the first try, what we call single shot accuracy. Right. Get it right the first time. Exactly. But, and this is the crucial downside, this sharpening comes at a cost, a massive collapse in generation diversity. Meaning? Meaning if you ask that RL-tuned model for, say, 10 different potential solutions, what we call pass at K, trying to solve it within K attempts, you often find those 10 solutions are just tiny variations of each other, maybe even variations of the same wrong path if the RL over-optimized.
3:07Ah, so it loses the ability to explore different valid approaches. Precisely. And paradoxically, the original base models, the ones before RL training, often do better in those multi-shot scenarios, like pass at 10 or pass at 100 because they still have that diversity. RL optimized for concentration, but it sacrificed variety. And for complex reasoning, that variety is probably crucial, right, for exploring different lines of thought, maybe correcting itself. Absolutely. You need that room to explore. Which brings us back to power sampling. The amazing thing is the sources show they can get that high single-shot accuracy without the RL just by cleverly using the model's own internal likelihoods when generating the answer.
3:49And the engineering benefit, you mentioned training-free, data set-free, verifier-free, that avoids all the headaches and costs of RL. It's huge. No unstable training runs, no need for complex reward functions or verifiers. It shifts the complexity from training to inference, but in a manageable way. Okay, so how does it work? What's the mechanism? You mentioned it's based on the model's own likelihoods. Yeah, the foundation is actually pretty elegant mathematically. It revolves around creating what's called a power distribution, usually written as tiles raised to the power of alpha. Oh, alpha.
4:22So palace is the base model's normal probability distribution for the next token or sequence. Exactly. Towers is just what the standard model thinks is likely. Now, you take those probabilities and you raise them to a power, alpha, where alpha is greater than one, say, alpha 4 ,000. Right, and raising probabilities less than one to a power greater than one makes them smaller. And it makes the higher probabilities relatively much, much higher. It systematically biases the whole distribution. So you're essentially amplifying the model's confidence in its best ideas, making the strong signals stronger and the weak signals almost disappear.
4:58That's precisely it. You're dramatically upweighting the sequences the model already thought were pretty good, and you're heavily penalizing the ones it was unsure about. You're telling it. Focus only on the highest likelihood reasoning path that you already know. Okay, this sounds intuitively similar to something people already do, which is low temperature sampling. Setting the temperature parameter tall low makes the output less random, more focused. Why isn't this just the same thing? Isn't low temperature tall one alpha? Ah, this is the crucial distinction. And honestly, it's the real aha moment in this work.
5:31They look related and people often conflate them, but they operate very differently on the sequence generation process. Okay. How so? Low temperature sampling is greedy. It looks at the probability of the very next token and sharpens that choice based on the immediate likelihoods. It picks the locally best next step. Right. Focused only on the current move, like you said. Correct. Tower sampling, PALFA is different because it inherently considers the likelihood of the entire future path that stems from a potential token choice. The whole sequence. How? Well, the math gets a bit abstract. It's about whether you exponentiate the sum of log probs or sum of the exponentiated probs, but the functional outcome is key.
6:10Power sampling effectively weights tokens based on the total probability of the completions they lead to. Okay, let me try an analogy. Is it like low temperature sampling is a chess player who just takes the highest value piece right now? Yeah. Even if it leads to a bad position later. Whereas power sampling is looking ahead, prioritizing the move that leads the overall highest probability of winning the game, even if the immediate capture isn't as flashy. That's a fantastic analogy. It captures exactly what the people call Observation 1. Power sampling tends to upweight tokens that might lead to fewer but high likelihood future paths.
6:48It helps the model avoid those critical decision points where choosing the immediately most likely token actually traps the rest of the generation in a low probability incorrect outcome. Power sampling has this implicit look ahead quality. That's really clever. But OK, the practical problem. If you're considering the likelihood of the entire future sequence for power filler, how do you actually compute that? The number of possible sequences is astronomical. You can't normalize that distribution. Exactly. Direct sampling from alpha for the whole sequence is computationally intractable. And this is where they pull out a really elegant tool from, well, old school statistics.
7:24Okay. What is it? They use the Metropolis-Hastings algorithm. It's a classic Markov chain Monte Carlo method, MCMC. MCMC. Okay, I've heard of that. Used in physics, Bayesian stats. Why here? Because MCMC is designed specifically for sampling from complex probability distributions where you don't know or can't compute the denominator in the probability formula, which is exactly our problem with the global alpha. Ah, so it's a way to explore that alpha landscape and find high probability samples without needing to map out the entire thing. Precisely. It's a statistical simulation trick. Instead of calculating everything, MCMC lets you effectively search the high likelihood regions.
8:04Okay, so how do they implement it? Are they running MCMC for the entire text generation? No, that would still be too slow. They use an approach called autoregressive MCMC. They generate a chunk of text, a prefix, using the normal model. Then, for the next chunk or block of tokens, say, a block-sized dollar, they run an iterative MCMC process. Iterative. Yeah. They run it for a certain number of steps. Let's call it net MCMC all. In each step, they propose a change to the token sequence within that block, maybe resampling a subsequence. Then they use the Metropolis-Hastings rule, which looks at the ratio of the alpha likelihoods of the proposed sequence versus the current one, to decide whether to accept the change or stick with the current sequence.
8:47Over several MCMC steps, this process guides the block towards a much higher likelihood configuration under alpha. Got it. So it's like refining a chunk of the output multiple times before moving on. This confirms that tradeoff you mentioned. Less compute upfront and training, but more compute, more search at inference time to get that higher quality sample. Exactly. It's an explicit choice. Spend more flops during generation to find a better answer that the base model already knew how to produce. And what do they find worked best for these parameters? The alpha, the block size I-law, or the number of MCNC steps?
9:19They tested various things. They found an intermediate sharpening value of alpha equals$4 worked well. For the block size, they used Y$ equals 122 tokens for typical sequence lengths. And crucially, they found they didn't need a huge number of MCMC steps. The accuracy seemed to stabilize pretty quickly after just none MCMC teller steps. Only 10 steps? That's surprisingly few iterations to find that better path. Yeah, it suggests the high likelihood paths aren't that hard to find if you have the right search method. 10 quick refinement iterations per block were enough. Okay, let's get to the results.
9:54The performance. This is where the rubber meets the road. How did power sampling actually compare against GRPO, the RL benchmark? The results are pretty striking. On the tasks that the RL models were specifically trained on their in-domain tasks, like the Match 500 benchmark, power sampling achieved comparable performance, roughly neck and neck with GRPO, which is already impressive for a training-free methyl. Just matching the specialized RL model is a win. But here's the really critical part. On generalization tasks, things the RL model wasn't explicitly trained for, power sampling often significantly outperformed GRPO.
10:28Outperformed? Okay, like what kind of tasks? Like HumanEVOL, which is a standard benchmark for coding ability, and AlpacaeVOL 2.0, which measures general helpfulness on non-verifiable instructions. Power sampling pulled way ahead there. Wow. Any specific examples jump out? Yeah. The paper highlights a really dramatic one with the PHY 3.5 mini-instruct model. On HumanEVault, using power sampling boosted the accuracy by, get this, plus 51.9 % compared to the version fine-tuned with GRPO. 50 % just from changing the sampling strategy. That's not incremental. Not at all. And it points directly back to that diversity issue we talked about.
11:06The RL model likely overfitted to its training data distribution. It became very good at that specific domain, but brittle outside of it. So that distribution sharping actually hurt its ability to generalize. Seems like it. They showed a concrete example from human evil. The task was to filter a list of strings based on a prefix. Power sampling generated a clean, correct Python list comprehension. Okay. And the RL model? The GRPO sample failed. It got stuck in some weird, redundant logic involving string concatenation like fprefix2. Just completely wrong. It seemed trapped in a low-likelihood reasoning path because its distribution had collapsed too much.
11:45That's a perfect illustration of mode collapse hurting generalization. Okay, so single-shot accuracy is comparable or better, especially out of domain. What about that diversity score, PASC-K, if you ask for multiple solutions? Ah, this is where power sampling really demonstrates its advantage. Remember how RL methods tend to have their PASC-K performance taper off quickly as TECOL increases? Yeah, because all the samples are too similar. mode collapse. Exactly. When they plotted the pass at K curves in the paper, the GRPO curve flattened out fast. It didn't get much better after the first few samples.
12:18Power sampling, though. Its curve kept climbing strongly for Keller dollars. Meaning it was finding genuinely different, correct solutions. Yes. It maintained that high single-shot accuracy without sacrificing the ability to explore multiple valid reasoning paths. It often reached performance close to the base model's theoretical upper limit based on its inherent diversity. It finds the good paths without destroying the variety that RL tends to eliminate. It gets the best of both worlds, essentially. High accuracy and exploration. That's powerful. Was there anything else interesting, maybe side effects?
12:53One interesting side effect was the response length. You know how sometimes RL models are rewarded, maybe implicitly, for generating longer step-by-step reasoning? Right. Verbosity can sometimes correlate with correctness in the reward signal. Well, power sampling, which has no external reward signal, naturally ended up producing responses of similar length. On MATH 500, the average length was around 679 tokens, very close to GRPO's 671 tokens. Huh. So the length wasn't just an artifact of the reward. It's inherent to actually solving the problem correctly within the model's own preferred paths.
13:27That seems to be the implication. The high likelihood reasoning paths that power sampling finds just naturally require that level of detail and length to be correct. Fascinating. Okay, let's wrap up with the bottom line. Cost efficiency. We know RL training is expensive. How does the inference time cost of power sampling stack up? Right. This is maybe the most practical argument. They calculated the overhead. Running power sampling with those settings, Nalan MCMC dollar steps, multiplies the token generation cost by about 8.84x compared to standard greedy single-shot inference. Okay, nearly 9x the compute per token generated sounds like a lot.
14:05It does, but compare it to the alternative. The researchers estimate that this controlled, predictable inference cost is roughly comparable to the compute needed for just one single epoch of GRPO training. One epoch. And GRPO training usually involves many epochs, right? plus data set curation, potential instability. Exactly. You're comparing a predictable, stable, one-time inference cost multiplier that delivers superior, more general results against a massive, uncertain, potentially unstable training cost just to get a model that might be worse at generalization. So the trade-off seems heavily in favor of smarter inference, doesn't it?
14:41That's the argument this research makes, very strongly. It really reframes the whole conversation around how we improve LLMs. Maybe the focus should shift away from just more and more complex training towards more intelligent ways of using the models we have at inference time. The big takeaway then is that this powerful reasoning ability, it's not something we necessarily need to painstakingly build into the models with complex RL. It seems it's already there, embedded in the high likelihood regions of the base model's distribution. Right. It's latent knowledge waiting to be unlocked. And techniques like power sampling provide a key.
15:15The fact that a sophisticated sampling method can achieve these results with better generalization, just by leveraging what the model already learned during pre-training, well, it's quite a humbling finding for the field, I think. It really is. It makes you wonder. If the biggest gains are now coming not from building entirely new models, but from getting smarter about how we prompt and sample from the ones we already have, does that signal a major shift? Perhaps the answers to many next-gen AI problems are already sitting there in the existing probability distributions, and we just needed the right statistical key, like power sampling, to actually find them.
15:51A powerful thought indeed. It suggests we might need to invest just as much in the art and science of asking as we do in the science of building. Definitely something to chew on. Thanks for breaking that down. We'll have to leave it there for this deep dive.
From the publisher
The research paper titled "Reasoning with Sampling: Your Base Model is Smarter Than You Think" by Harvard researchers introduces a novel, training-free iterative sampling algorithm inspired by Markov Chain Monte Carlo (MCMC) techniques to enhance the reasoning capabilities of large language models (LLMs) at inference time. This method, termed "Power Sampling," leverages the base model's own likelihoods to simulate sampling from a "power distribution," which sharpens the distribution toward higher-likelihood sequences without additional training or the need for a reward signal. The authors argue that this technique successfully elicits latent reasoning skills in base models, demonstrating performance on par with, and sometimes exceeding, models post-trained with Reinforcement Learning (RL), particularly the Group Relative Policy Optimization (GRPO) method, across diverse benchmarks like MATH500, HumanEval, and GPQA. Crucially, Power Sampling maintains greater generation diversity compared to RL-posttraining, which typically suffers from a collapse in multi-shot performance.




