Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems

3 Dec 2025 · 14 min · 8 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Pass@K Policy Optimization (PKPO) for reinforcement learning of LLMs on objective, high-stakes tasks (code/unit tests, theorem correctness). It addresses the “pass@1 trap,” where standard policy gradients optimize expected reward of a single sample and underuse diverse attempts needed for rare successes.

Guest backgrounds

No guests mentioned; this is a solo “Deep Dive” episode.

Key claims

Maximizing expected maximum reward over K samples encourages high-variance, joint-utility exploration. PKPO uses low-variance, unbiased gradient estimators via reward transformations (“best-shot transformer”) so the max objective is trainable with standard policy-gradient updates. It decouples batch size from K, allowing arbitrary K optimization.

Notable examples

ARC-AGI-1—baseline pass@1 RL stalls; PKPO with K>1 boosts cumulative solve rate for Gemma 2 9B from ~12% to ~82%. Annealing K (high K early, then KL-style consolidation) preserves pass@1 quality while retaining diversity.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding the Pass at One Trap

0:45 to 2:16

Explaining the limitations of traditional reinforcement learning methods in solving complex AI tasks.

“Because when you move to those objective really high bar challenges, it turns out that traditional RL methods, well, they start to fail.”

Introducing Pass at K Policy Optimization

2:16 to 3:52

An overview of the Pass at K Policy Optimization method and its advantages in AI training.

“It's not about making sure your average attempt is just slightly better than random.”

Technical Breakthroughs Behind PKPO

3:52 to 7:00

Delving into the mathematical foundations and innovations making PKPO effective.

“it seems like the missing piece for complex reasoning.”

Real-World Performance of PKPO

7:00 to 10:17

Examining how PKPO performs on complex real-world AI tasks and its impact on exploration.

“into a stable standard policy gradient problem that we already know how to solve.”

Trade-offs and Consolidation in PKPO

10:17 to 13:05

Discussing the balance between exploration and efficiency in PKPO training stages.

“and the single sample policy gradient gives it no reward signal for trying anything risky.”

Core Principles of PKPO

13:05 to 13:27

Highlighting the fundamental change in success definitions for LLM training.

“That speed of adoption is absolutely key.”

Implications for Future AI Algorithms

13:27 to 14:00

Exploring broader applications of the PKPO approach to enhance AI reasoning capabilities.

“that seems like it could apply whenever an LLM uses a search process.”

Exploring Advanced Goal-Directed Reasoning

14:00 to 14:22

Learn about advanced reasoning techniques in AI and their implications.

“You could unlock truly advanced goal-directed reasoning.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to the Deep Dive. Today we're tearing into some really cutting edge research that's aimed at solving one of AI's biggest limitations right now. How do we get large language models to solve problems that are genuinely hard? And we're not talking about just writing a nice email. We mean verifiable, high-stace tasks, like generating flawless code or proving complex mathematical theorems. That's right. I mean, a lot of the big leaps we've seen in LLMs come from these post-training methods. You know, you start with a huge pre-trained model, then you use supervised fine-tuning and critically reinforcement learning RL to align it.

0:35But the frontier is shifting. Early RL was all about soft subjective signals like human feedback. Now it's about objective grounded reward signals. Can the code actually pass the unit tests? Is the math proof logically sound? Yes or no. And that's where things get tricky. Because when you move to those objective really high bar challenges, it turns out that traditional RL methods, well, they start to fail. They hit a wall that the sources for today's deep dive call the pass at one trap. Exactly. Just imagine a really challenging puzzle. For any system to succeed, it often needs to try multiple different approaches, right?

1:08In modern LLM training, the model samples many different solution attempts for a single problem. Let's say, you know, another attempts, maybe 16 or 32. The problem is that standard policy gradient RL algorithms only optimize for the expected reward of a single sample. We call that pass at one. Okay, so let me get this straight. The model generates, say, 16 different possible solutions, but the training algorithm is only rewarding the probability that the first one, or maybe the average one, is correct. What happens to all the other ideas, the high-risk, diverse concepts in the other 15 samples?

1:41They're just drastically underutilized, even though one of those diverse attempts might actually hold the key to the solution. Optimizing strictly for pass at one prioritizes safe, conservative individual performance. It ignores the collective utility of the entire set. And what the sources show is that this fundamentally limits exploration, especially on tasks where success is rare. Learning just stalls. The model is never incentivized to try a brilliant but high-variance approach that's often wrong. That makes so much sense. If you have a really difficult task, success means maximizing the chance that at least one of your attempts is brilliant.

2:17It's not about making sure your average attempt is just slightly better than random. That is the crucial shift in perspective. And that brings us to the really elegant solution presented in this research. It's called Pass at K Policy Optimization, or PKPO. And PKPO is designed to directly optimize for the expected maximum reward over a set of K samples. This shifts the entire training signal toward prioritizing what you could call joint utility. Okay, let's unpack that. Joint utility. Because that sounds like it's at the heart of this whole technical breakthrough. First, can you just define PASECK for us?

2:50What does it mean for a simple problem? And then how does that scale to something complex like code generation? Absolutely. So if we're in a simple binary world, you know, the answer is either right or it's wrong, then PASECK is simply the probability that at least one of your dollar independent samples is correct. But LLMs often deal with continuous rewards. For instance, a piece of code might pass, say, 73 % of its unit tests. In that case, the objective generalizes to max at K, which is the expected maximum reward you get across those taller attempts. So if I'm training a model and I set taller to$10, I'm basically telling it, look, I don't care if the average of your 10 attempts is 50%.

3:27I only care if the single best shot among those 10 hits 95%. Exactly right. By optimizing for max at K, the policy gradient actually encourages the model to generate sets of samples that jointly maximize that best outcome. And this fundamentally rewards the policy for high-variance exploration, because the optimization signal only depends on the best sample in that batch. The others don't drag it down. This idea of prioritizing joint utility over individual utility, it seems like the missing piece for complex reasoning. If a problem is really hard, a good strategy is to hedge your bets, right?

4:01You generate a few safe, predictable answers, but you also throw in a handful of really risky, out-of-the-box ideas that might just unlock the solution. That's the core tension. The sources actually explored this with a small, one-dimensional toy problem, and the findings were really striking. They found that for larger values of$2, the optimal policy becomes increasingly risk-tolerant. It learned to accept a much higher rate of zero-reward samples, that junk you mentioned, if that trade-off increased the probability of getting at least one brilliant, maximally-rewarded sample. But is that actually desirable in the long run?

4:36I mean, generating a bunch of risky, diverse samples, doesn't optimizing for a high dollar a risk making the average output quality worse? Like, does the model become wildly inefficient when you finally just ask it for a single answer in production? That is the perfect question, and it's really what this research was designed to answer. And yes, if you only look at the model's entropy or its individual sample quality, optimizing for a high dollar does make the average sample a little bit worse. But for these frontier tasks where traditional RL stalls at, say, a 12 % solve rate, this diversity is the only way forward.

5:08We're consciously prioritizing finding the needle in the haystack over just, you know, avoiding the hay. And as we'll get into, PKPO actually offers a way to have your cake and eat it, too, through some smart training schedules. Fascinating. So it sounds like we're switching from training a highly consistent, accurate model to training a policy that can manage a diverse, high-risk portfolio of solutions, knowing only the best one matters. That is a perfect metaphor. The model learns to distribute its policy mass in a completely different way to maximize its expected best outcome. Okay, let's switch from the why to the how.

5:41From the concept of RIF to the technical difficulty of actually achieving it. The word maximum in math is, well, it's tricky. When you try to differentiate a maximum function to sign the gradient for optimization, you run into all sorts of discontinuities. So what's the technical breakthrough here that makes PKPO robust and actually feasible for policy optimization without just adding tons of noise or bias? This is PKPO's primary contribution. It's really the key. They derived a set of novel, low-variance, unbiased estimators specifically for the pass at K and maxed at K gradient. Just a level set for everyone.

6:17An unbiased estimator means that, on average, the gradient we calculate points exactly where the true optimal policy should go, right? Exactly. If your estimator's biased, you're just going to consistently drift away from the true solution. But just being unbiased isn't enough. RL training is inherently noisy. Calculating a gradient based on a maximum value across a batch of samples just adds another layer of instability. So what the authors did was find a way to transform the rewards within the batch, such that calculating the standard policy gradient on these new transformed rewards gives you the correct, unbiased, and crucially low-variance update for that pass at K objective.

6:54Ah, so it's a clever mathematical trick. It essentially transforms a really difficult, non-differentiable optimization goal, the maximum, into a stable standard policy gradient problem that we already know how to solve. Precisely. They introduce what we can think of as a best shot transformer. This transformation uses a specific formula to calculate the contribution of each sample to the overall best outcome of the set. So it doesn't just look at the raw reward. It looks at the reward in the context of all the other samples generated at that exact moment. And this context-aware calculation naturally includes a baseline that reduces variance and stabilizes the whole learning process immensely.

7:32I think you mentioned earlier that this derivation has a big practical advantage over previous attempts, which were often pretty limited. What was that limitation? It was often a coupling issue. Previous methods really struggled with the math. Unless your batch size, 1x die, was exactly equal to your optimization target,$2. So that means if you decided you needed 64 samples to find a hard solution, you were forced to optimize for pass at 64. It was really restrictive and inefficient. The PKPO changes that. Yes. PKPO's transformations are the first to enable robust optimization of PATH, a K, for any arbitrary dollar up to no samples.

8:06You can collect a big batch of, say, 128 attempts for efficiency, so$128, but tell the policy to only optimize for the success rate over the best eight samples, where a K equilitated$8. This decouples computation from the policy objective. It gives the trainer complete control over the degree of exploration they want to inject. That's a huge operational advantage. Okay, that covers the theory. Now, this is where it gets really interesting for me. How does this actually perform on real LLMs and messy real-world tasks? You don't get bonus points for generating 10 ,000 bad answers in the real world, after all.

8:40The validation was really strong. They used large open source models like Gemma 2 and LMA 3.1 across some really complex grounded benchmarks. They focused on tasks that demand true multi-step reasoning. The math beta set, code generation with human evil and MBPP, and the extremely challenging Arse AGI 1 reasoning task set. What were the key metrics they were tracking to prove that PKPO actually improves exploration? The most important one was the cumulative solve rate. So that's defined as the fraction of tasks for which the model has sampled a correct solution at least once across the entire training process.

9:13It's a direct measure of successful exploration of its ability to find those difficult solutions. And what was the pattern that emerged when they started dialing up that optimization target that transferred? That's a very, very clear and convincing pattern. As they increased during PKPO training, the model showed a consistently higher cumulative solve rate. They were just finding solutions to more unique problems. And crucially, they also saw a big increase in model entropy, which is just the technical way of saying the policy learned to be much more diverse and explore a wider range of high-quality samples.

9:45So yeah, optimizing for joint utility demonstrably expands the model's exploration boundaries. Let's focus on that ARC AGI-1 task set. My understanding is that's a benchmark designed to test abstract reasoning, and it's known for being brutally difficult, often requires complex search. What happened there specifically? This is the headline result. It really demonstrates the core limitation of traditional RL. On ARC-AGI 1, the standard pass at one optimization, so the baseline RL approach, it just stalls. The model makes almost no progress because it can't find a foothold on those first few difficult steps, and the single sample policy gradient gives it no reward signal for trying anything risky.

10:25And the difference when they introduced PKPO. It was transformative. Deploying PKPO with a dealer greater than one completely unblocked learning. For the Gemma A29B model, the cumulative solve rate jumped from a dismal starting point of just 12 % with standard RL to an astonishing 82 % when they trained it with PKPO. Wait, wait. From barely solving 1 in 10 to solving 8 out of 10, despite changing the reward objective. That suggests the failure wasn't necessarily a lack of knowledge in the model, but a failure of the training incentive. Exactly. It proves that for these advanced reasoning tasks, the bottleneck is often the optimization objective itself, not the raw capacity of the underlying transformer.

11:04If you reward conservatism, you get conservatism. If you reward diverse, high-potential exploration, the model finds solutions that were completely inaccessible before. You mentioned earlier there might be a way to resolve that trade-offs, getting all that diversity, without sacrificing the single-sample quality. How did they pull that off? They used the flexibility of PKPO to implement a strategy called annealing dollars. So they would start training with a high pollard out, let's say, at early stages. This maximizes exploration and rapidly expands the space of known solutions. Then, in the later stages of training, they reduced tallers all the way down to not collar.

11:41So high tellers for pure exploration, and then you switch to KL1 at the end to sort of consolidate the policy and polish that single sample quality. Precisely. And the result was really the best of both worlds. The final models got the high pass at K rates for all Taylor dollars greater than one that you'd expect from that aggressive exploration phase. But, and this is the key part, they achieved this with no observed sacrifice and pass at one performance compared to the baseline. The consolidation step locks in the efficiency and quality without losing all the solutions it discovered during that high risk phase.

12:13That's definitely the dream outcome for any real world deployment. Yeah. Okay, so to bring this all together for everyone listening, standard reinforcement learning, by focusing so myopically on average individual performance on pass at one, is just fundamentally unsuited for teaching LLMs to tackle these truly complex frontier reasoning tasks. Yes. That pass at one trap just limits the model's ability to explore nonlinear solution paths. PIA-KEO resolves this by offering a robust, mathematically sound, and efficient way to optimize for the expected best outcome within a set of samples that pass at K-objective.

12:45It's really a fundamental change in how we define success in LLM training. And because PKPO delivers all this through these robust, easy-to-calculate reward transformations, it's not some esoteric new RL algorithm you have to build from scratch. It's basically a drop-in replacement for the reward signals in existing policy-gradient RL algorithms. It's immediately deployable. That speed of adoption is absolutely key. You don't need a whole new infrastructure. you just need to calculate the reward a little differently based on the batch of samples you already have. That brings us to a really compelling final thought.

13:19This research focused on optimizing for the best result out of multiple independent samples. But the core principle here, maximizing the expected best reward rather than the average reward, that seems like it could apply whenever an LLM uses a search process. So what other, maybe more complex, inference-time search algorithms do you think could immediately benefit from this insight into joint utility? Oh, I immediately think about methods like Monte Carlo tree search or maybe various forms of beam search. Those are already designed to generate and evaluate pools of high potential paths or tokens.

13:52But right now, the policies guiding those searches are often trained with traditional RL signals. If you could integrate the PKPO objective, if you could optimize the parameters of the search itself to prioritize paths that maximize the probability of an eventual high reward outcome, You could unlock truly advanced goal-directed reasoning. The learning signal would finally align with the search strategy itself. A powerful idea. Training the policy not just to pick the next best token, but to manage an entire decision-making portfolio. Something for you to think about as you listen to that next batch of AI news.

14:25Thank you for joining us for this deep dive. We'll catch you next time.

From the publisher

This paper introduces Pass-at-k Policy Optimization (PKPO), a novel Reinforcement Learning technique that shifts the focus from individual sample performance (pass@1) to optimizing the collective utility of a batch, quantified as the maximum expected reward (pass@k). This method is necessary because conventional RL under-utilizes sample diversity, limiting exploration and leading to stalled learning on difficult problems. PKPO's primary technical contribution is the derivation of novel, low-variance unbiased gradient estimators for the pass@k objective, which work robustly for any arbitrary $k$ with both binary and continuous rewards. The authors validate their transformation using open-source models like GEMMA2 and LLAMA3.1 on mathematical and coding benchmarks. Crucially, PKPO enables improved exploration, demonstrated by its ability to unblock learning and achieve superior performance, especially when the optimization target $k$ is annealed during training.

More from Best AI papers explained

All 475 episodes
Pass@K Policy Optimization: Solving Harder Reinforcement Learning ProblemsBest AI papers explained · 14 min
Listen in VO