Representation-Based Exploration for Language Models: From Test-Time to Post-Training

18 Oct 2025 · 17 min · 11 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Representation-based exploration (REPEX) for language models, combining RL with a novelty bonus to overcome the “sharpening constraint” and “diversity collapse,” improving discovery beyond just optimizing known strategies.

Guest backgrounds

No guest names or bios appear in the transcript; only hosts discuss the work.

Key claims

Standard RL (e.g., PPO/GRPO) boosts easy success but fails to discover skills with near-zero initial success; REPEX adds correctness plus representation-based novelty to prevent diversity collapse. Novelty is computed from hidden-state “conceptual fingerprints” and an “elliptical” distance bonus in representation space.

Notable examples

Inference-time selection for math/code/reasoning (N≈10^24 pool; select K≈8–256 diverse candidates) improves verifier efficiency by 50%+ on GSM8K, MBPP+, MILTAF, Game of 24; hardest 20% yields up to 3x. Post-training: on AIME 2024, QUEN 2.5B with REPEX+GRPO reaches pass@80 equivalent to GRPO needing pass@256 (≈3x sample efficiency); beats “unlikeliness” by 2.1–4.1x and standard GRPO by 3.2–13.4x.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Reinforcement Learning and LLMs

0:45 to 2:33

Exploration of the intersection between reinforcement learning and language models, questioning their capabilities and limitations.

“Our mission today is really to understand how this technique, REPEX, is specifically designed to push LLMs out of that, well, that frustrating rut, what we're calling the sharpening constraint.”

Representation-Based Exploration (REPEX)

2:33 to 4:17

Introduction to representation-based exploration and its significance in overcoming model constraints.

“So if the existing methods are kind of trapping the model, we need something to explicitly force it to look elsewhere.”

Mechanics of REPEX: Novelty Bonus

4:17 to 6:07

Detailed explanation of how REPEX uses internal language representations to calculate a novelty bonus.

“Those are the vectors inside the network.”

Testing REPEX: Inference Time Selection

6:07 to 8:11

Discussion on using inference time selection to isolate the effects of REPEX in practical scenarios.

“But what's really neat about the engineering here is the scalability.”

Results and Implications of REPEX

8:11 to 10:13

Analysis of the results from REPEX integration and its impact on efficiency and problem-solving.

“It trusts the prior knowledge baked into those representations to distinguish novel and maybe right from novel and just nonsense.”

Diversity Collapse in RL Training

10:13 to 12:15

Exploration of diversity collapse in reinforcement learning and how REPEX addresses this issue.

“Okay, and we should probably briefly touch on that other approach mentioned, token-level exploration, trying to build diversity right into the generation process.”

Sample Efficiency with REPEX

12:15 to 14:01

Demonstration of how REPEX enhances sample efficiency in RL training compared to traditional methods.

“It is a fundamental problem with naive optimization in these complex spaces.”

Exploring the Effectiveness of Repax

14:01 to 14:48

Learn how Repax outperforms other methods in sample efficiency.

“It directly translates into compute savings, faster inference times for these really demanding reasoning tasks.”

Key Takeaway on Guided Exploration

14:49 to 15:40

Understand the practical implications of guided exploration in language models.

“If someone listening needs the single most important takeaway today, what is it?”

Future Applications and Challenges

15:41 to 16:51

Explore the potential and challenges of representation-guided exploration.

“Which naturally leads to the final provocative thought, right?”
Show all 11 chapters

Concluding Thoughts on Research Frontiers

16:52 to 17:03

Reflect on the future landscape of AI research and its implications.

“Using the model's own mind to guide it, but ensuring that guidance leads somewhere genuinely productive, not just somewhere different.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to the Deep Dive. We are taking your collection of sources and cutting straight through the noise. Delivering those critical insights. Exactly. Today we're diving deep into that frontier where reinforcement learning RL meets large language models, LLMs. It's a really hot area right now. It really is. And the core question driving this deep dive, it feels fundamental. Are we actually teaching these models something new with RL? Or are we just making them maybe obsessively good at the few things they already kind of knew how to do? Right. Is it true discovery or just better execution? Sharpening the existing tools.

0:36That's the million-dollar question. So your source material, and this is based on recent work from Princeton and Microsoft Research, it tackles this head-on. They introduced this novel mechanism, representation-based exploration, REPEX for short. Our mission today is really to understand how this technique, REPEX, is specifically designed to push LLMs out of that, well, that frustrating rut, what we're calling the sharpening constraint. And that constraint is absolutely the critical context here. You know, RL algorithms, they're fantastic tools. They promise to let LLMs find complex reasoning path, novel code, all on their own.

1:14That was the dream, wasn't it? Economist discovery. It was. But what we've actually found, especially in tricky domains like advanced math or, you know, sophisticated coding, is that standard RL methods, things like PPO or GRPO, they often just, well, they struggle. They don't really move past what was already baked into the pre-trained model. Why is that? What's holding them back? Well, these algorithms, they're designed to reward immediate success, right? So they find the quickest, most reliable path to getting that reward signal. Okay, so they take a skill the model might have, say, a 1 % success rate on.

1:47Yeah. And they push that up to 99%, which sounds great. It does sound great, but here's the catch. They never discover the skill the model has a 0 % chance of finding initially, even if that skill is crucial for the really hard problems. Okay, so it's like taking a student who's okay at geometry and making them perfect at quadratic equations. Exactly. But never actually teaching them calculus because they never tried it in the first place. That's a perfect analogy. If the model, when it first explores, only tries strategy A and strategy B, RL will just optimize A and B. It won't even consider strategy C even if C is the breakthrough needed for the next level.

2:24That's the sharpening constraint in a nutshell. And data scale isn't the whole answer anymore. No. It's becoming clear the search strategy itself is the bottleneck now. Okay. So if the existing methods are kind of trapping the model, we need something to explicitly force it to look elsewhere. We need deliberate exploration. Precisely. Deliberate exploration. It's about adding an incentive not just for being right, but for being different. Yeah. For novelty. So reward correctness, fine, but also reward newness. Correctness plus novelty. That's the mix that encourages genuine discovery, not just optimization.

3:02The challenge there seems huge, though. I mean, language is just vast, combinatorially explosive. It is. How do you even quantify novelty in, say, a 100-token response? How does the algorithm know if this math explanation is genuinely different from when it generated like a thousand samples ago? That is the central problem this research gets at. And the guiding hypothesis is, well, it's kind of brilliant in its simplicity. Okay. The idea is the LLM already has this immense amount of prior knowledge. It's all encoded in its network layers, right? Its representations. Right. It's internal understanding.

3:40So can we use that internal understanding? Yeah. The model's own representation of meaning to guide its external search for new solutions. Can the model's own brain be the map? Okay, interesting. Using its internal map to explore new territory. So let's unpack representation-based exploration, the mechanism they built for this. Let's do it. It sounds abstract, maybe a bit academic, but your sources say it's surprisingly practical. It avoids some common RL headaches. Yeah, that's a key point. It sidesteps the need for complex auxiliary networks that you often see in deep RL. It's more self-contained.

4:13What's the core trick? The key innovation is using the LLM's own internal language, those hidden states, to calculate a novelty bonus. Hidden states, right. Those are the vectors inside the network. Exactly. When the LLM processes text, those hidden states are essentially the feature vectors the model uses to understand the meaning, the concepts. Okay, so instead of just looking at the output text, the words on the page, and asking, is this different? They look inside the model's brain and ask, is the internal conceptual representation of this answer significantly different from previous ones?

4:47You got it. For any given prompt and the response it generates, Rebex extracts a single feature vector. Typically, it's the average of all the token representations from the last hidden layer. So like a compressed conceptual fingerprint of the whole response? A conceptual fingerprint. That's a great way to put it. Highly compressed captures the essence. And then how does that fingerprint translate into a bonus? Right. So this is where the diversity measure comes in. They call it elliptical bonuses. It sounds fancy, but think of it like this. The system keeps track of a sort of historical conceptual memory cloud.

5:22All the different types of solutions that's generated before represented in that high dimensional hidden state space. A conceptual GPS remembers where it's been conceptually. Exactly. Knows where it's been. Now, if the model generates a new answer and its conceptual fingerprint lands right in the middle of that existing memory cloud. Then it's not really new conceptually. Redundant. Precisely. The bonus is small. But if that new fingerprint lands far outside the existing cloud. Uh-huh. That's conceptually novel. Big bonus. Big bonus. It explicitly encourages moving into new conceptual territory.

5:57It penalizes redundancy in that representation space. That sounds computationally intensive, though, keeping track of that whole cloud, calculating distances for every new generation. You'd think so. But what's really neat about the engineering here is the scalability. They use some clever linear algebra, specifically something called rank one updates. It allows them to efficiently manage and update that memory cloud, that inverse covariance matrix, without doing massive recalculations every single time. So they can handle the sophisticated history-aware diversity measure, even with today's huge LLMs.

6:30That's the key. It makes the sophisticated exploration strategy actually practical to deploy. Okay, that makes sense. Before we get to how this plugs into the full RL training, let's talk about that first test case they used, the inference time selection problem. Ah, yes. Seems like a smart way to isolate the effect, prove the concept without the messiness of training dynamics. Absolutely brilliant move. Inference time selection really isolates the power of REPEX diversity signal itself. The setup is pretty straightforward. You get the model to generate a big pool of responses. Let's say N equals 10 to 24, a lot of possible answers.

7:07Okay. Now, the traditional way might be just to randomly pick, say, KK8 samples from that big pool and check if any of them are correct. That's standard sampling. Makes sense. But with REPEX, you don't sample randomly. You use that diversity bonus calculation we just talked about to select the K-8 responses that are the most conceptually diverse subset within that pool of 1024. Ah, so you're picking the eight answers that cover the widest range of ideas or approaches. Exactly. And the measure of success here is verifier efficiency. Basically, how few of those K samples do you need to check on average before you find a correct one?

7:41OK, but here's my pushback. If you're actively rewarding things that are different, aren't you just asking the model to generate more weird, wrong answers? Yeah. More conceptual garbage. That's the crucial question. And the answer hinges entirely on the quality of the base model's hidden states. Remember, REPEX rewards what the LLM itself registers internally as meaningfully different. So it's trusting the pre-trained model's own internal compass for what constitutes plausible different versus just weird different. Precisely. It trusts the prior knowledge baked into those representations to distinguish novel and maybe right from novel and just nonsense.

8:20And did that trust pay off? What did the results show? They really did, especially with stronger base models. On tough reasoning benchmarks, GSMA-K, math, MBPP +, MILTAF, Game of 24, REPEXP, delivered over a 50 % improvement in verifier efficiency compared to just random selection. 50 percent. Wait, so you only need to check half as many samples, roughly, to find a correct answer? In many cases, yes. That's a massive operational gain. Think about the compute budget for verification. Cutting it potentially in half is huge if you're running these things at scale. No kidding. And you mentioned it works better with stronger models.

8:54Yeah, the sources highlight a really clear correlation there. Repax really starts to shine when the base model itself is already pretty capable. So if the model's internal map is bad, the representations are weak, the GPS is broken, basically. That's right. They tested weaker models, like a QUIN 2.5 with only 0.5 billion parameters. It saw little benefit, maybe even got slightly worse. Yeah. But as the models got stronger, up to 32 billion parameters, the benefits grew almost uniformly. It really confirms the idea. REPEX helps the model search its existing knowledge creatively. It doesn't magically create knowledge that isn't there.

9:30Interesting. It unlocks latent potential, maybe. That seems to be the case. And crucially, the biggest impact was on the hardest problems within those data sets. Oh, like the real head scratchers. Exactly. Take the top 20 % hardest questions on the math data set. Repex showed its biggest gains there. Or the hardest problems in Game of 24, using Repex with a model like 5.4 gave a 3x improvement in verifier efficiency. Three times better on the toughest stuff. Yeah. It helps the model really stretch its capabilities right where it matters most at the edges on the problems it would normally fail on.

10:06So it's not just polishing the easy wins, it's actually helping tackle the previously unsolvable ones. That's what the inference time results strongly suggest. Okay, and we should probably briefly touch on that other approach mentioned, token-level exploration, trying to build diversity right into the generation process. All right, that was an interesting side experiment they reported. Instead of selecting diverse outputs after generating a big pool, they tried modifying the probabilities, the logits, during the autoregressive generation, token by token, to encourage diversity right from the start.

10:37Sounds computationally heavy. It is, I know. But it showed promise. They found that if you gave it a really large sampling budget like K512 or more, it significantly boosted the solve rate on the hardest problems. So the potential there is, maybe eventually you don't need to generate huge pools, just generate a smaller number of inherently diverse solutions. That's the revolutionary hope, yeah. Generate diversity up front, potentially saving a lot of computation down the line. Still early days, but promising. Which brings us neatly to the main event, putting REPEX into the full RL post-training loop.

11:10Right, the ultimate test. Now, standard RL training, like GRPO, apparently suffers from this thing called diversity collapse. Can you explain what that is and why it's bad news? Yeah, diversity collapse is a known issue. It happens when the RL training process, in its relentless drive to maximize that reward signal, it basically preens away almost everything except the absolute highest reward paths it finds early on. The model becomes hyper-focused, hyper-optimized on just a few ways of solving things. It loses its breadth. Its exploratory spirit gets crushed by optimization. Exactly. It forgets how to try different things.

11:46And the result can be pretty catastrophic, especially for metrics like pass at K, where K is large. Pass at K, meaning the success rate when you check different answers. Right. So if you train a model with standard RL and then test it by asking for, say, 256 different solutions, K256 act, that RL trained model often performs worse than the original base model you started with because it's lost the ability to generate a wide variety of attempts. It's become sharp and incredibly narrow. Wow. So the training actually degrades its ability to find a correct answer if you give it enough tries. It seems fundamentally flawed.

12:21It is a fundamental problem with naive optimization in these complex spaces. But here's the kicker. Integrating Revex directly into the RL training pipeline. Did it fix it? It completely solved the problem according to the results. By training the model to value both correctness and novelty via the RepEx bonus, they eliminate diversity collapse entirely. So the pass at K's scores didn't drop off anymore for large K. They were maintained or even improved across the board, even for very large values of K. It's strong evidence that RepEx was fundamentally pushing the model beyond just sharpening.

12:55It's enabling genuine discovery within the RL framework. Okay. We absolutely have to talk about the numbers here. The efficiency gains you mentioned earlier, especially on that AIMA 2024 competition task. Can you lay that out? Right. This is probably the most striking result for anyone thinking about deploying these models. It's about sample efficiency at test time. On that AID 2024 math competition data set, they took a QUEN 2.57 billion parameter model and post-trained it using REPEX integrated with GRPO. They found that this REPEX-trained model achieved a certain level of performance, its pass at 80 score, was equivalent to the performance level that the same base model trained with standard unmodified GRPO only achieved a pass at 256.

13:40Hold on. Let me make sure I understand that. To get the same success rate, with the standard RL method, you'd need to generate and check 256 different solutions. Right. But with the RepEx enhanced RL method, you only needed to generate and check 80 solutions. That's exactly what the results showed. A 3x improvement in test time sample efficiency. Three times fewer samples needed for the same result. That's not just incremental. That changes the economics. It absolutely does. It directly translates into compute savings, faster inference times for these really demanding reasoning tasks. It makes achieving high performance much more practical.

14:12Were the game consistent compared to other methods, too? Yeah, they compared it against another baseline called unlikeliness, which also tries to reward novelty, but in a different way. Repax was 2.1 to 4.1 times more sample efficient than unlikeliness, and compared to the standard unmodified GRPO, Repax was between 3.2 and 13.4 times more sample efficient, depending on the specific task and metric. 13 times, wow. It really underscores how effective using the model's own internal representations is for guiding this kind of exploration compared to more external or heuristic approaches. Okay, this has been a fantastic deep dive.

14:50Let's try and synthesize this. If someone listening needs the single most important takeaway today, what is it? Hmm, the big picture. Yeah, the core message. I'd say it's that deliberate guided exploration isn't just some theoretical nice-to-have. It's practical, it's scalable, and it works. This REPEX method effectively turns the LM's own internal knowledge, its hidden states, into a kind of GPS. Yeah, conceptual GPS. Right. And that lets it navigate beyond just optimizing what it already knows and actually discover novel, useful ways to solve problems. Which means, for you, the user, cheaper, faster, more reliable results on tough tasks.

15:26I think that's spot on. Yeah. The key insight really is that these LLMs often have the knowledge. They just need help searching that knowledge space creatively. REPEX provides that help by using the models and understanding of concepts to guide the search. Yeah. And we see the payoff. Big improvements in performance, huge gains in efficiency across math, code, reasoning. It seems broadly applicable. Very much so. Which naturally leads to the final provocative thought, right? Where does this go next? Always the fun part. What's nagging at you? Well, all this success relies on having clear, verifiable rewards.

15:59You know, the math answer is right or wrong. The code compiles or it doesn't. Binary feedback. Exactly. But we saw that early hint of using exploration during generation. How could we take this powerful idea of representation-guided exploration and apply it to domains where the feedback is much fuzzier, subjective, like creative writing or generating genuinely engaging dialogue or even complex scientific hypothesis generation? Where there's no single right answer to verify. Precisely. And the flip side of that coin, As we get better at encouraging novelty, how do we make sure we don't fall into reward hacking?

16:36Where the model finds clever, novel ways to get the exploration bonus without actually producing anything useful or high quality. Exactly. Balancing true, valuable novelty with quality and avoiding gaming the system, especially in domains without clear verification. That feels like the immediate frontier this research opens up. Fascinating. Using the model's own mind to guide it, but ensuring that guidance leads somewhere genuinely productive, not just somewhere different. Lots to think about there.

From the publisher

This paper investigates the effectiveness of deliberate exploration in enhancing the reasoning capabilities of large language models (LLMs) trained with reinforcement learning (RL). The authors propose and evaluate a novel representation-based exploration (RepExp) strategy, which uses a bonus derived from the LLM's hidden states to encourage the discovery of diverse and novel behaviors. The study employs a two-pronged evaluation methodology, first testing RepExp in an inference-time setting for selecting diverse responses and then integrating it into the RL post-training pipeline. Key findings indicate that this exploration method significantly improves verifier efficiency and mitigates the "diversity collapse" phenomenon observed in standard RL methods, suggesting that the approach moves beyond merely sharpening existing model capabilities. The results show RepExp provides substantial improvements in pass@k rates and is especially beneficial for stronger models and harder reasoning problems across various tasks like MATH and GSM8K.

More from Best AI papers explained

All 475 episodes
Representation-Based Exploration for Language Models: From Test-Time to Post-TrainingBest AI papers explained · 17 min
Listen in VO