Representation-Based Exploration for Language Models: from test-time to post-training

12 Jan 2026 · 14 min · 5 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Representation-based exploration (REPEX) for language models to enable deliberate, conceptually novel discovery instead of “sharpening” via RL; covers inference-time selection and post-training reward shaping to prevent diversity collapse.

Guests/backgrounds

No named guests; only host/interviewer voices.

Key claims

Standard RL fine-tuning improves optimization but causes diversity collapse and “sharpening.” REPEX uses the model’s hidden states as an internal map to score novelty via elliptical bonuses, guiding targeted exploration without retraining.

Notable examples

Inference-time “core set” selection reduces verifier calls; reported 50%+ verifier efficiency gains on math/code/logic tasks with Qwen2.5-?; hardest 20% yields max gains, including 3-fold verifier efficiency on hardest Game of 24 cases. Post-training with REPEX preserves/improves pass@K and improves sample efficiency 3.2–13.4x on AE2024 math tasks. Token-by-token diversity is promising but costly.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Dilemma of Exploration vs. Optimization

0:45 to 4:38

Discussion on the limitations of current language models and the need for true discovery.

“You give them feedback on how to make a chair faster.”

Representation-Based Exploration (REPEX)

4:38 to 8:14

Introduction to REPEX and how it enables diversity in language models.

“And that verifier, that's the bottleneck.”

Efficiency Gains in Language Model Verification

8:14 to 10:54

Exploration of how REPEX improves efficiency in verifying model outputs.

“So the key takeaway is that guided exploration is only useful if the guide is already smart.”

Diversity Collapse and Training Integration

10:54 to 13:16

Explanation of the diversity collapse issue and how REPEX can mitigate it during training.

“They just augmented the rewards in the standard RL process with these sequence-level novelty bonuses.”

Future of Knowledge-Guided Exploration

13:16 to 13:30

Discussion on the future applications of knowledge-guided exploration in AI.

“where we have to rely on subjective human feedback and where reward hacking is a constant danger?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to the Deep Dive, where we take complex research and distill it into the powerful insights you need to stay well informed. Today, we are putting on our explore hats and diving deep into the mind of the machine itself. Our mission is all about a core dilemma in modern AI. How do we get these large language models to go past just being excellent refiners of what they already know and turn them into genuine discoverers of new ways to solve problems? It really is the difference between just optimization and true innovation, and that's the frontier. here. When researchers use something like reinforcement learning RL, which is basically feedback training, to fine tune these models, they often seem to hit a wall.

0:42Think of it like a highly skilled craftsman. You give them feedback on how to make a chair faster. They'll sharpen their tools, they'll refine their techniques, and maybe they shave a few seconds off the time it takes to carve a leg. We call that sharpening. They just perfect what they already do. So the feedback is like a coach who's just yelling faster, faster on the same play, but not coming up with a whole new strategy. Exactly. What we want is for the model to look at the problem, maybe the need for a new way to sit. And instead of just making a better chair, it discovers the whole concept of, I don't know, a stool or a beanbag.

1:12True discovery. Right. The quest for solutions that are conceptually novel, not just little variations, but the here complexity of language makes that kind of exploration incredibly difficult. It usually just falls apart into, well, gibberish. Okay, so let's unpack that. If current RL methods are mostly just sharpening the saw, how do you even begin to introduce deliberate exploration that hunt for novel behaviors in this massive, complex world of LMs without just getting nonsense? The key, and this is where the work is so brilliant, is using the model's own internal compass. An internal compass.

1:48Yeah. They're not just relying on random chance. They're using what's called representation-based exploration or REPEX. It's a method that uses the model's hidden states. You can think of those as its internal map of concepts to guide the search specifically toward diversity. That idea of an internal compass is fascinating. Yeah. It almost implies the model knows where its own blind spots are. In a way, yeah. We're moving way beyond just, you know, counting how many times a word shows up. We need to measure conceptual distance. That's the whole point. We need some kind of metric that tells us, is this new response just a small tweak on solution A, or is it a completely new approach, solution B?

2:27The core strategy uses something they call elliptical bonuses to give each response a novelty score. Okay, an elliptical bonus. How does that work in practice? Let's use an analogy. Imagine your model is a massive library, and every response it generates is a new book. You're the librarian trying to build the most diverse collection possible. And I've got a limited budget, so I can only buy a few books. Precisely. So you track the features of every book you already have, genre, narrative style, historical period, you name it, and you plot them in a huge conceptual space. When a new manuscript arrives, you check where its features land on that map.

3:03If it's in a conceptual area where you have basically no other books, it gets a massive novelty bonus. Ah, I see. So that bonus is the mathematical version of the elliptical bonus. It measures the empty space on the model's internal knowledge map. You got it. The bonus is huge when the idea is conceptually far away from everything else it's tried. That makes perfect sense. But how does the model define that feature vector? How does it even know the conceptual genre of its own book, so to speak? It uses its own deep, pre-trained knowledge. The features, the representations, are pulled directly from the language model itself.

3:41Specifically, they average the model's deepest internal processing states, the hidden states, across the whole response. So it's not looking at surface-level words. Not at all. You don't need to get bogged down in the linear algebra. The important thing is that this average captures the holistic meaning of the entire response. It's using its own map of reality, built from training on, you know, huge chunks of the internet, to decide what novel actually means. It turns exploration into a targeted search. Exactly. A knowledge-guided search. Okay, so if this internal map is so powerful, how can they use it, like, right now, without even retraining the model?

4:16Let's talk about the immediate win here, the inference time selection problem. This is all about efficiency. It's about saving time and money right out of the gate. Today, a lot of LLM systems will generate, say, a hundred or even a thousand possible answers to a tough question, like a coding problem. Then they have to run this expensive external process, a verifier, to check which one is actually right. And that verifier, that's the bottleneck. That's what costs the most time in compute. It is. So picture yourself as a CEO. You run a massive brainstorm. You get a thousand ideas back. You can't possibly afford to fully test every single one.

4:52You need a small, diverse, high-quality subset, maybe 10 ideas, a core set, that's most likely to contain the best, most conceptually different solutions. So the REPEX algorithm selects this core set. It iteratively picks the generation that has the highest novelty bonus compared to the ones it has already selected. Oh, so it's history aware. Yes. So if two ideas look different on the surface but are conceptually the same, like two different ways of writing the exact same code, you will only pick the first one. Redundancy gets penalized in that deep feature space. That sounds incredibly efficient.

5:24Yeah. But are you sacrificing quality for diversity? What's the bottom line here? We're talking about verifier efficiency, right? How many samples you need to check before finding a correct answer? The gains are massive, and they translate directly into dollars saved. When researchers applied this to powerful models like the Quinn 2.5 Fortin B model, they got over a 50 % improvement in verifier efficiency. 50%. 50%. Across really hard reasoning tasks, two math, code generation, logic puzzles. Wait, that's not a marginal gain. that fundamentally changes the cost of running these models for high-stakes reasoning.

5:59Finding the right answer with half the work is, it's huge. It is. I mean, if you're a company using an LLM to generate and check complex code for clients, cutting your verification costs in half is a total game-changer. I have to ask, though, does this just come from the model being really clever in how it generates that first big pool of answers? Or is the selection method really the hero here? That's a great question, and they checked for that. They found that RepEx still improved efficiency even when that big pool of candidates was generated with different methods, like low-temperature sampling, which tends to make very similar safe answers.

6:37So it's definitely the selection process. It confirms that the knowledge-guided selection is the key innovation here. The only time it struggled was when the initial pool came from super-high-temperature sampling. Which is just too random. Yeah, it just produces too much noise, too much nonsense. the compass can only guide you if the map has at least some coherent information on it. These results are really compelling. It works. It saves money. But who benefits the most from this? Does every language model get the same kind of boost? Interestingly, no. And the answer tells us something really critical about the quality of that internal compass we were talking about.

7:12The benefits correlate really strongly with the model's initial strength. Okay, so let's go back to our Explorer analogy. Right. Imagine you give this novelty compass, RepX, to two people. One's a total novice, the other is an experienced mountaineer who already has a detailed map of the mountain range in their head. My first instinct is that the novice would benefit more. They have the most new territory to discover. That's the logical assumption, but the data showed the complete opposite. Weaker models, like a little 0.5b parameter model, they saw almost no benefit. Sometimes their performance even got worse.

7:48Why? Because their internal map, their hidden states, was too weak. It was messy. So the compass guided them toward novelty, but that novelty was just incoherence. It's like a novice following a compass they can't read right off a cliff. But the experienced mountaineer? The experienced mountaineer, the stronger models, like the 32B parameter one, they almost always benefited. Their internal map was rich, it was detailed, and that allowed the compass to guide them efficiently to new frontiers that were still logically sound. So the key takeaway is that guided exploration is only useful if the guide is already smart.

8:23That's a perfect summary. Okay, so we know which models benefit. But what about the problems themselves? Is it just making easy problems easier, or is it actually cracking the really tough nuts? Another vital question. Because if it only speeds up things we could already solve, it's useful but not, you know, transformative. The data shows RAPEX shines brightest when the task is hard. The researchers sorted all the questions by difficulty. And RPEX showed the biggest improvements on the hardest questions. For the hardest 20 % of problems on that famous math data set, the efficiency gains were at their maximum.

8:58Wow. And the results on the game of 24 puzzles were even more striking. For the absolute hardest questions in that set, RPEX gave them a three-fold improvement in verifier efficiency on the 5-4 model. A three times improvement on the problems the model is most likely to fail. Exactly. So that implies that when the model really needs to try a completely new approach, when all the obvious paths are dead ends, this internal compass is what finds the conceptual detours that actually work. You've got it. Exploration isn't a luxury. For complex problems, it's a necessity. That's a powerful case for using REPEX at inference time.

9:33But what happens if you actually bake this into the training process? Can you fundamentally change how these models learn during that RL post-training phase? This is the ultimate goal, right? To move from just picking better answers to actually training the model to generate novel, correct answers from the get-go. And we talked about how traditional RL can lead to that sharpening effect. Well, it also creates this other major problem called diversity collapse. Okay, let's define diversity collapse clearly for everyone listening. What does that actually mean? Imagine you're training an AI chef.

10:06Traditional RL rewards it for making a perfect chocolate souffle. At first, the chef experiments, but as it gets more training, it figures out that only one specific recipe, whipping the eggs for exactly 5 minutes and 30 seconds, gets the absolute maximum reward. So it just throws out every other good technique it learned, even if those other techniques made slightly different but still delicious souffles. Precisely. That is diversity collapse. The model gets incredibly good at one specific path to success, and its pass at killer performance. Its ability to solve a problem if you give it, say, 100 tries, actually gets worse than the original base model.

10:44Because it's forgotten how to be flexible. It's fantastic at one thing, but it's lost its creativity. So integrating re-pecs during training must be like giving the AI chef an extra bonus point for being creative or original on top of just making a perfect souffle. That's exactly the mechanism. They just augmented the rewards in the standard RL process with these sequence-level novelty bonuses. So the model gets a little reward for being right and another little reward for being conceptually different from the other solutions in its training batch. You're actively incentivizing breath while it's learning.

11:17You are. And did it work? Did it stop the diversity collapse? It did, completely. Policies that were trained with REPEX preserve or even improve that pass at K-Low Performance for large numbers of tries. The diversity collapse phenomenon was just gone. So the chef still makes the perfect souffle, but it also remembers all its other great recipes. Exactly. And it can invent new ones. This is where the true discovery comes in. For really complex, unseen tasks, like math problems from the AE 2024 competition, the Rebex method was just profoundly more sample efficient at test time. How much more efficient?

11:51It was somewhere between 3.2 and 13.4 times more sample efficient than the standard method. 13 times? 13 times. What that really means is that Rebex was helping the model discover truly novel behaviors. It's pushing it past that sharpening phase into the realm of real conceptual discovery. It learns new ways to win, not just how to run the old plays faster. So if we zoom out and connect all this to the bigger picture, this work really shows that deliberate exploration, as long as it's guided by the model's own internal knowledge, is a practical and scalable path to expanding what these models can do.

12:28Whether it's picking better ideas on the fly or changing how the model learns from the ground up. Right. That internal compass is just essential. And they also looked at taking this guidance even deeper, right? Trying to encourage diversity, not just at the whole response level, but token by token. They did. Initial results are promising. It seems to improve the solve rate on the absolute hardest questions, but the computational cost of checking at every single step is, well, it's pretty high. It still needs a lot of optimization to be practical in terms of raw speed, but the potential is there.

13:00We've spent this whole deep dive looking at how incredibly effective this internal compass is when the reward is verifiable. You either solve the math problem or you didn't. You fix the code or it's still broken. But here's the thought I want to leave you with. How can we incentivize the same kind of high quality knowledge guided exploration in areas without a clear right answer, where we have to rely on subjective human feedback and where reward hacking is a constant danger? How do you reward true creativity there? That is a deep dive for another time.

From the publisher

This paper introduces representation-based exploration, a method designed to help language models discover novel behaviors rather than just refining existing ones through reinforcement learning. The researchers propose using elliptical bonuses derived from a model's internal hidden states to explicitly reward diversity and novelty during both inference and training. Their experiments demonstrate that this approach significantly improves verifier efficiency and pass@k rates across complex reasoning and coding tasks. Notably, the technique mitigates the common problem of "diversity collapse," where standard reinforcement learning causes a model’s responses to become repetitive. By integrating these bonuses into the GRPO post-training pipeline, the authors show that models can achieve superior performance with fewer samples. Ultimately, the work suggests that leveraging a model's own internal knowledge is a practical and effective way to advance its autonomous reasoning capabilities.

More from Best AI papers explained

All 475 episodes
Representation-Based Exploration for Language Models: from test-time to post-trainingBest AI papers explained · 14 min
Listen in VO