RLAD: Training LLMs to Discover Abstractions

29 Oct 2025 · 16 min · 15 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

RLAD (reinforcement learning with abstraction discovery) trains LLMs to solve hard, unseen reasoning problems by first generating concise “reasoning abstractions” (high-level strategies/guardrails) and then solving while being forced to use them. It targets “degenerate exploration,” where long chain-of-thought over-focuses on a first path.

Guests/backgrounds

No specific guest names or credentials are provided in the transcript; it’s a host-led “Deep Dive” discussion.

Key claims

RLAD’s two-player setup (abstraction generator + abstraction-conditioned solver) uses reward shaping so the solver gets zero reward if it solves without an abstraction. This improves breadth (multiple strategies) over depth (more steps).

Notable examples

A modular arithmetic prime congruence problem; abstractions include checking multiplicative inverses (co-prime requirement), transforming using modular quadratic-formula logic, and other abstraction types (caution alerts, productive launch points, blind follow trajectories, structural shortcuts). Results: On AME 2025, RLAD reaches 42.45% average over four abstractions (48.33% best), versus DPO 37.92%; also improves even without test-time abstractions (38.04%).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding LLM Challenges

0:45 to 1:29

Exploring issues LLMs face with complex reasoning tasks and degenerate exploration.

“The model keeps elaborating on that initial path, going deeper and deeper, because current methods often reward that depth.”

Introducing R-L-A-D

1:29 to 2:14

An overview of the R-L-A-D method aimed at improving LLM reasoning.

“It's called reinforcement learning with abstraction discovery, or R-L-A-D.”

Math Problem Example

2:14 to 3:08

A math problem is used to illustrate how R-L-A-D operates with abstractions.

“The source material uses a math problem example.”

Two-Player Game Mechanism

3:08 to 4:38

How the two-player mechanism in R-L-A-D helps generate solutions with guidance.

“It's meta-knowledge that guides the detailed calculation.”

Reward System and Player Interaction

4:38 to 6:06

Detailed explanation of how the reward system enforces the use of abstractions.

“So its entire thought process is guided by that initial strategy?”

Mandatory Guidance in Learning

6:06 to 6:34

The clever design makes ignoring abstraction a losing strategy during training.

“So the only way player two can get positive reward for solving the problem is if it does so while following an abstraction provided by player one.”

Performance Comparison with DPO

6:34 to 8:14

Comparative results of R-L-A-D against traditional reinforcement learning methods.

“It builds the cooperation right into the reward structure.”

Generalization Findings

8:14 to 9:18

Unexpected results showing R-L-A-D models perform better even without abstractions.

“But here's the part that I found genuinely surprising, the real kicker, the generalization finding.”

Strategic Abstraction in Various Domains

9:18 to 9:40

How the R-L-A-D approach improves performance across different tasks.

“It's like it learned how to think about problems more effectively, not just execute steps.”

Optimizing Compute for Performance

9:40 to 10:36

Discussion on balancing compute costs with performance improvements in LLMs.

“That leads perfectly into the practical side.”
Show all 15 chapters

Comparing Depth vs. Breadth Strategies

10:36 to 12:39

Exploring the efficiency of breadth strategies over depth in generating solutions.

“Rather than just generating more solutions for a single strategy.”

Types of Abstractions and Their Impact

12:39 to 14:00

Overview of different types of abstractions used in R-L-A-D and their effectiveness.

“And crucially, player two, the solver, is actually paying attention.”

Understanding Structural Shortcuts in RLAD

14:00 to 14:42

Learn how RLAD training uses structural shortcuts to simplify problem-solving.

“These are more conceptual leaps, like you have a perimeter constraint A plus B plus C P2.”

The Generalization Effect in AI Models

14:43 to 16:00

Explore how RLAD training influences generalization and strategic planning in models.

“Okay, so stepping back, the big picture here with RLAD seems to be adding this explicit high-level strategic planning layer to LLMs.”

Future of AI Reasoning and Strategic Learning

16:01 to 16:18

Discuss the implications of generalization in AI and the future of strategic reasoning.

“How do we really teach these models not just to execute complex steps, but to truly strategize like an expert from the outset?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to the Deep Dive, where we shortcut you to the core of complex topics. Thanks. Today we're tackling something really fascinating at the edge of AI research. We're looking at a big challenge for large language models, LLMs, how they handle, well, really hard reasoning problems, especially ones I haven't seen before. Right. We all know LLMs can generate incredibly long, detailed responses, these chains of thought. You ask it a question, it outputs step after step after step. Which sounds great, but the problem comes with truly complex tasks, like, say, advanced math problems. Sometimes that long chain of thought just goes off the rails.

0:37Exactly. It's something researchers call degenerate exploration. Think of it like getting stuck on the first idea you have for solving a puzzle, even if it's leading you nowhere. The model keeps elaborating on that initial path, going deeper and deeper, because current methods often reward that depth. So it digs itself into a hole, essentially. It's underthinking the overall strategy because it's so focused on the next tiny step. Precisely. Standard reinforcement learning, or RL, tends to encourage finding one successful path, however convoluted. This is good for refining known procedures, but bad for exploring different strategic angles.

1:14Which limits how well these models can generalize to new problems where the old tricks don't work. Yeah, it limits the breadth of their thinking. Okay, so that sets the stage for our deep dive today. We're looking at a new approach designed specifically to fix this strategic weakness. It's called reinforcement learning with abstraction discovery, or R-L-A-D. R-L-A-D. The core idea is pretty intuitive, actually. It teaches the LLM to first propose a high-level strategy before it even starts trying to work out the detailed steps. Like creating an outline before writing the essay. Exactly. And these high-level strategies are called reasoning abstractions.

1:50They're basically concise, natural language descriptions of useful knowledge. It could be procedural know-how, like use this specific formula, or maybe a strategic framing like treat this variable as a parameter, or even just a factual reminder, a check you need to do. Sort of like the hints you might get on a tough exam question, the ones that point you in the right direction. That's a great analogy. They act as high-level sub-goals or guardrails. Let's make this concrete. The source material uses a math problem example. Determine the smallest positive prime p zesha, which satisfies the congruence p plus p1 equiv 25 p monocosin.

2:26Okay, that's not trivial. Requires some specific knowledge about modular arithmetic. Right. A standard LLM might just start plugging in numbers or applying theorems somewhat randomly. But with our LAD, this system would first generate an abstraction, maybe something procedural-like. Transform the congruence using the quadratic formula adapted for modular arithmetic. That's a clear strategic path. Or maybe something preventative. Definitely. A really crucial abstraction might be a warning. Hold on. Before you use Exodon modular math, you must check if a multiplicative inverse actually exists. Make sure$6 and the modulus 143 are co-prime.

3:03Ah, okay. That avoids a huge potential dead end if the inverse doesn't exist. Exactly. It's meta-knowledge that guides the detailed calculation. It prevents wasted effort. So how does this RLED system actually work? You mentioned it's like a two-player game. Yeah, they frame it as a cooperative two-player game, which I think is a helpful way to picture it. You have two LLMs trained together. Let's call them Player 1 and Player 2. Okay, Player 1. Player 1 is the abstraction generator. Its job is to look at the problem,$6, and propose one or more of these reasoning abstractions, these strategies,$6.

3:36Does it just pull these out of thin air? How does it know what a good strategy looks like? Good question. It doesn't start from scratch. They warm start this generator. Basically, they pre-train it using supervised fine-tuning SFT. Meaning they show it lots of examples. Right. They fed it problem-abstraction pairs. And crucially, these initial abstractions were created by summarizing multiple successful solution attempts from a more powerful model. Ah, so it learns what useful hints look like by seeing strategies that actually led to correct answers. Exactly. That warm start is key to ensuring player one proposes things that are, you know, actually helpful and not just random sentences.

4:15It grounds the strategy in known successful patterns. Okay, makes sense. So player one proposes a strategy. What about player two? Player two is the abstraction condition solution generator. This is the model that does the heavy lifting, generating the step-by-step solution. But here's the crucial part. it generates the solution condition on both the original problem result and the specific abstraction Zollers it received from player one. So its entire thought process is guided by that initial strategy? Supposedly. But here's where the cleverness comes in. You might ask, what stops player two from just ignoring the hint if it thinks it knows better?

4:52Or if the hint isn't perfect? Yeah, that seems like a potential problem. How do you force it to actually use the abstraction? They tackled this with a really interesting twist in the reward system during training. It's designed to make ignoring the hint impossible, or at least unrewarding. Okay, how so? First, let's look at Player 1, the abstraction generator. Its reward isn't based on how nice the abstraction sounds. It's rewarded based purely on the average success rate of Player 2 when Player 2 uses that specific abstraction. So, if Player 1 suggests a strategy that helps Player 2 solve the problem correctly more often, Player 1 gets a higher reward.

5:29Exactly. It's rewarded for the utility of its strategic advice. Simple as that. Okay, that makes sense for player one, but how do you enforce it on player two, the solver? How do you stop it from just going its own way? This is the key mechanism. The reward for player two, the solution generator, is modified. If player two generates a solution without being given an abstraction, its reward for that solution is explicitly set to zero. Wait, zero. Even if it gets the right answer? Even if it gets the right answer. If no abstraction was provided for that training run, the binary success reward is zeroed out.

6:03It still gets, like, regularization penalties, but no positive reinforcement for the correct answer itself. Wow. Okay. So the only way player two can get positive reward for solving the problem is if it does so while following an abstraction provided by player one. That's the core idea. It structurally forces player two to learn how to use the abstractions. actions. It learns that the path to reward requires engaging with the strategic guidance. It basically makes ignoring the hint a losing strategy during training. That is really clever. It builds the cooperation right into the reward structure.

6:37It's not just suggesting strategy, it's making it mandatory for reinforcement. Yeah, it creates this environment where strategic planning isn't just a nice to have, it's fundamental to the learning process. Okay, So we have this elegant two-player setup with this mandatory guidance system. Did it actually work? How did RLAD stack up against methods that just focus on making the chain of thought longer or better, like the DPO method mentioned in the sources? The results are pretty stark, actually, especially on tough benchmarks. They tested it on AME 2025. That's a set of really difficult math competition problems.

7:11RLAD showed an average improvement of 44 % over state-of-the-art RL methods like DPO that just focus on that deep single chain of thought. 44%. That's not incremental. That sounds like a step change. It really is significant for this kind of complex reasoning task. Let's look at the numbers from Table 2 in the paper. The base model they used, QUIN 3 1.7B, scored about 33.75 % on AME 2025. Okay, baseline. Then they applied DAPO, that standard RL fine-tuning focused on depth. That pushed the score up to 37.92%. An improvement, but, you know, modest. Right. But the RLAD model. When tested using the average result from four different abstractions, it generated, so exploring four strategic paths, its score jumped to 42.45%.

7:58Wow. And get this. If they cherry-picked the best result obtained from those four strategies, the success rate hit 48.33%. So exploring different high-level strategies is way more effective than just trying harder down one path. The breadth beats depth here. It strongly suggests that, yes. Yeah. But here's the part that I found genuinely surprising, the real kicker, the generalization finding. What do you mean? Well, you'd expect the model to do better when it gets these helpful abstractions, right? Yeah, that's the point. But the gains weren't only when using abstractions at test time. Even when they tested the RLED-trained model without giving it any abstraction, the WOAB's condition, it still performed better than the DPL model.

8:35Hold on. So the model trained with RLAD, even when it wasn't explicitly using an abstraction for a specific problem, was still a better solver than the model trained just on refining solutions. That's right. In that WOABs test on AIM 2025, the RLAD model scored 38.04%. Now that's only slightly better than DAPIO's 37.92%, but the implication is huge. What's the interpretation there? It suggests that the training process itself, being forced to constantly generate and then use diverse high-level strategies, actually changed the model's underlying reasoning capabilities. It didn't just learn to follow instructions.

9:11It seems to have internalized some aspect of strategic thinking. So the scaffolding helped build a stronger foundation, even when the scaffolding is taken away later. That seems to be the case. It's like it learned how to think about problems more effectively, not just execute steps. And this isn't just for math. They mention this approach, improved performance by about 30%, on average, across a whole range of tasks. There's the seven different ones, spanning things like healthcare, legal reasoning, even web security. It seems this idea of strategic abstraction is pretty domain general. That leads perfectly into the practical side.

9:46When you're actually using these models, there's always a trade-off between performance and compute cost. More compute usually means better answers. Right. And traditionally, more compute meant sampling more solutions. You just run the model many times, hoping one attempt works out. That's optimizing for depth. But RLED gives us another option. Exactly. It introduces this idea of optimizing for breadth. You can spend your compute generating more different strategies, more abstractions, rather than just more solutions following the same implicit strategy. It's presented as a new sort of orthogonal axis for scaling.

10:20A different dimension to improve performance. So if you have a fixed compute budget at inference time, where's the best place to spend it, according to this research? More solution attempts or more strategic exploration? The findings in Figure 5 are pretty clear on this. As your total compute budget goes up, you get more bang for your buck by allocating that extra compute towards generating more diverse abstractions first. Rather than just generating more solutions for a single strategy. Yes. Investing in breadth exploring different strategic pathways via abstractions seems to yield greater improvements than putting the same extra compute into depth refining solutions within one pathway.

10:57They had that two notters only two versus new notter comparison, right? Can you walk us through that? Sure. Imagine new R is, say, 16. You have a compute budget roughly equivalent to generating two dollars, which are 16 times 16 equals 256 cylinder solutions. Okay, 256 units of compute. Right. The depth approach, like DPO, would just sample 256 solutions directly from the problem, hoping one works. The RLAD breadth approach is something different. It uses part of the budget to sample no 106 and 106 different abstractions, strategies. Okay, 16 potential plans. Then for each of those 16 abstractions, it samples 116-cellular sunny solutions conditioned on that specific abstraction.

11:36So$16 times 16, it involves 256.66 total solutions generated, the same overall compute cost. Got it. Same compute, different allocation. And the results. The RLAD strategy sampling strategies first, then solutions consistently did better. Using NL16TOL1 as the example from Table 3, RLAD achieved a PASIC case score of 0.71. Compared to? Compared to only 0.65 for the pure solution sampling approach, DAPO, using the same total compute. So guiding the search with abstractions was significantly more efficient. It really focuses the search effort, and we know these different abstractions genuinely lead to different solution paths.

12:17It's not just generating fluff. They check that too, looking at semantic similarity. When you generate solutions conditioned on, say, four different abstractions, the resulting solutions are much less similar to each other based on text embeddings compared to just generating four solutions without any specific abstraction. Meaning the model is really exploring distinct reasoning pathways prompted by the different strategies. Exactly. And crucially, player two, the solver, is actually paying attention. They measured the adherence rate. How often the solution actually followed the given strategy.

12:47Yep. And that rate was highest around 43 % when the solution generator was explicitly conditioned on an abstraction. It confirms that the training mechanism we talked about, the zero reward for ignoring hints, actually worked. The solver learned to adhere to the plan. Okay, this is really compelling. Let's maybe quickly touch on what these useful abstractions actually look like. the paper categorized them right yeah in appendix D they broke them down into roughly four types which gives you a good feel for them first you have caution alerts like warnings pretty much things like always check for forbidden values and denominators before and after you manipulate an equation basically avoid common algebraic mistakes okay useful guardrails what else then there are productive launch points these are about reframing the problem or suggesting a smart starting move.

13:34For example, pick one variable like$6, set it to a parameter, a dollars, and try to express everything else in terms of dollars. That can simplify things a lot. A strategic opening move. Right. Third is the blind follow trajectory. This is more like a recipe, a step-by-step procedure. Something like, remember, the mean is the sum divided by the count. Use this relationship to switch back and forth between sums and averages to solve for the unknowns. Almost like a mini algorithm? Kind of. And finally, structural shortcuts. These are more conceptual leaps, like you have a perimeter constraint A plus B plus C P2.

14:10Use that immediately to eliminate one variable and simplify the whole system. It leverages a high-level insight. Interesting categories. Did the RLAD training favor certain types? Yes, that was another finding. The training process, by rewarding success, naturally shifted the distribution of abstractions the generator proposed. It started producing more of those effective blind follow trajectory types, the ones that lay out a clear, actionable path to the solution. So it learned not just to propose any strategy, but strategies that are actually likely to work and be followed correctly. Precisely.

14:42It optimized for implementable success. Okay, so stepping back, the big picture here with RLAD seems to be adding this explicit high-level strategic planning layer to LLMs. Yeah, it's creating this, as we said, orthogonal axis for improvement. It's not just about longer chains of thought or more parallel attempts anymore. It's about smarter guided exploration. Which brings us back to maybe the most thought-provoking piece, that generalization effect. the fact that training with abstractions makes the model better, even when it doesn't get an abstraction later on. Right. It really sticks with you.

15:17It suggests that the training isn't just teaching the model to follow instructions, but embedding some kind of deeper procedural or strategic knowledge into the model itself. The staff holding comes down, but the building is stronger. It learns something fundamental about problem solving. It certainly seems that way. And that leaves us with a really interesting question, maybe for you, the listener, to think about, what does that generalization really mean? Is the model developing a truly general strategic planning ability? It can apply anywhere. Or is it more like it's built up a really good internal cheat sheet of techniques and heuristics that it now accesses more effectively?

15:55Is it learning to plan or just learning better plans? Exactly. Figuring out the mechanisms behind that generalization and how to amplify them, that seems like a huge next step. How do we really teach these models not just to execute complex steps, but to truly strategize like an expert from the outset? A fascinating glimpse into the future of AI reasoning. Thanks so much for unpacking that for us. My pleasure. It's really exciting work.

From the publisher

This paper introduces a novel two-player reinforcement learning (RL) framework, RLAD, designed to enhance the reasoning capabilities of large language models (LLMs). This framework jointly trains an **abstraction generator** and an **abstraction-conditioned solution generator** to propose and utilize **concise natural language descriptions of procedural and factual knowledge** called "reasoning abstractions." The core objective is to move beyond conventional chain-of-thought methods, which often result in degenerate exploration, by teaching models to discover **high-level subgoals or strategies** that guide the solution process. Experimental results on various math and non-math reasoning benchmarks demonstrate that RLAD significantly **improves accuracy and exploration diversity** compared to prior RL approaches, with performance scaling more efficiently when compute is allocated toward generating diverse abstractions rather than solely increasing solution length or count.

More from Best AI papers explained

All 475 episodes
RLAD: Training LLMs to Discover AbstractionsBest AI papers explained · 16 min
Listen in VO