Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMs

27 Nov 2025 · 31 min · 14 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Prompted Policy Search (PropES) uses an LLM as the optimizer in reinforcement learning, merging linguistic “manuals” and numerical reward history to speed policy search and add interpretability (“textual gradients”).

Guest backgrounds

No guests are named in the transcript; it’s a solo host-style episode (“The Deep Dive”) discussing the research.

Key claims

PropES bypasses gradient-based optimizers (PPO/SAC) by having the LLM propose new policy parameters from in-context history (theta-reward pairs). PropES+ adds environment/policy semantics and expert hints to guide optimization. Textual justifications are claimed to reflect the LLM’s actual optimization reasoning. Language can help or hurt depending on whether it matches true dynamics (e.g., Frozen Lake).

Notable examples

Inverted pendulum reached maximum reward; Swimmer improved over baselines; NIMH (misère stick game) became top performer with rules in plain English; Frozen Lake got worse due to an incorrect deterministic “common sense” bias. Mountain car hints improved rewards and reduced variance; cliff walking required state-action-specific guidance. Scaling uses random projection with QR decomposition to optimize compressed parameters for deep RL (Swimmer: 150 vs PPO 92.5).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding RL Limitations

0:45 to 1:40

Discussion on the shortcomings of traditional reinforcement learning methods.

“And this is a massive hurdle in the real world.”

Introduction to Prompted Policy Search

1:40 to 2:36

Exploration of the Prompted Policy Search (PropES) framework and its significance.

“It unifies the numerical rewards with that linguistic context.”

Policy Optimization with LLMs

2:36 to 4:23

Detailed explanation of how the LLM drives policy optimization in PropES.

“If the LLM isn't just suggesting a reward and it's not making real-time decisions, what is it actually doing to steer the agent's learning?”

Textual Gradients and Their Importance

4:23 to 6:08

Introduction of textual gradients and their role in making RL transparent.

“Is the LLM involved in real-time control?”

Comparing Baseline PropES and Language Integration

6:08 to 8:06

Contrasting the performance of Baseline PropES versus PropES Plus incorporating language.

“Because the LLM works with language, it can actually justify its own mathematical decisions.”

Performance of PropES in Numerical Optimization

8:06 to 10:34

Results of PropES as a numerical optimizer across various benchmarks.

“We've got the baseline, which is called PropES, which is purely numerical, and then PropES plus a CARO, which adds in all the rich language.”

Impact of Language in Optimization Tasks

10:34 to 12:31

Assessment of how semantic language improves the performance of PropES Plus.

“This is the really crucial integration step.”

Performance Gains with Prop.ES Plus

14:00 to 17:40

Learn how adding semantic descriptions improves performance in LLM tasks.

“In that one, Prop ES showed really substantial and sustained performance gains over methods like A to C and PPO across the entire training run.”

The Impact of Linguistic Context

17:40 to 23:00

Explore how linguistic context can enhance or hinder performance in RL environments.

“So moving on from just the basic environment descriptions, the researchers also tested what happens when you give the LLM explicit expert advice, what they call hints.”

Memory and Cost Considerations

23:00 to 27:20

Understand the role of memory in optimization and cost implications of using LLMs.

“And the big question is, did it just learn Mountain Car or did it learn a more general transferable skill?”
Show all 14 chapters

Future of Prop.ES and Deep RL

27:20 to 28:05

Discuss the challenges and opportunities in scaling Prop.ES to deep reinforcement learning.

“Those textual justifications, the textual gradients, they provide an inherent audit trail.”

Exploring Reward Shaping and Its Challenges

28:05 to 29:55

Learn about reward shaping in reinforcement learning and the challenges of scaling to complex models.

“With PropYes, we can potentially just bypass all that complexity.”

The Role of LLMs in Policy Optimization

29:55 to 30:28

Discover how large language models can aid in policy optimization and their implications for AI.

“It's moving us toward creating human-aligned RL agents that can reason not just with reward numbers, but with the entire context of human knowledge, all expressed in natural language.”

Trust and Transparency in Autonomous Systems

30:28 to 30:59

Understand the importance of explanation and transparency in future autonomous AI systems.

“So as you go about your week, think about this.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Welcome to The Deep Dive, the show where we take a stack of dense research, unpack the most revolutionary ideas, and distill them down into the core insights you need to be perfectly informed. Today, we are tackling one of the most persistent and frankly frustrating bottlenecks in the world of reinforcement learning, or RL. Think about how a typical RL agent operates. It's basically a learning machine that's fueled exclusively by numbers. It tries an action. It gets a scalar reward, just a single number telling it if that was good or bad. But when you or I learn a new skill, say learning a new game, or even just putting together furniture, we don't just use numbers.

0:35We read the manual. We follow instructions. We use common sense. We use all this rich linguistic context. So the challenge we're really confronting today is a huge one. How do you redesign the entire RL system to actually use that human knowledge? How do you let an AI read the manual? And this is a massive hurdle in the real world. Right now, if you're trying to deploy an autonomous system, maybe a robot in a warehouse or some kind of manufacturing control system, you're starting with just mountains of domain knowledge. You've got safety manuals, engineering specs, expert strategies, all of it written in plain English.

1:10And standard RL just can't see any of that. Exactly. It's completely blind to it. It can't read the safety protocol. It can't use those expert tips. So it has to throw all that human wisdom away and just learn everything from scratch, which is this incredibly slow, brute force trial and error, this inefficiency and the lack of transparency. That's what's holding back a lot of these high stakes deployments. Right. So our mission for this deep dive is to explore a, well, a pretty revolutionary new framework that's designed to solve this exact problem. It unifies the numerical rewards with that linguistic context.

1:43It's called prompted policy search or Propies. And what makes Prop ES such a big deal and what our sources really hammer home is that it goes way beyond just, you know, simple augmentation. A lot of prior work used large language models, LLMs, to maybe suggest a better reward or translate what the agent is seeing. Prop ES is so much more aggressive. It puts the LLM right at the center of the whole optimization loop. So the LLM becomes the optimizer itself. It literally becomes the policy optimizer. This is a fundamental shift from just helping out to being in core control. We're going to get into how this fusion speeds up learning, how it performs across dozens of tasks, and maybe most importantly, how it gives us a level of transparency that we just never had before in RL.

2:29Okay, let's unpack that architecture then, because this sounds like it fundamentally changes the relationship between the policy and the thing that's optimizing it. If the LLM isn't just suggesting a reward and it's not making real-time decisions, what is it actually doing to steer the agent's learning? The core innovation is really a conceptual one. Propies treats the whole problem of policy optimization as an in-context reasoning problem. In a traditional RL setup, you have these specialized optimizers, right, like PPO or SAC. They use numerical data, specifically gradients, to mathematically tweak the policy's parameters, which we call theta.

3:06Right. It's all calculus, basically. It is. Prop. Yes just bypasses that entirely. It doesn't need to calculate an explicit gradient. Instead, it uses the LLM to directly propose a new, hopefully improved, set of parameters for that policy. So the whole loop changes. It's not just about the agent doing something and getting a reward. It's more like the optimizer suggests a blueprint, the agent tries it out, sees how it goes, and then the optimizer reasons about a better blueprint for next time. That's a perfect way to describe it. Let's walk through that cycle. The LLM, it gets an initial prompt, and it generates an initial set of parameters.

3:39Let's call it theta1. That policy, using theta1, runs in the environment for an episode, maybe a few episodes, and it gets a reward, R1. Okay. Simple enough so far. Here's the key step. The LLM doesn't just get that one reward number back. It gets the complete history of every single parameter suggestion it's ever made and the reward that came from it. The researchers call this history gamma. So it sees this whole ledger of theta-i post pairs. And based on that entire history, it's prompted to reason about it and then suggest an improved set of parameters, theta i plus one. Wow. So that history, that gamma, is like the agent's entire long-term memory of every experiment it's ever run.

4:18That lets the LLM search globally, not just the LLM. But let's be crystal clear on one thing, because this comes up a lot with LLM research. Is the LLM involved in real-time control? If the agent is balancing a poll, is the LLM getting an ATI call every tenth of a second to decide push left or push right? No, and that is an absolutely critical distinction. The LLM is strictly the high-level optimizer. Once it generates the parameters, the policy just runs on its own in the environment, using a simple function to map states to actions. The LLM is only called once per optimization step, say every 20 episodes or so in the experiments, to update those parameters based on all the history it's seen.

4:55So it's the architect suggesting a new blueprint. Exactly. It's not the foreman on the ground executing every single move. That really makes the comparison to traditional optimizers so much starker. PPO, SAC, they're just pure number crunchers. They're great at following the steepest numerical paths to an optimum. They are. They are all about finding that local gradient. Prop ES, on the other hand, is designed to integrate two really powerful kinds of reasoning at the same time. First is numerical reasoning, you know, looking at the trends and that historical reward data in gamma. And second is linguistic reasoning, understanding the high-level goals from the prompt, the constraints, any expert hints you've given it.

5:34So if traditional RL is like trying to find the bottom of a valley while you're blindfolded, just feeling for the steepest slope. Right. You only know the immediate local slope, the gradient. Then prompt ES is like having that same feedback, but also a high-res map of the whole valley. That's the language. And a detailed logbook of every path you've already tried and how far down it got you. Yes. That combined intelligence allows for a much more strategic, a much more global search. And this is where that human alignment piece, which is so important for actually using these things in the real world, comes in.

6:07Interpretability. Because the LLM works with language, it can actually justify its own mathematical decisions. This is probably the most significant outcome of the whole Prop ES architecture. For every single set of parameters it proposes, the LLM also provides a natural language explanation, a textual justification for why it chose those specific values. The researchers have this great name for it. They call them textual gradients. I love that phrase, textual gradient. It completely transforms policy learning from this opaque black box thing into a process you can actually audit. But let me push back a little here.

6:43Does the research show that these justifications are actually accurate? Or are they just, you know, plausible sounding text that the LLM is good at generating? That's a great question. And the research really suggests they're highly reflective of the process. Because the LLM is the one doing the optimization reasoning itself by looking at gamma. And the prompt, the text it generates, is a direct reflection of the internal logic that led to that parameter choice. So instead of an auditor just staring at a matrix of numbers from a gradient calculation. Right. Instead of that, they can read a sentence that says something like, I saw that previous attempts with parameter X above 0.8 led to the pole being unstable.

7:19So I'm reducing X to 0.55 and compensating by slightly increasing Y, which the PROM says controls dampening. That level of detail is a game changer. I mean, think about a high stakes robot or a healthcare application where something goes wrong. With traditional RL, you get a crash log. It's useless. With Prop PS, you get a step-by-step narrative explaining exactly how and why the policy got to that faulty state. Precisely. It gives the human in the loop the context they need to step in to correct a bad assumption or to just debug a bad outcome. You're not just getting an optimal policy. You're getting a documented reason and transparent path to that policy.

7:57And that's essential for any kind of regulated or safety-critical field. Okay, so the PropES framework is split into two main versions, which lets the researcher sort of isolate and test the LLM's different skills. We've got the baseline, which is called PropES, which is purely numerical, and then PropES plus a CARO, which adds in all the rich language. Let's start with that baseline. Right. The baseline, PropES, is really designed to test one thing. Can an LLM, with zero specific training, perform global mathematical optimization just using in-context information? The prompt they use is deliberately sterile.

8:31It's task agnostic. It just tells the LLM to act like a good global optimizer, and its only goal is to maximize the output of some function, which, for RL, is just the episodic reward. And the sources show they didn't just jump straight to RL environments with this. They benchmarked it against pure math problems first, which feels like a really important step to prove the basic skill. It was crucial. they ran these rigorous tests on well-known, very difficult numerical optimization benchmarks, functions like Ackley and Rastrogen, which are designed to be tricky, with lots of local minima that can trap simple optimizers.

9:05They tested this across different dimensionalities, from simple 2D problems up to 16 dimensions, and they used top-tier models like Gemini 1.5 Pro and GPT-40. And the results against methods that have been engineered for decades specifically for this, like Atom or Neldermede, were, well, they were pretty shocking. They were startling. When you look at the mean objective values across 20 of these different optimization tasks, the LLM-based optimizers were incredibly competitive. Often they were just plain superior. Gemini 1.5 Pro got the best meaning, the lowest mean objective value in 12 out of those 20 hard math problems.

9:39Hang on. A text model beat Adam the algorithm at the heart of pretty much all modern machine learning on a majority of these benchmarks just by being told the history of its previous guesses. That's right. And this is a fundamental proof point. It shows that LLMs have these powerful implicit priors about what optimization procedures look like. It's not just generating text about math. It's actually executing complex numerical reasoning that can rival and sometimes beat specialized tools. The researchers think this comes from being exposed to just massive amounts of technical data. You know, everything from optimization tutorials and hyperparameter tuning logs to financial market data.

10:14It absorbs the procedure of optimization as a linguistic concept. This zero-shot numerical skill is the bedrock that all the later RL success is built on. Okay, so we've established a solid numerical optimizer. Now let's add the language. Let's talk about Prop ES Plus there. This is where we go beyond just raw math and let the LLM use all that rich, task-specific human knowledge. This is the really crucial integration step. We go from that sterile task-agnostic prompt to one that's semantically informed. Prop ES Plus leverages the LLM's symbolic side by embedding detailed linguistic descriptions right into the prompt.

10:52This linguistic input is, for all intents and purposes, the user manual for the environment. It includes a full description of the environment, often just adapted from the official docs, you know, how cart-poll works. And it includes detailed definitions of what each policy parameter actually represents. And I want to circle back to the key advantage here. Why is this so much more powerful than just running a normal optimizer? Because a traditional method like gradient descent is completely agnostic to meaning. To a standard optimizer, a parameter value of 0.5 is just a number. It has zero semantic meaning.

11:23But to the LLM. But to the LLM, if the Prop PS Plus prompt says, parameter P1 controls distance scaling and it's measured in kilometers, while parameter P2 controls velocity dampening measured in meters per second, that's a huge clue. A traditional optimizer might see that 0.1 works well for P1 and then blindly try 10 or 100 next. But the LLM, informed by the language, can use its world knowledge to understand that jumping to 100 kilometers is probably unphysical for this environment. It can use that linguistic context to guide its numerical search and avoid wasting time on ridiculous guesses.

11:57So the change is actually pretty elegant. You're just inserting a descriptive paragraph, the policies instruction manual, right into the prompt, mixing the pure numbers with natural language context. That's it. the prompt template has a dedicated section for defining the environment, describing the state space, the action space, the policy structure, all of it in clear human language. And this blend is what lets the LLM apply these linguistic inductive biases to the numerical search. That's the core mechanism that makes PropYes Plus so much faster. All right, let's get to the results. The team didn't just test this on a couple of toy problems.

12:29They ran a massive evaluation across 15 different gymnasium tasks. This is a really tough test suite. It covers classic control like mountain car, high-dimensional Mujoko stuff like swimmer and hopper, and even games like NIMH. And to make it a fair fight, they benchmarked Prop ES and Prop ES Plus against seven really well-established RL algorithms, ATC, PPO, TRPO, SAC, you name it. And to make sure the results were solid, every single trial was run 10 independent times with 8 ,000 episodes per run. The LLM made 400 optimization steps in each of those trials. Okay, so let's start with the baseline again.

13:05Just Prop Yes, the LLM acting as a pure numerical optimizer with only the reward history gamma to go on. How well did that zero-shot numerical approach hold up? It was remarkably robust. Prop Yes consistently found effective policies without any kind of environment model or explicit gradients. Now, it did struggle with some of the really complex high-dimensional locomotion tasks like Walker, but it was superior or at least competitive in most of the other environments. What were some of the standout successes for the numerical-only Prop ES? Prop ES actually beat all the other baselines in 7 out of the 15 environments.

13:38The inverted pendulum task was a perfect showcase of its power. The goal there is just to balance a pole. And Prop ES consistently, in every single experiment, reached the maximum possible reward, the global optimum. Flawless convergence. Flawless convergence, just from looking at the numerical history. Another big win was in Swimmer, which is a continuous control Mujoko task. You have this little linked body that has to learn a snake-like motion to move forward. In that one, Prop ES showed really substantial and sustained performance gains over methods like A to C and PPO across the entire training run.

14:11So that really solidified the idea that the LLM's built-in optimization priors are powerful enough for continuous control, even without language. Okay, so now we add the secret sauce. Okay. We turn on Prop.ES Plus and include that rich semantic description of the task, what happens to performance. The overall findings are pretty overwhelming. They completely validate the unified approach. Adding the semantic information improved performance in most of the tasks compared to the vanilla Prop.ES. It led to faster convergence and higher final rewards. All told, Prop.ES Plus achieved the top performance, the highest average reward, in 8 out of the 15 tasks.

14:49The next best was TRPO, which only won 5. The NIMM game result seems like a really striking example of this synergy. Can you break down what NIMM is and why the language helped so much? Sure. NIMM is a classic strategy game. In the version they tested, two players take turns removing one, two, or three sticks from a pile. The player who takes the last stick loses. The optimal strategy is actually quite mathematical and strategic. Vanilla Prop. Yes, just using numerical trial and error, you know, did I win or lose? It really struggled. It was one of the worst performers because brute force search is a terrible way to discover a sophisticated game-winning strategy.

15:23But when you add the rules in plain English... It completely transformed the search. When the Prop PS Plus prompt included the explicit rules, the goal, the constraints Prop PS Plus immediately became the top performer in that task. It's a crystal clear example of how linguistic context gives the LLM this crucial, high-level inductive bias a map of the rule space to guide its numerical search and find the optimal policy way faster. Okay, but this is where we really need to dig in, because it's not all good news for adding language. In most cases, reading the manual helps, but for the frozen lake environment, adding semantics actually made performance worse than the baseline prop ES.

16:03That feels totally counterintuitive. What happened? The key thing about Frozen Lake is that it's a grid world with stochastic transitions. And the agent decides to move up, there's a good chance it might slip and go left or right instead. It's the slippery environment. And the problem was, even though the prompt explicitly said the environment was stochastic or slippery, the LLM seems to have defaulted to a policy based on an inherent linguistic assumption of departmentistic dynamics. So it's own common sense. The idea that when you decide to go up, you go up. That was a stronger signal than the explicit instructions in the prompt.

16:36Exactly. It generated a policy that would have been perfectly logical if the world wasn't slippery, a policy that just tried to move straight to the goal. But in the real stochastic environment, that policy failed over and over because it wasn't accounting for that uncertainty. The LLM's powerful linguistic reasoning, instead of helping, introduced a misleading inductive bias that was actively harmful. That is a fascinating and really important lesson. It highlights the danger when an LLM's generalized common sense oversimplifies a technical specialized environment. Absolutely. The semantic context that's so powerful in rule-based or deterministic tasks becomes a liability if it misrepresents reality.

17:17Vanilla Prop Yes, the one just using numerical rewards, actually did better in Frozen Lake because it had no preconceived ideas. It just reacted to what it observed and eventually learned to navigate the slippery ice through pure trial and error. It's a crucial tradeoff. Language can speed up learning dramatically, but only if it accurately captures the domain's weird nuances, especially things like randomness and uncertainty. So moving on from just the basic environment descriptions, the researchers also tested what happens when you give the LLM explicit expert advice, what they call hints. If we give it a specific strategy, how big of a boost do we see?

17:53The results were consistently and strongly posited. Hints just drastically enhance performance across the board. They essentially give the LLM a really informed starting point for its search, or they significantly narrow down the space it needs to explore. And the benefits are twofold, right? Yep. First, you get much faster learning, especially right at the beginning of training. And second, you get higher final rewards with way less variability between runs. So the policies it finds are more robust and reliable. Take the mountain car task. The goal is to get an underpowered car up a steep hill by rocking back and forth.

18:27The hint they provided was the classic strategy. When the car's velocity is negative, the force should also be negative to push it back and build momentum and vice versa. That's the textbook human heuristic for that problem. It is. And just putting that in language immediately improved the final rewards and cut the standard deviation of the results way down. In another domain, navigation, adding hints, literally halved the number of optimization steps needed and doubled the final reward compared to vanilla prop. That's amazing. It means the LLM can translate a high-level strategic idea like build momentum into very specific low-level policy parameter changes.

19:04And they proved this with a really important ablation study. They tested how specific the hints needed to be in a task like cliff walking. They found that if you took away the directional guidance, the bit that explicitly linked the hint to specific state action pairs performance just cratered. So it's not enough to say the cliff is dangerous. You have to say something like in states near the cliff, prioritize the action that moves you away from it. The LLM needs advice that can translate into concrete parameter tweaks to be really effective. We've seen that Prop ES works really well on these medium-sized tasks, but now we have to get into the mechanics.

19:40What did the research find out about using these massive, expensive, and kind of slow LLMs as core optimizers? I'm talking about memory speed and scaling. Right. The first critical finding is all about the in-context history, that gamma we talked about. The LLM uses the history of all the parameter-reward pairs to make its next guess. So the researchers asked, how long does that memory actually need to be? Does it need to see its entire path to do an effective global search? The answer is a resounding yes. They used the mountain car task to track performance against the history size, which they called N.

20:14And the data shows this clear, almost linear improvement in the average reward as you give it more history. The key comparison was between a really myopic search, where N equals 1, and a full history. When the LLM could only see the single most recent parameter reward pair, its performance just plateaued at a pretty mediocre reward of about 100. It got stuck in a local minimum. But when it could see its whole exploration history? When it had the full, unbounded history, it could see all its past successes and failures. And that allowed it to effectively escape those local traps and reach the maximum reward of 200.

20:47It confirms that the LLM is using its huge context window as this explicit memory buffer that's essential for a global policy search. It's not just reacting, it's reasoning over its entire trajectory. Okay, what about cost? The intuitive fear is that making an API call to GPT-4 for every optimization step would make Prop ES way too slow compared to just running PPO or SAC locally. Is that what happened? Surprisingly, no. The API calls do add some latency, of course, but the overall time requirements for Prop ES were actually pretty modest compared to the standard RL baselines. The total time they measured included both the CPU time for the simulation and the API call duration.

21:27Since you're only calling the LLM once every 20 episodes or so, the overhead isn't that bad, especially if the environment simulation itself is fast. But the choice of LLM mattered a lot. Absolutely crucial. The big proprietary models, GPT-4O, Gemini, Claude, they all demonstrated that high-quality, effective policy search capability. They could follow the instructions. They could do the numerical reasoning. And when they tried smaller open-source models. They struggled a lot. The lightweight LLMs, like Quinn, just had much more limited numerical optimization skills. And worse, they often failed to follow instructions.

22:02They'd generate responses in the wrong format or, and this is a killer, they would just repeatedly suggest the exact same parameters over and over again. Which just stalls the whole optimization process. Completely stalls it. The training just never moves forward. It really highlights that for Prop ES to work, the LLM has to have this high-quality emergent capability for in-context reasoning and instruction following. And right now, that's still strongly tied to model size and training quality. Which suggests that Prop ES could be locked behind these really expensive proprietary models. Did they find any way around that to maybe democratize this for smaller models?

22:37They did. They explored fine-tuning as a potential fix. They used a technique called GRPO fine-tuning, which basically trains a smaller LLM on the procedure of policy search itself. They used a data set of optimal parameter suggestions that were generated for the mountain car task. So they took QN 2.514B Instruct and used this data to teach it what good optimization looks like. And the big question is, did it just learn Mountain Car or did it learn a more general transferable skill? It learned a general skill, which is really exciting. The fine-tuned QN model got much better on the training task, Mountain Car.

23:11But critically, it also showed performance gains on totally untrained domains like inverted pendulum and Pong. It suggests that by fine-tuning on this kind of optimization data, you're not just teaching the model-specific answers. You're actually enhancing its general ability to reason about optimization trajectories. And that's a very promising path for people who want to use smaller, more efficient models for this kind of thing. Okay, so we've got great results for policies with, say, up to 100 parameters. But the holy grail of modern AI is deep RL, where you're optimizing neural networks with thousands, sometimes millions of parameters.

Read the full transcript

23:48How does Prop PS even begin to handle that kind of dimensionality? This is the acknowledged frontier of the research right now. And LLM's context window, even though it's huge, just can't handle directly optimizing that many parameters. The input prompt, gamma, would become gigantic and too sparse for the LLM to reason about effectively. So to get around this, the team explored a really clever mathematical solution, reparameterization using random projection via QR decomposition. Okay, let's break down that mouthful. How do random projection and QR decomposition help? You can think of it like a really efficient compression technique, but for complexity.

24:23You've got your high dimensional policy parameter theta, all the thousands of weights in the neural network, and you need to compress that down into a small vector that the LLM can actually handle. So the policy, theta, is mapped down to a fixed, low-dimensional latent vector, let's call it Z. The LLM only ever sees and optimizes that small vector Z. Once the LLM suggests a better Z, that vector gets mapped back up to generate the full, high-dimensional neural network policy. And why the complex QR decomposition part? Why not just a simple projection? The QR decomposition is essential because it guarantees the mapping is orthonormal.

24:59In simple terms, that means it preserves all the important geometric information of the policy when it compresses and decompresses it. If you just used a simple projection, the search space for Z could get all distorted, and a tiny change in Z might lead to a huge catastrophic change in theta. QR decomposition keeps the search space stable and structured. It makes sure the LLM is optimizing a meaningful compressed representation, not just a jumble of numbers. And did it work? Did this let Prop ES tackle deep RL environments? It did. The initial results were incredibly encouraging. They tested this on the swimmer environment, using the exact same neural network architecture for all the algorithms to keep it fair.

25:38Prop PS, using this projection method, got a mean reward of 150. And the traditional methods. They were way behind. PPO only got 92.5, and TRPO only reached 71.8. That's a huge win just from swapping out the optimizer from PPO's gradient descent to the LLM's in-context reasoning with this smart projection technique. And they took it even further. They successfully applied Prop ES to high-dimensional robotics using something called dynamic motor primitives or DMPs. These are often used to encode complex motor skills like swinging a robot arm. They used it for a robotic table tennis task, which involved optimizing a complex policy with 70 parameters.

26:1870 parameters for a task that requires that level of coordination. Yes. And in that task, Prop ES achieved a significantly lower distance from goal than traditional methods like OpenAI ES and PPO. It's a really promising demonstration that Prop ES can scale beyond simple simulations into complex real-world motor control, and that this projection technique is a viable path toward broader deep RL application. So to kind of pull all this together, what Prop ES represents is a really critical development because it successfully integrates linguistic reasoning and numerical optimization into one single framework.

26:55It completely changes the role of the LLM, right? It goes from being a sidekick to being the central engine of optimization. And in doing that, it lets the user and all their existing domain knowledge directly shape how the policy learns and evolves. And that has some massive practical implications for you, the listener, especially if you're involved in deploying AI. The two biggest takeaways here are all about governance and ease of use, transparency and flexibility. Absolutely. Transparency is the killer app. Those textual justifications, the textual gradients, they provide an inherent audit trail.

27:27For any field that needs strict safety standards, autonomous cars, medical robotics, this level of interpretability isn't a nice-to-have. It's non-negotiable. You can now trace exactly why a policy parameter was changed, linking it back to a specific failure it saw in the history and a linguistic rule in the prompt. This is how we move toward AI systems that are actually verifiable and accountable. And the flexibility part is just as big a deal. I mean, think about traditional RL. You want a robot to follow some complex safety rule, like never exceeding a certain torque. You have to go through this painful process of trying to mathematically encode that into a brittle reward function.

28:05We call it reward shaping. With PropYes, we can potentially just bypass all that complexity. Instead of trying to code the constraints, you just describe them in plain English. You provide expert hints directly to the LLM optimizer. That dramatic drop in complexity makes this whole approach way more adaptable and should speed up deployment in new areas where we already have a lot of human knowledge. Of course, the research is revolutionary. But the authors were also very clear about the current limitations. Right. The first and biggest one is still scaling up to really complex deep RL. As we said, the random projection method is promising.

28:37But the work so far focuses on fixed policy architectures. Getting this to work on massive neural networks where the challenge is discovering optimal representations, not just tuning weights, that's still a huge open research problem. And then there's that more philosophical question about where this skill even comes from. Yes. The mystery of the origin of the optimization capabilities. We can see that LLMs are powerful optimizers, but we don't fully understand why. Did they just read every optimization tutorial on the Internet? Or is this an emergent ability of large models to reason over numerical sets?

29:11We need a deeper analysis to go from just observing this skill to being able to reliably engineer it. And finally, prompt sensitivity. You have to structure that knowledge in the right way. Exactly. While the LLM was pretty resilient to specific phrasing, you know, you could use synonyms and it was fine, it was sensitive to the order of information. If you put those long tables of numerical history too early in the prompt, it could kind of overwhelm the LLM's focus on the actual instructions. So optimizing the prompt structure itself is still an area that needs refinement to make Prop ES totally robust.

29:47This entire deep dive has really shown that the large language model can be this competent central force in policy optimization. It's moving us toward creating human-aligned RL agents that can reason not just with reward numbers, but with the entire context of human knowledge, all expressed in natural language. Indeed. And it's worth pointing out that this research, Unifying Semantics and Numerics for Efficient Transparent Policy Search, it was supported by a grant from Procter & Gamble. That really underscores the huge commercial interest in AI that is not just effective, but also, and this is crucial, transparent.

30:19Absolutely. This work is showing that autonomous systems can start to derive these textual gradients plain language explanations for really complex parameter changes. So as you go about your week, think about this. If future autonomous agents in high-stakes environments, systems controlling factories, driving platforms, or hospital robots, if they can explain their parameter changes in natural language, what impact will that have on our ability to audit them, to regulate them, or ultimately to trust them as they become a bigger part of our world. That ability to trust the explanation behind the action, that's arguably what's going to define the next generation of safe AI.

30:57A deep dive for another time. Thanks for joining us.

From the publisher

The source material details the Prompted Policy Search (ProPS) framework, a novel approach that positions a Large Language Model (LLM) as the core policy optimizer in reinforcement learning tasks. This architecture operates by having the LLM iteratively propose new policy parameters after reasoning over the **history of previous numerical reward feedback** and corresponding parameter settings. The advanced version, **ProPS+**, significantly improves performance by integrating rich semantic information, such as task descriptions and expert hints, directly into the learning process via prompts. Empirical testing across 15 standard control environments demonstrates that **ProPS+ is highly effective**, often surpassing traditional RL algorithms by capitalizing on this linguistic context. Furthermore, the research validates the concept by showing that the ProPS technique is **robust to variations in prompt phrasing** and can scale to more complex, high-dimensional neural network policies through methods like random projection. This methodology establishes a paradigm for transparent, human-aligned optimization by unifying **linguistic reasoning with standard policy search**.

More from Best AI papers explained

All 475 episodes
Prompted Policy Search: Reinforcement Learning through Linguistic and Numerical Reasoning in LLMsBest AI papers explained · 31 min
Listen in VO