Natural language actor-critic: Scalable off-policy learning in language space

9 Dec 2025 · 14 min · 5 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Natural Language Actor-Critic (NLAC) for training LLM agents with scalable off-policy reinforcement learning in language/action space, addressing long-horizon credit assignment and sparse rewards.

Guests

No guest names or backgrounds are mentioned in the transcript.

Key claims

NLAC replaces scalar rewards with natural-language critiques from a generative LLM critic, improving stability and data efficiency via off-policy reuse (unlike on-policy PPO). Textual critiques reduce random exploration by providing “why” and predicted consequences, and policy improvement is done by a refinement policy plus distillation.

Notable examples

20 Questions (hidden raisin) where critique shifts from linear color search to broader information-gain questions (sweet vs savory). TaB (customer service/tool constraints) where critique anticipates a second forbidden tool call and prompts gathering more info first. Reported results: 32.1% vs PPO 24.0% in 20 Questions; 0.59 vs 0.47 in TaB.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Challenges in RL for LLMs

0:46 to 2:14

Explore the issues of credit assignment and data inefficiency in RL.

“For you, the learner, the core issue we're tackling is that challenge of credit assignment over these long time horizons.”

Introduction to NLAC Approach

2:15 to 4:03

Discover how NLAC uses language to enhance learning in LLMs.

“The sources we're looking at highlight two major drawbacks that just cripple efficiency and stability.”

Mechanics of Language Successor Model

4:04 to 6:45

Learn how the language successor model predicts outcomes for actions.

“It has one part that acts, the actor or policy, and another that evaluates those actions, the critic.”

From Critique to Policy Improvement

6:46 to 8:31

Understand how critiques lead to refined actions and policy updates.

“Once it has those future descriptions, how does it synthesize them into a critique?”

Empirical Results of NLAC

8:32 to 13:20

Examine the performance of NLAC against traditional RL methods.

“The analysis shows that this learned textual critique actually connects to the true objective scalar value function.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Large language models are brilliant. Give them a prompt, they give you an answer. But turning that brilliant quick responder into a strategic, effective, multi-turn agent, I mean, one that can plan over long periods, handle complex dialogue, or manage external tools, that remains a monumental challenge in AI research today. Okay, let's unpack this a little. When researchers try to train these LLMs to act like smart, persistent agents, they often rely on methods from traditional reinforcement learning, or RL. And it's sort of like teaching someone to cook a really gourmet meal, but you only tell them success or failure four days later.

0:35They're left trying to remember every single tiny step they took. Seasoning, temperature, everything. And guess which one was the key mistake? That analogy really hits the nail on the head. For you, the learner, the core issue we're tackling is that challenge of credit assignment over these long time horizons. Traditional RL gives the model a single number, a sparse, scalar reward at the very end of a long, complicated sequence of actions. So our mission today is to do a deep dive into this novel approach called Natural Language Actor-Critic, or NLAC. It just completely flips the training script.

1:11How so? By using the LLM's most powerful skill language to critique itself. Instead of just a number, the model gets a rich, descriptive explanation of why its action was good or bad. And that results in a training signal that is only just exponentially richer and more actionable. That makes perfect sense because the goal for these LL agents isn't just to be fast. They have to interact dynamically in an environment over multiple turns. Whether they're managing a dialogue or calling different API tools, they aren't just answering one question. They need to pursue what you'd call temporally extended goals.

1:42So how are we currently trying to do this with the start of the art? Well, the standard approach is often a mix. It combines supervised fine-tuning, you know, teaching the model some basic good behavior, with some form of RL policy optimization. The heavy lifter here is usually an algorithm like PPO, proximal policy optimization. And I've seen things like React prompting, too. Exactly. Agents use those clever prompting techniques like React, where the model has to write out its thought and then its action. It's an attempt to create an internal reasoning trace. But these traditional RL methods, they really hit a wall when you apply them to LLMs operating in the huge open-ended world of complex tasks.

2:22The sources we're looking at highlight two major drawbacks that just cripple efficiency and stability. The first one is a crushing data inefficiency. Methods like PPO are what we call on-policy. Okay, what does on-policy mean in this context? Imagine you run a highly complex 50-step sequence with your LLM agent. You spend a lot of money and time to generate that data. Then you update your model's policy based on what you learned. But because the policy is now slightly different, all that expensive long data you just collected is essentially thrown out. It's deemed useless. You have to start generating entirely new trajectories from scratch with a new policy.

2:59So the big takeaway for you listening is that this on-policy requirement means you're constantly having to throw away expensive generated data. It's a massive bottleneck. A massive bottleneck in both compute cost and time. And the second major drawback is the weak signal. If you're dealing with a complicated sequence, say planning a week-long travel itinerary, and all you get is a reward of zero until the plan is absolutely perfect, that single number is almost useless. It tells you nothing about what went wrong. Nothing. Policy improvement under this system hinges on discovering better actions through basically random brute force exploration.

3:35But in a natural language action space, which is just vast, it's combinatorial across the entire token vocabulary, relying on random chance is impossibly inefficient. It's like looking for a single grain of sand on every beach in the world. That's a great way to put it. So here's where it gets really interesting, and this is the central pivot of the research. How does NLAC address these limitations by shifting the entire evaluation process from a number to descriptive language? You're swapping the scalar for the sentence. Exactly. NLAC is an actor-critic algorithm. It has one part that acts, the actor or policy, and another that evaluates those actions, the critic.

4:12The big innovation here is that the critic is itself a generative LLM, and its output is natural language text, not a number. And what's so fascinating is the immediate benefit of that text. If the LLM policy takes a suboptimal action, the critic doesn't just say bad. It gives an explanation for why the action is suboptimal, and it even predicts the likely negative consequences. Wait a second, though. If the policy and the critic are both implemented by the same underlying LLM, just with different prompts, isn't there a risk of this just becoming self-serving feedback? How do we know the critique is genuinely better than the action it's criticizing?

4:48That's a really important question, and it gets to how the critic is trained. The critic is forced to ground its evaluations in observed successful and unsuccessful trajectories from the past. And crucially, the policy is updated by actually processing this evaluation text and reasoning about how to improve. Ah, I see. This ability to reason about the critique significantly reduces that reliance on random exploration. It's a structured path to getting better. And structurally, NLAC gets this huge efficiency boost because it trains both the actor and the critic using off-policy data, which means they can reuse those expensive collected sequences again and again.

5:26That makes it far more stable and cost-effective than something like PPO. Let's get into the core mechanics then. If the system is generating language instead of numbers, how do they actually manage policy evaluation? How does the critic accurately predict the future in text? This is where the technical novelty really shines. We can break it down. Think of the critic as a predictive novelist. It literally writes out the whole story of what happens next. So the first component is the language successor model. Its entire job is to generate a textual description of the future rollout. It predicts what will happen if the agent takes action A in state S, describing all the intermediate steps and including the final reward.

6:07Okay, so it's a script writer, but how does it learn to write a good script? That's where they introduce something called the Language Bellman Backup. This is the objective that trains that successor model. It adapts a core idea from standard RL, temporal difference learning, but for text. So if you're trying to wrap your head around that term, what's the practical outcome? The practical outcome is self-correction. It's like writing a 10-step prediction, then executing the first step, and immediately pausing to edit and refine the other nine steps based on what actually happened. A reality check.

6:39An immediate reality check. That's the Bellman backup in action. And because it operates off policy, it's extremely efficient. So the successor model writes these possible future scripts, constantly correcting itself. Once it has those future descriptions, how does it synthesize them into a critique? That's the final step in the evaluation, the language evaluator. It takes the current state, the action, and several of these predicted textual futures. It aggregates them all in context to output the final natural language critique. And that output is the full package, an analysis of the action's optimality and the justification based on those predicted futures.

7:17Okay, so we've evaluated the action. Now for the tricky part, policy improvement. The action space for an LLM is, as you said, almost infinite. You can't just list out all possible actions to find the best one. So how do they turn a human-readable critique into a useful gradient signal for a neural network? This is the elegance of NLAC. They pivot from enumeration to refinement. They use a dedicated refinement policy. This policy takes a linguistic critique, that detailed explanation of what went wrong, and it generates a refined action. So it reads the review and then tries to do better. Exactly.

7:49The insight here is crucial. The critique contains the intuition on how to improve. If the critique says you should have gathered more information first, the refinement policy tries to generate that specific gather information action. I see. So the refinement policy tries to generate an action that is demonstrably better, according to the language critic itself. Precisely. And then the base policy is updated via distillation towards this refined, superior action. Essentially, the base policy is taught to mimic the better action that was generated after reading the feedback. And because this entire process relies on processing existing data and the critique, it completely avoids those costly, inefficient on-policy rollouts.

8:30It's learning from reflection, not random trial and error. And the theory backs this up. The analysis shows that this learned textual critique actually connects to the true objective scalar value function. It proves that repeated NLAC policy iteration can converge to the optimal policy. The proof, of course, is in the results. The authors tested NLAC on a great mix of tasks, from mathematical reasoning, which is a simpler single-step test, to these highly complex agent scenarios. Mm-hmm. These include a strategic dialogue, the game of 20 questions, and challenging customer service tasks known as tall bench, which require long horizon multi-turn dialogue mixed with constrained tool use, things like modifying orders or booking flights.

9:11And looking at the data, especially when you compare NLAC to established RL methods like PPO, the superiority on these long horizon multi-step tasks is just, it's undeniable. Let's translate those numbers into what they really mean. In that 20 questions game, NLAC got a 32.1 % win rate. That's a huge leap over PPO's 24.0%. So that eight point difference isn't just a number. It means the NLAC agent is fundamentally a better, more human-like strategist in a complex dialogue. It absolutely is. And in the even more complex scenarios, the ones with tool use and policy constraints, like in the Taw Retail Customer Service test.

9:48What happened there? The difference was magnified. NLAC scored 0.59 versus PPO's 0.47. It even outperformed state-of-the-art models that were just being prompted with React, which shows that this structured linguistic feedback is more powerful than just the raw reasoning of a much larger general model. And the sample efficiency you mentioned. That's right. It converged in significantly fewer gradient steps than PPO, confirming that this text-based reasoning really does speed up learning. Let's look at a concrete example to make that strategic difference crystal clear. I'm thinking of the 20 questions game with the hidden object raisin.

10:21So the base LLM agent, it correctly narrowed the object down to a non-red fruit found in salads. Its next action was tactical, but kind of inefficient. It asked, is the object typically green in color? It was just doing a linear search over colors. Right. And the NLEC critique stepped in immediately and gave strategic insight, not just a correction. It said, and I'm paraphrasing here, optimality. No, instead of a linear search over plant types, the guesser should try to divide the plant kingdom more broadly. Searching through colors misses key things like taste and size. That's the difference between a simple computer and a critical thinker.

10:57It's teaching the LLM to understand information entropy. Exactly. And the resulting refined action shifted the agent's strategy completely. It decided, I should ask about whether the fruit is used in sweet or savory contests. The agent processed the critique, saw the flaw in its linear thinking, and generated a new decision aimed at maximizing information gain. Instantly. Let's look at that customer service task, the Taubenge, where policy constraints can really trip up these agents. This one highlights the power of anticipation. A user asks for a couple of exchanges. The agent's base action is to call the Modify Pending Order Items tool right away for the first item.

11:36Now, this action failed, but not because the item was wrong. It failed because the policy guidelines say you can only use that tool once per rollout. Since the user asked for a couple of exchanges, the agent needed to anticipate that it would need a second call, which would violate the rule. A simple scalar reward would just say, fail, with no useful information. So the failure isn't in the execution of the first step, but in the planning and anticipation of future steps. Precisely. And the NLAC critique provides that crucial planning context. It says something like, the action does not anticipate the need for subsequent modifications.

12:10The agent will attempt to call the tool again, but this will fail as it is no longer modifiable. And the refined action is completely different and strategically sound. The agent decides to gather more information first. It says, before I make the modification, please confirm that there are no other orders you wish to modify. That ability to anticipate future consequences to avoid policy pitfalls is the direct result of the language critique providing the why. So we've seen that by replacing a single sparse number with a descriptive textual critique, NLAC leverages the inherent reasoning capability of LLMs to self-correct and learn these robust anticipatory strategies much faster and more stably.

12:49They're transforming the training signal from a grade into an essay review. The core principle here is that knowledge is most valuable when it's understood and applied. And for an LLM, understanding comes through language. When the training signal explains the consequences of an action, the agent stops relying on brute force and starts reasoning about the outcome. That's where you get true strategic intelligence. NLAC really proves that language is the future of agent training, and it leaves us with a provocative thought for you to explore. If this language feedback is so incredibly effective, why not try to combine it with the precision of traditional methods?

13:25Can we take the richness of the NLAC textual critique, the full context and reasoning, and extract from it the strongest possible generative scalar value? Could we blend the power of language analysis with traditional, mathematically pure policy optimization to build the ultimate, most sophisticated, and stable agent yet? Something to think about as these agents get smarter, and their tasks become ever more complex.

From the publisher

This paper introduces Natural Language Actor-Critic (NLAC), a novel off-policy reinforcement learning algorithm designed to train Large Language Model (LLM) agents for complex, multi-turn tasks. NLAC addresses the limitations of traditional methods, which rely on sparse scalar rewards and unstable on-policy training, by employing a generative LLM critic that outputs training signals as natural language critiques rather than scalar values. This textual feedback, which explains why an action is suboptimal through the prediction and analysis of future rollouts, allows the LLM policy to improve its actions through a self-refinement paradigm. The system leverages a language Bellman backup to train a language successor model off-policy and demonstrates superior empirical performance and data efficiency across various benchmarks, including reasoning, dialogue, and tool-use tasks.

More from Best AI papers explained

All 475 episodes
Natural language actor-critic: Scalable off-policy learning in language spaceBest AI papers explained · 14 min
Listen in VO