Multi-Turn Reinforcement Learning from Human Preference Feedback

10 Jul 2025 · 17 min · 11 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Multi-turn reinforcement learning from human preference feedback (RLHF), arguing that training LLMs with feedback on whole conversations beats single-turn feedback because good actions can look bad in isolation. It introduces contextual Markov decision processes and a multi-turn preference optimization (MTPO) method that compares two full dialogues, not individual turns, using preference-based cue functions, mirror descent, and self-play; MTPO is shown to converge to a Nash equilibrium.

Notable examples

a negotiation “car dealer” task (maximize final sale price) and an “education dialogue” task where an LLM teacher guides a student using only preference judgments from a judge with a constitution (clarity, encourages questions, comfortable environment).

Key claims

multi-turn preference learning outperforms single-turn baselines; MTPO variants beat multi-turn RLHF in preference-only settings, and MTPO matches reward-based performance in the car task despite weaker signals.

Guests

none mentioned in the transcript.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Reinforcement Learning from Human Feedback

0:30 to 2:45

Discussing the principles of RLHF and its importance in AI training.

“multi-turn interactions, you know, like real human conversation.”

Challenges of Single-Turn Feedback

2:45 to 4:52

Examining the limitations of single-turn feedback in conversational AI.

“It could actually be part of a, well, a complete and successful strategy to maybe anchor high and increase the final sale price over several turns.”

Introducing Multi-Turn Preference Optimization (MTPO)

4:52 to 6:15

Overview of the new MTPO algorithm and its significance in reinforcement learning.

“Each turn or state and the chosen action are influenced by the context, basically, everything that came before in the chat.”

Long-Term Consequences and Preference-Based Cue Function

6:15 to 8:00

How MTPO handles long-term consequences in conversations using a novel cue function.

“And that's a significant theoretical result for these complex multi-turn interactions.”

Experimental Setup and Data Generation

8:00 to 9:31

Discussing the innovative methods used to generate data for testing the algorithms.

“More diverse examples are generally very beneficial for learning, especially in these complex spaces.”

Testing Environments: Education Dialogue and Car Dealer

9:31 to 11:01

Exploring the two main testing domains for the algorithms developed.

“A really practical way around the data bottleneck for research.”

Results and Implications of the Experiments

11:01 to 13:30

Highlighting the key findings from the experiments and their implications for AI development.

“The second environment was more traditional in a way.”

Discussion on Generalization and Limitations

13:30 to 14:01

Analyzing the potential advantages of preference models and acknowledging limitations.

“This really reinforces the idea that when the environment, the goal, is truly preference-driven and subjective, directly optimizing for those preferences yields better results.”

Insights from the Car Dealer Task

14:01 to 15:00

Learn about surprising results from MTPO in reward-based scenarios.

“Directly learning from, I prefer this conversation over that one, might be the superior path in those cases.”

Limitations and Future Directions

15:01 to 15:49

Explore the limitations and future potential of learning algorithms in AI.

“Well, the experiments used relatively smaller LLMs, T5-based models, not the absolute giants.”
Show all 11 chapters

The Evolution of AI Conversations

15:50 to 17:06

Discover how multi-turn learning could revolutionize AI interactions.

“Natural, goal-oriented, defective conversations.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:28Welcome back to the Deep Dive. these, well, brilliant new methods that are teaching LLMs to really excel in those complex multi-turn interactions, you know, like real human conversation. And learning directly from our preferences, which is key. Right. It's all about moving LLMs beyond just spitting out single responses towards truly conversational intelligence. Yeah. And to really appreciate this deep dive, I think we need to kick things off with reinforcement learning from human feedback, our LHF. Okay. It's really become the sort of go-to method for aligning these models with what humans actually want or expect.

1:03Right, the standard approach now. Exactly. Now, think about traditional reinforcement learning, RL. It learns from a numerical reward signal, right, like a score in a game. Simple enough. Get points, do more of that. Yeah, but imagine trying to assign a perfect numerical score to something as nuanced as, say, conversational quality. It's incredibly challenging. I can see that. How do you put a number on a good chat? Precisely. And that's where RLHF steps in. Instead of a hard number, it learns directly from human preferences. So like, I prefer response A over response B. Exactly that. It often uses a statistical method, something like the Bradley-Terry model, to convert those simple human choices, I prefer A over B, into an, let's say, implied reward score.

1:48Okay, the Bradley-Terry model. How does that work, roughly? Well, you can think of it like a system for assigning a skill score or maybe preference strength to different LLM responses, it's a bit like how sports rankings are figured out from who beats whom. But here it's all about which response humans consistently prefer over others. That makes sense. And what's particularly insightful about the research we're looking at today is that it really highlights a major challenge, right, with how most existing RLHF methods and even the newer direct preference learning techniques have typically been applied.

2:22Yeah, they've overwhelmingly focused on single turn scenarios. Single turn. So the LM gives one response. And gets immediate feedback on just that single action. Then the process repeats. Isolated turns. Which really begs the question, why isn't that single turn feedback enough? Especially when we're trying to build AI that can handle, you know, longer adaptive conversations. That's the core issue. a key insight here is that the quality the true quality of an LLM's action might not be immediately obvious or even seem good when you look at it in isolation can you give an example sure the paper we read mentions a seller agent in a negotiation let's say it initially asks for a price that seems way too high okay sound like a bad move on its own right in a single turn view you might give that negative feedback but if we connect this to the bigger picture the whole conversation.

3:13Ah, I see. It could actually be part of a, well, a complete and successful strategy to maybe anchor high and increase the final sale price over several turns. It's about playing the long game, not just the immediate reply. Exactly. Playing the long game. So this explains why it's actually incredibly hard for a person, a human raider, to judge an isolated turn accurately. Very difficult. Imagine, like, a chatbot asks for more details. Maybe it says, could you clarify that? because it doesn't have enough info yet. A single-turn feedback system might judge that negatively, right? Because it's not a direct answer.

3:49It's a question. Seems like stalling, maybe. Yeah. But if you consider the whole conversation, that clarifying action, that question, could be absolutely crucial for getting to a much better overall outcome down the line. Precisely. And this research really drives home that point. Multi-turn RL needs planning ahead. It can't just rely on that short-sighted myopic reward maximization that's common in the single-turn approaches. Okay, so they identified this limitation. How did they tackle it? What did the researchers develop? Well, to address this fundamental limitation, they came up with novel methods for reinforcement learning that use, and this is the key bit, preference feedback between two full multi-turn conversations.

4:32Not single turns, but whole conversations compared. Exactly. They formalize this complex multi-turn setting using a framework called a contextual Markov decision process. Okay, that sounds technical. What does it mean in practice? It's a standard way in reinforcement learning to model sequences of decisions like a conversation. Each turn or state and the chosen action are influenced by the context, basically, everything that came before in the chat. But the crucial difference here is the feedback. It's not for the individual actions within the conversation, but for the entire trajectory, the whole conversation from start to finish.

5:09You compare conversation A with conversation B. Okay, so comparing entire dialogues, how did they actually make that work? What's the core algorithm they came up with? Their main algorithm is called multi-turn preference optimization, or MTPO. MTPO, got it. It's built on some pretty powerful optimization principles like mirror descent and self-play. Mirror descent. Self-play. Can you unpack those briefly? Sure. Mirror descent, think of it like a clever way to find the bottom of a valley. You take steps, but they're sort of reflected based on the shape of the valley, ensuring you move efficiently towards the best spot, the optimal solution.

5:45Okay, efficient optimization. And self-play is where the agent learns by playing against itself, or maybe different versions of itself, like AlphaGo learning Go by playing millions of games against itself. Right, learning through competition. Exactly. And by combining these, MTPO is actually proven mathematically to converge to what's called a Nash equilibrium. A Nash equilibrium. It means it finds a strategy, a policy for conversing that's preferred over any other possible strategy, assuming the other strategies stay the same. It's a very stable, optimal point. And that's a significant theoretical result for these complex multi-turn interactions.

6:22Wow. OK. And how does it handle those long term consequences we talked about? Crucially, MTPO tackles that by using a new form of something called a preference-based cue function. A cue function. I know that from RL. It usually tells the agent the expected future reward for an action. Right. That's the standard cue function, yes. But this new preference-based cue function is different. It directly accounts for how individual actions contribute to the overall conversation's success, judged by preference. How does it do that? It goes beyond just measuring value from a specific point in the conversation.

6:53it considers how a policy, a way of talking, from a given point compares to any other policy, even ones that might not have reached that exact same point. Huh. So it's comparing potential futures more broadly. In a sense, yes. It allows it to capture that long-term impact, the strategic element that was really missing from just looking at single turns. That approach seems particularly well-suited for handling that long-term context, that long game. Did they find that this core MTPO algorithm was the final word? or were there variations, maybe things that worked even better in practice? They did explore variations.

7:26There's one called MTPO. MTPO, okay. Yes. This version uses something called a geometric mixture policy. Now, theoretically, it converges to the same Nash equilibrium, the same optimal point. Right. But it tends to perform better in their actual experiments in practice. Interesting. Why is that? Well, the researchers believe, or they conjecture, that it's because this mixture policy introduces more stochasticity, more randomness and diversity in the conversation paths it explores during learning. So more exploration helps it learn better. It seems so, yes. More diverse examples are generally very beneficial for learning, especially in these complex spaces.

8:06They also developed a separate multi-turn RLHF algorithm, one that learns an explicit reward function first. Okay. It showed a similar optimization process, but MTPO's adaptive self-play mechanism, where it learns directly from preference comparisons, seemed to be a key differentiator, especially in certain tests. This all sounds incredibly complex to actually test. How on earth did they get the huge amount of feedback needed? Comparing whole conversation sounds like it would need endless human effort. That's a very, very practical question, and the researchers came up with a really clever experimental setup to handle it.

8:42What did they do? Essentially, they mimicked the RLHF process, but they replaced the human parts with other LLMs, specifically prompted state-of-the-art LLMs. LLMs judging other LLMs. Exactly. It involved a few steps. First, creating an initial data set, then using that data to fine-tune smaller LLMs to act as the conversational agent and the environment it interacts with. Okay, so smaller models doing the talking. Right. And then, crucially, they used high-capacity LLMs, think things like Gemini Ultra or Flanti 5XL, to act as the preference oracle. The judge. The judge, yes. Evaluating the quality of the conversations generated by the smaller models, this let them generate vast amounts of preference data comparing whole conversations without needing real humans for every single interaction.

9:30That's ingenious. A really practical way around the data bottleneck for research. It's a huge leap forward for scalability in this kind of research, definitely. So using LLMs as proxies for human feedback is clever. What were the specific environments they created? Where do they put these algorithms through their paces? They focused on two main domains. The first one, and I think maybe the most compelling, is a novel multi-turn task they designed called education dialogue. Education dialogue. This environment is truly preference-based. And what that means is there are no explicit numerical rewards defined beforehand.

10:02None at all. No score to maximize. Nope. Instead, you have a teacher agent, LLM, whose goal is to guide a student environment, LLM, in learning some random topic. An LLM teaching another LLM. Yeah. And the quality of that teaching conversation is evaluated solely through preferences. Preferences generated by that powerful LLM judge we mentioned. How did the judge decide what's good teaching? It uses a predefined constitution. Basically, a set of principles defining effective learning things like, is the presentation clear? Does it encourage the student to ask questions? Does it foster a comfortable environment for the student?

10:39Fascinating. So it's judging based on pedagogical principles. Exactly. It's a crucial test case for learning purely from preferences without any predefined reward signal. And really great for the community. They've publicly released the data from this. That's fantastic. Open data really helps move the field forward. Absolutely. An LLM teaching an LLM judged by another LLM. That's quite a setup. What was the second domain they tested? The second environment was more traditional in a way. It was the car dealer domain from an existing benchmark called LMRL Gym. Car dealer. Okay. Sounds more straightforward.

11:12It is. This one is a reward-based setting. The goal for the car dealer agent is very clear. Maximize the final sale price of a car during the negotiation. So a clear numerical goal. Right. This allowed them to see if their preference-based algorithm, MTPO, could perform well even when the underlying objective was clearly numerical. Could it compete with traditional RL that gets the exact reward signal? Ah, I see. A great way to compare sort of apples and oranges preference learning versus direct reward learning. Exactly. A very useful comparison. So after all this setup and testing, what are the practical implications?

11:49What did the experiments actually show? What does this mean for the future of LLMs and how we interact with them? Well, the results were, I'd say, quite clear. And they really validated the researchers' core hypothesis. Which was? That all the multi-turn algorithms, both their MTPO variants and the multi-turn RLHF algorithm, they tested significantly outperformed the single-turn baseline. Significantly better. Yes. That's really the central finding here. When it comes to these complex ongoing conversations, getting feedback on the entire trajectory, the whole dialogue, is vastly superior to just judging individual turns in isolation.

12:26So if you've ever felt frustrated because your chatbot seemed to forget the context or kept asking you things you'd already answered, this research basically explains why that happens with single-turn training. Exactly. It shows that single-turn feedback often leads to inaccurate models of what's good or bad because the ripple effect of a single decision on the whole conversation is just too hard to capture in one go. It leads to those short-sighted local decisions, not good global strategies. Precisely. You optimize for the next best reply, not the best overall conversation. Okay, so multi-turn is better.

12:59But here's the core question, I think. How did MTPO, which learns directly from preferences, stack up against the multi-turn RLHF approach, which still involves learning that intermediate reward model, especially when the goal is subjective? Right. That's a critical comparison. And in the preference-based education dialogue environment. The teaching one. Yes. The teaching one where there's no number score, the MTPO variance actually outperformed the multi-turn RLHF approach. Outperformed it. So directly optimizing preferences was better. It was. This really reinforces the idea that when the environment, the goal, is truly preference-driven and subjective, directly optimizing for those preferences yields better results.

13:46That feels important. It's not just a small improvement. It suggests that for genuinely subjective human-like goals, like having a truly engaging teacher or a really helpful assistant, maybe trying to quantify everything into a rigid reward score isn't the best way. It might not be. Directly learning from, I prefer this conversation over that one, might be the superior path in those cases. Okay, what about the car dealer task? The one with the clear reward? And this was perhaps the most surprising result. In the reward-based car dealer environment, MTPO. Which only gets the weaker preference signal, right?

14:17Just, this negotiation was better than that one? Exactly. It only knows which final outcome was preferred, not the exact price difference. Despite that weaker signal, MTPO achieved comparable performance to the algorithm learning directly from the explicit price rewards. Wow. Comparable performance with less information. Yeah. This suggests something really interesting. That these preference models might actually generalize better or be more robust than traditional reward models, even in contexts where explicit rewards could be defined. Maybe learning the underlying why of preference is more powerful than just learning the what of a numerical score.

14:56That's a fascinating possibility. Better generalization. It is. Now, while these findings are very promising, it's definitely important we mention the current limitations the researchers themselves point out. Okay, what are those? Well, the experiments used relatively smaller LLMs, T5-based models, not the absolute giants. And as we discussed, the preference data itself was generated by other LLMs, not by real human raters in large numbers. Right, the LLM judges. So the true goal they achieved in this specific setup was alignment with the preferences of a highly capable LLM, which served as the proxy for a human.

15:30It's a proof of concept for the method. Understood. So it's a proof of concept demonstrating the potential of the approach. But it sets the stage for some potentially massive advancements down the line. Absolutely. Because this research isn't just about, you know, tweaking an algorithm slightly. It feels like it's about fundamentally changing how LLMs learn to have conversations with us. Natural, goal-oriented, defective conversations. I agree. Imagine an LLM that truly understands the long game of a discussion. One that anticipates your needs based on the flow, guides you through something complex over many exchanges.

16:05Because it's learning from the entire conversational flow, not just the last thing said. That's a powerful idea. It really is. And that brings us to the end of our deep dive into multi-turn reinforcement learning from preference human feedback. We really hope this exploration gave you some serious aha moments. And maybe a shortcut to being well informed about this cutting edge of LLM development. Yeah, understanding this research helps us appreciate that intricate dance between AI capabilities and human interaction, doesn't it? It really does. It's not just about what LLMs say, but how they learn to say it in a way that truly serves our long-term goals and, well, the natural flow of conversation.

16:43So looking ahead, as LLMs get better and better at these multi-turn interactions, at planning ahead in a dialogue, what new kinds of conversations or tools or maybe learning experiences do you, our listener, think will become possible? That's a great question. How will this reshape how we interact with AI when it can truly play the long game of dialogue, just like humans do? Something to ponder until our next deep dive.

From the publisher

This academic paper introduces Multi-turn Preference Optimization (MTPO), a novel approach to Reinforcement Learning from Human Feedback (RLHF) for Large Language Models (LLMs). Unlike existing RLHF methods that evaluate single conversational turns, MTPO focuses on multi-turn interactions, where feedback is provided for entire conversations to capture long-term goals and planning. The paper presents theoretical guarantees for MTPO's convergence to a Nash equilibrium in a multi-turn preference-based RL problem. Experimental results in a new "Education Dialogue" environment demonstrate that MTPO and its variant, MTPO-τ, outperform single-turn baselines and traditional multi-turn RLHF in aligning LLMs with human preferences, even when relying on a weaker preference signal compared to explicit rewards.


More from Best AI papers explained

All 475 episodes
Multi-Turn Reinforcement Learning from Human Preference FeedbackBest AI papers explained · 17 min
Listen in VO