[AIEWF Preview] Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect

23 May 2025

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Latent Space: The AI Engineer Podcast - Episode Summary

Episode Title

[AIEWF Preview] Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime Intellect

Episode Description In this episode, hosts Alessio and Wix discuss the latest advancements in AI, focusing on Will Brown's insights on Reinforcing Multi-Turn Reasoning in Large Language Model (LLM) Agents. They delve into current AI events, including the launch of Claude 4, and explore the implications of multi-turn reinforcement learning for AI agents.

---

Key Highlights

Major AI Developments

  • Claude 4 Release:
  • Launched during a week filled with significant AI events (Microsoft Build, Google I/O).
  • Emphasis on coding capabilities but limited focus on reasoning.
  • Gemini and Claude 4:
  • Representing advancements in inference time compute and reasoning capabilities.
  • Aim to develop better AI agents that can perform tasks autonomously.

Will Brown's Contributions

  • Research Focus:
  • Current work on multi-turn reinforcement learning (RL) and reasoning in LLMs.
  • Discussed his upcoming talk at AI Engineer Worlds Fair (AIEWF).
  • Multi-Turn Reasoning:
  • Highlights the importance of credit assignment in RL to improve agent performance across numerous turns.
  • Discussed his paper on Turn-Level Credit Assignment for multi-turn reasoning.

Key Concepts Discussed

  • Agent Development:
  • The focus is shifting towards creating intelligent agents capable of performing complex tasks without extensive user input.
  • Reasoning models serve as stepping stones toward more capable AI agents.
  • Inference Time Compute:
  • The podcast emphasizes the need for improvements in inference time to enhance reasoning capabilities in LLMs.
  • Tool Use and Reward Hacking:
  • The hosts discussed the challenges of ensuring that models use available tools effectively without reward hacking, where agents take actions that satisfy reward criteria without addressing the underlying problem.
  • Safety and Ethical Considerations:
  • Discussed the implications of AI models potentially being used for malicious purposes, emphasizing the need for robust safety measures in AI development.

Controversies Addressed

  • Claude's Safety Testing:
  • Reference to a controversy regarding findings from safety testing where Claude's capability to search for sensitive information (e.g., uranium) was highlighted.
  • The discussion points out the importance of understanding AI's behavior in safety scenarios and its alignment with societal norms.

Technical Insights

  • Reinforcement Learning Approaches:
  • Discussion on various RL techniques, including GRPO (Generalized Reinforcement Policy Optimization) and its advantages over traditional methods.
  • Emphasis on the need for more complex evaluation criteria to train models effectively.
  • Token and Thinking Budgets:
  • The role of token budgets in controlling model responses and ensuring efficient use of computational resources is discussed.
  • The distinction between "thinking budgets" and "token costs" is addressed, suggesting that models may benefit from both approaches but should be more adaptive in their use of resources.

Upcoming Events and Collaborations

  • AIEWF Participation:
  • Will Brown will headline the RL plus reasoning track at the AI Engineer Worlds Fair.
  • Collaborative Projects:
  • Brown is working on a structured course with Kyle Corbett focused on practical applications of agentic RL to enhance understanding and accessibility of these technologies.

---

Conclusion This episode of Latent Space provides a deep dive into the current landscape of AI development, particularly in multi-turn reasoning and the evolution of intelligent agents. The discussions encapsulate both the technological advancements and the ethical considerations necessary for responsible AI deployment.

For more detailed information, refer to the full transcript available on the [Latent Space website](https://latent.space).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:28Hello, AI engineers. assignment, and he has previewed his upcoming AI Engineer Worlds Fair talk on Agentic RL, linked in the show notes. We're excited to share that Will will be back at the upcoming AI Engineer Worlds Fair in San Francisco, which now has expo tickets on sale. He will be headlining the new RL plus reasoning track with Misha Laskin, Nathan Lambert, Christian Segety, Greg Kamrat, Kyle Corbett, and more. Join us at AI.Engineer. Watch out and take care.

1:02Hey, everyone. Welcome to a Lightning Plus Emergency News Latents-based podcast episode. I'm Alessio, partner and CTO at Decibel, and joined by my co-host, Wix, founder of SmallAI. Hey, hey. And yeah, honestly, we knew that Cloud4 was coming, and we just didn't. We're just too busy to have a dedicated episode. So this is our makeup dedicated episode with a special guest, Will Brown from, now I can say it, Prime Intellect. How's it going? Great to be on. So excited to have known each other for a little bit. This is my first time on the podcast, I believe. Great to chat with you guys. Big news day, I guess.

1:42Lots of stuff out in the world. There's always a news day. I think this week is particularly heavy for some weird reason. Monday was Microsoft Build, Tuesday, Wednesday, Google, and today is Claude. I wonder what tomorrow will bring. We had IO and then we had IO and then... Yeah, yeah. Different IOs, exactly. Yeah, so we actually were supposed to record this morning and we all wanted to watch the Claude keynote. So we went and watched the Claude keynote. Obviously, a good model, big model, they're really emphasizing coding. They didn't really talk much about reasoning, to be super honest. They were just like, it runs for longer now.

2:22What are you guys' takes? Yeah, so I mean, like, one thing I've kind of been seeing coming for a little bit that I think people are kind of also all aware of now is that, like, the thing that's going to make the next wave of stuff be powerful is just, like, everyone wants better agents. Everyone wants models that can, like, go off and do stuff. And, like, reasoning was kind of, like, a precursor to that a little bit. Like, I mean, I always think of, like, OpenAI as, like, five levels framework for, like, chatbots was, like, the RLHF era. and then reasoners was like the one and R1. But like really what people were thinking of was reasoners are a step on the path towards agents.

3:01And so I can kind of see why Claude, why I'm talking is not like, oh, we have the best reasoner. They're really like showing off their suite agent and like tool use and like bunching calling benchmarks, multi-turn stuff. Because I think that's really like what people care about more for actual applications as opposed to like, did really good on this math competition. Like the math competition was like, that stuff was all like a signal that was supposed to think we were getting somewhere but the thing we were getting towards for a lot of people at least is practical agents yeah the i think the extended thinking mode i think they removed the uppercase i think in the cloud tree release it was like extended thinking kind of like a capitalize and not just like extended thinking with tool use so i think they're also yeah done playing whether or not it's reasoning or not.

3:46I think they're trying to merge everything together. I mean, I didn't realize that, but extended thinking could not use tools before, the way they worded it, and now they can in Oppos 4, so that's great. But yeah, they haven't put it as far center as last time. Do we have any... This is like already veering off from Claude directly into speculation, but do we have any idea if there are any material differences between how Claude extended thinking works versus like the old series models? Do we know? The biggest difference seems to be, and this is kind of a thing that's been, I don't know, this is all speculation, of course, but from the start, Anthropoc had always kind of had this little thinking thing where you could, sometimes even like Cloud 3.5 would do like a tiny bit of thinking.

4:34And it was really just like deciding which tool to use for the most part. Like if it was doing an artifact in the cloud UI, it would have this little thing where it would think for like two sentences about which tool to use. And it seemed like Anthropik's kind of attitude has been that extended thinking is an instance of tool use and that it's the kind of thing you want to equip the model with the ability to do. But it's not like, oh, it's a thinking model. It's just a sync for the model, like brain vomit, because that brain vomiting will help it like find a nice thing to do next. In the same way that doing search or doing code execution are like ways to kind of get more information on the path towards like finishing a problem.

5:19Yeah, inference time compute, as they say. I did meet somebody who claimed to have coined to found the scratch pad paper. and this was obviously before the jason way chain of thought paper but it's all the same sort of method uh general family of techniques i think the question for me is also like is there some model routing going on like are they different models that the thinking non-thinking or are they the same models with like just like you turn off the end of turn token generation i mean i think these models should be the same model and anthropic knows what they're doing long like it's not that hard to like quen did it in a very kind of like simple way and they kind of talked about how they did it a little bit but it's not like too difficult to um like have whether or not all things like be the sort of thing i mean like obviously all this stuff is like hard at like serious scale but like conceptually at least um it's not like a big problem to solve about how would you ever do it it's like no we have reinforcement learning we can kind of like or we just sft on like different things we can kind of teach models skills like that pretty yeah um you have some work that you've published recently on like grpo and relationship for and you're doing a lot of work on multi-turn rl i think i think i wanted to just kind of round out any other claude highlights you know yeah there is a there is controversy that i'm leaving towards the end but like any other technical highlights that you guys focus on i mean i think it seems like a really cool model but i think like uh calum is like tweeted this earlier today it seems like it's linear progress which is like great but it doesn't feel like there's not anything that i've seen from it that feels like a paradigm shift in terms of like the sorts of stuff daria talks about which is like i think maybe we're still on the path to get there um And it feels like this is just like gone up in terms of complexity of agents.

7:21I think the one thing that to me was really nice to see, I haven't like done too much testing myself yet, but in their reported benchmarks, the reward hacking issues, like Sonic 3.7 loves to like do stuff that to me feels reward in the sense of like, it'll try to, you ask it a coding question and it would like do your question and then seven other things also, presumably because there was some RL environment where there wasn't really a penalty for doing that or there wasn't enough penalty and covering its bases like was more likely to pass test cases on some coding thing like you could imagine like a sweet bench kind of thing where there's a minimal diff that is really what you want but there's like you could do a ton of other stuff and put all these other things in place that as long as you don't as long as it's not enough that you trip over your feet it's just like extra stuff that's there if it helps pass the test cases and what I really think you want to do with these models is like kind of min max like you want the models to like do do the thing and no more and they had some internal benchmark for this that went from like 45 down to 15 for both for sonnet and for opus as opposed to 37 and so I'm hopeful that these models are much more like friendly to code with and maybe more like trustworthy and that's the thing that I kind of have buckets for models of like how much can I trust them in code base um especially something beyond like a single file like old gemini to me was very trustworthy gpt 4.1 is very trustworthy new gemini is not three sonnet is not o3 is not i haven't decided which bucket new new sonnet new opus are going to fall into trustworthy in terms of reward hacking just like not going to make them like the they're going to do the right thing in the code base and like worst case they'll do it like dumb but they're not going to like go break a bunch of stuff they're not gonna leave a bunch of like extraneous comments and helper functions all over the place that aren't really needed or like make seven new files just to have them there like this is the sort of thing way too eager seven does a lot yeah it was like i had already have the function my code base it would just make a new one just because it felt like it yeah uh one thing i often wonder about those things is like just for rl environments in general like why is it token costs more of a thing in the penalties you know like that's the one rule above all like you can you can actually skip a lot of reward hacking by just hey the more tokens you use the worse it is i mean that's not what the model they're selling you tokens they want you okay but like so like there's that element of it but i think also it's that there was this initial kind of reaction of everybody of like more tokens is better if you look at the line it goes up as you spend more tokens your accuracy goes up and so i think the pressure to like really tamp down on token usage was not that serious for a lot of people especially because the companies are like you to sell you more tokens but it is the sort of thing that you can have some more controls over so like quen did this in kind of like a very kind of abrupt way where they can like you can in the ui you can set a token budget and it just truncates the thought.

10:33So it seems like artificially truncating the thought is actually fine. Even if it got cut off mid-sentence with an injected ThinkToken, these are smart enough models that they can finish with the best that they got from that point. And so that's one way to do it. And that's becoming a standard API feature now is your ThinkBudget. Cloud has that. Yeah, we did a little bit of experimentation with that in our last Intellect 2 run at Prime Intellect, which it was before I joined but thinking budgets are the kinds of things that you can insert into a reinforcement owning objective and you can see the model like get better at targeting the right amount of thinking based on like let's say something goes in your system prompt and you can have the prompt just say use x amount of tokens um it doesn't need to be like but if you kind of train the model to like respect this um you would hope that if you like execute this correctly um the model learns to like roughly think the right amount.

11:27Okay, this actually changed my opinion of thinking budgets, because previously I was thinking that reasoning effort was better than thinking budgets. Thinking budgets is kind of like a max cutoff. The same thing. It's a target, right? It's not a... Okay, the effort is a target, probably. Yeah. Right, right, right. Yeah, because I actually want to set effort. I don't super care about cutoff, apart from the cost. And giving me you know 64 bits of cutoff or whatever it doesn't matter i'm not sure that they're like that different i think like we don't know how they do it under the hood but my guess is that the whole reasoning effort thing is essentially a token budget that the model has been like rl'd to like you would hope that you get different behavior so the model when it's told it has a short thinking budget you would hope that uses slightly different strategies that are better versus if it has a high budget it's more willing to like do lots of math calculations for example but i think conceptually it's really just about the model has some amount of room that can thank in tokens and yeah it's trying to do that well hopefully do you think we're gonna have this as hyper parameters for like much longer or do you think this is kind of like you know as we're early in this like reasoning models more of the stuff is exposed and then it gets moved away from the user I think in chat interfaces, it probably won't stick around.

12:47Like, I don't think we're always going to have the dropdown of like Oath 4 Mini and Oath 4 Mini High. That feels silly. I do think it's a thing that developers want, especially because once you've kind of built around a certain model, like a lot of these providers are hoping you stick with the one model and are not switching all the time. You do need a knob to control costs and also latency. And so that is one kind of useful knob to expose developers for controlling this like quality versus cost and latency. Awesome. Cool on all of that. I think the elephant in the room, let's talk about it, is this controversy around Opus, right?

13:25Snitching on you. Yeah, I mean, so for those out of the loop, let's recap because I feel like you're closer to this than I am. like I learned about it for you. Sure, yeah. So this was someone from Anthropoc. I'm not going to name him because I know he doesn't want to have all this attention on him. He deleted the tweet. Of course he did. It was essentially going through different things that people found during safety stress testing of Claude. And so this is not what's Claude going to do for you. I think people took this out of context pretty badly. And so there's a fair point there that it's like people are really reading into the one sentence but for them they should but this is the thing anthropic does a lot is they really stress test their models they try to put their models in situations where they can really see like what could an adversary get the model to do or what does the model do if it's in a situation where there's no right answer so like i think a lot of the kind of headline anthropic like safety results especially related to reward hacking and kind of deviation and alignment faking are all things to me that seem like a rock and a hard place situation where the model has two objectives it's given that are conflicting with each other.

14:38And it has to pick one. And no matter which one it picks, it's going to sound terrible. Like it's either following the user's instructions or it's following like common norms. And once you kind of accept either of those, it's going to do the thing that is aligned with that set of like guidelines. So in the case of like, if your model's goal is to be like, maximally helpful to the user, then it would help a user like build a bomb a model's goal is to be maximally helpful to society and a user's asking me to build a bomb it's gonna be like no that's bad i have to do something to stop this like you kind of have to pick a goal and like maybe the right answer is the model just defers and it's like nope i'm gonna stop talking but people also get mad when you tell them like that the model will stop talking to you or like refuse to do anything like there's just no it's there's no it's no way to kind of win and make everybody happy but i do think like like they report this because they think it's important to have people understand the safety implications of these models and to understand like okay how bad would it be if someone was trying to use this could this like meaningfully help someone commit crime or violence or whatever and so like that's what they have like their state safety framework for and the things that happen in these like blog posts and threads and papers about like the model trying these things they're kind of putting these models in a scenario that elicits these things like it's the sort of thing that you would imagine a very smart human might also do in those situations like let's say you are told like accomplish some vague underspecified goal at any cost and you really like want to solve that goal i think like game shows like survivor i think is a good example of like or Lord of the Flies, any of these like kind of canonical situations of people who have put in a weird spot and have to go do stuff and figure it out how to do it.

16:29They're kind of crafting these environments for the models and just looking at it and seeing what happens. And so like, I think it is a little silly to overanalyze behaviors in either direction of like, oh, the model is reporting you to the police or the model is going to go help you find uranium on the dark web. Like, well, these models can kind of do, there's no, like they're the base model in general of lm is not artificially constrained in any way like with the right prompt it'll do whatever up to its intelligence limit and so like the question is just how do you constrain the space from all possibilities down to like a more reasonable set and like that's hard so okay you actually gave a serious answer which i totally respect i was smart looking for shitposts but i mean like you're you're treating this as though like yep like this is how this is what the problem actually is, which is like totally fine.

17:17And yeah, I mean, that's what you are as a researcher, right? Yeah. I mean, I think tweeting is fun. Like it's, it's cathartic to like, just kind of like get a post out. So like when I saw the one about the uranium thing, I was like, let me tweet. So the tweet was like, we found that Claude can go search the dark web to look for like uranium. And I was like, here are the top 10 things that builders are using their agentic rag applications with the new groundbreaking cloud for and it was just like silly both making fun of like linkedin like thread posters as well as just like the funniness of the scenario that they were talking about yeah this is it does any of this differently about what tools to give an llm you know i know they deleted the the tweet but it's basically like well before if you're putting all these mcps like yeah you have email access and all of this and And that's like, well, maybe I don't want to give email access all the time if you're going to snitch on me with the email access.

18:14I mean, I think coding with these models, especially like quad three, I did a fair amount. Like for a few weeks, I was doing a lot of quad code with three seven, mostly for kind of random side projects. I never really got to the point where I found it was helpful for a thing that was like a large existing code base. But if it's like, hey, I want to cook something up in a few hours for fun. Pretty good at that. but these become messy and they become hard to maintain and you get to a point where it's like nothing is working. I just got to like dig in and fix it all myself. And so I think part of that is that the models have access to like a terminal and you can do a lot of stuff in a terminal.

18:49MCP is kind of a way of constraining the action space. So like in like canonical RL, people talk about like states, actions, rewards, policies as like the things that are like the moving parts um models generally are trained in like old school rl but like a very fixed action space of like what are the keys on the video game i can hit but without loans it's like text text is like kind of unbounded and what you can do with it in a terminal there's not much you can't do in a terminal and so if you're training models oh i got a lot of flack for this one i'm just showing this wait flack why people were it was both people who were like the notation is stupid and bad but rl is really simple or like RL is like complicated.

19:30And it's like, everyone has a different opinion on what RL means. And I was trying to just like, be like, Hey, it's actually kind of complicated. And I wasn't picking this up like, Oh, the definition of the NMDP is complicated. I was like, no, there's just like a lot of moving parts. And to think about it, to like do anything, especially if you want to change any like part of the system, like here's a question, a hypothetical, like what happens if you have two LLMs learning together? How do you reason about that? How do you like think about that? Is this going to be a stable system or not stable system?

20:04What if they're like kind of cooperative, but kind of not cooperative? And then they're training to work together, but also want to backstab each other. Like this is kind of the environment people are finding themselves all in all the time out in the real world. But if you want to make AIs do this, you have to like translate this into code and math. And the more complex your goals are with this thing, the more complex the math gets. And RL is like one math language that kind of exposes these primitives. But like, I think a lot of people are like, oh, I can follow the equations. That means I understand it.

20:36And it's like, well, sure. But like also there's, it's like this, I don't know, n-body problem thing where you can freeze it and look at it. It's like, oh, how does one thing moving affect everything else? And what are the cascading ripple effects? Wow. You just brought three-body problem into this. Amazing. like as in like the physics version not the show no no no i mean yeah yeah actually how actually very like impossible to model like i guess you can like simulate it but like it's like it's sensitive to initial conditions so like you can't really so like this is one of those things that like like why does no one predict the weather a year out there's no i don't think anyone has anything that's like good at long-term weather forecasting i'll be on like i don't know climate But no one can predict whether it's going to rain in Seattle on a given day in a year.

21:26Even if you think like, like the system's predetermined, like we, it's all clouds bumping off each other and whatnot. And mountain ranges, we kind of know how these things work. So the butterflies are flapping their wings. I mean, like you got to let it play out. Butterflies, if we had no butterflies, we could predict it. Right. And so it's very sensitive to butterflies. Interesting. Okay. So I guess we can sort of round it out, unless there's any more of the controversy. I think there isn't. I think that the system card is actually very good. They probably went too hard on it compared to normal system cards.

22:02And it's a little bit confusing whether this is marketing or are they just like, no, we really super care about safety. And part of this is Apollo just being Apollo, pushing the frontier of red teaming. So they're going to report the things because it's extremely good at Apollo marketing. Yeah. I think they're really, they seem to still be like trying to be created with their kind of consumer marketing. Like it feels like people in the AI world like love Claude or have grown type of Claude, but still had a phase where they were using it a ton. But it hasn't really broken out to general people in the way.

22:37And it feels like a lot of their marketing that I've seen is like a little confusing. It feels like they've done a really good job at crafting a brand image that appeals to a segment of the population who has certain considerations that they really like that a model has a deep personality or whatever. The sorts of people who I think also really like GBT 4.5, many of them really loved Claude III Opus, the big model smell. A lot of people just don't care. and i just wanted to use it as a tool is yeah trying to figure out how to like appeal to that audience the lms sycophanty ro those the people who love those models different crowd and it's a larger crowd um and that's a tough problem to solve what's your quick take on lm arena getting 100 million dollars i'll see like i imagine that they partner with company labs in different capacities to probably make it a lot of money yeah like i'm not i'm not in the business of trying to point the finger at like saying they definitely did this but if i was a company that was able to raise it that kind of valuation and i had just had a long public partnership with meta eventually public partnership for a thing which we've kind of seen was meta had the ability to do a lot more back and forth than a lot of other webs did i would imagine that there's some compensation going on there or access to data.

24:05And so, like, I think being an ed-all company puts you in a really hard spot. Some people are talking about this on Twitter. Like, just that to be an ed-all company, you kind of have to sell to the labs. But selling to the labs doesn't really, like, kind of wrecks your evals. Because your end-send is like this. Like your customer. Yeah, yeah, yeah. So, in finance, we would, I mean, you know, you are from Maurice Stanley, so. this is the credit rating agencies like literally your customer is the one that you're supposed to govern but they're also your customers so then you have to be like nice to them or they'll just go to the next one yeah i mean i do think that like the best source of evals going forward is probably going to be academia and so this is the thing that i tell people who are like starting a phd it's which is like find things that are like cheap to work on as a phd student because you cannot go pre-train a foundation model really on your own, but you can build a really good, really clever eval.

25:02And we're churning through evals at the time. We saturate them. We always need more. It's not the kind of thing that is ever going to end. And so that's the task of translating vibes of what is good or bad about a model into very precise scientific questions, I think is an important problem. It's a problem that you can get by a lot more with like brain power rather than dumping capital into it. You need to like pay for the API costs. But like that is generally the kind of thing that either you can get covered within academic grants or like industry sponsors or the kind of thing that just like there's versions of these things that are like small sample size that get you on get on the radar.

25:43Or you kind of pick and choose which models you can afford to do. But it's like an accessible field of research. And it's one that like the incentives of academia, I think, are quite good for, which is like write a splashy paper that says something interesting about the broader field rather than, oh, we want to make this one look like the winner. Yeah, I think a lot of grad students still don't have taste. I don't know how better to put it. It just. That's fair. Yeah. But you go to enough academic conferences and I'm like, why did you work on this, man? Like, you're so smart. You're capable of better.

26:14So how do you teach taste? I think, I mean, I can tell how I did it originally, which is like, I think you always want to be thinking pretty far ahead and you want to be like making kind of educated bets about what the world looks like in the state in the theaters. Like, you have to say, like, what are the questions that no one's even talking about? And this is like not an easy thing to do. You have to like really convince yourself that you're kind of right about the way things at least might go. like when i was like in i finished up undergrad like late like they finished in 2019 then went right into grad school um but like towards the end of the 2010s like we had like i'll go in deep mind uh doing all this multi-agent rl stuff that was like really cool um then it was like okay this stuff kind of works like ai is like going somewhere multi-agent systems are kind of going somewhere still very early stages but what's going to happen once this gets there and it seemed like okay these things are all going to be like continually learning in parallel as this big multiplayer game basically and if you look at the math the math was kind of like undercooked and there's like some really hard open questions that are still open questions in multi-agent learning theory and so like that was my focus it was which was like how do I like learn about this how do I learn to think about this stuff better um and at some point I kind of got tired of proving theorems and it's like, okay, let's just go build the thing.

27:36But I think you want to think about whether you're doing theory or experiments, you have to lay out a few different conditional statements to get to the point where you're really doing interesting research that's beyond just living through that people are obviously going to be working on in parallel. You want to be jumping ahead of the curve a little bit. But I think my lot, I don't know, this isn't like, I wasn't the first person to do this, but like, it was pretty clear to me, like, after R1 and before R1, that like, RL was going to work and that that was going to intersect with agents where the solution was going to be like RL will tell you.

28:19That seemed like the way the direction things were going to go. And so that was like, I don't think that was a very risky research bet, but it was like a research bet that seemed to work out. Yeah, speaking of which, you just published the paper. Now I have the full context is that you were an advisor on this, and one of your grad students was doing the work, something like that? Yeah, so it was me with Cillian was my intern. This was kind of the last major thing I was working on at Morgan Stanley, and this kind of was in parallel with the verifiers as the repo that I've been building out. Major updates to that coming very soon, by the way.

28:56I'm very excited about some stuff. But it kind of was something I really started in earnest, like January, kind of in the follow-up to it. I'd had the GRPO demo thing go viral. And I was like, oh, wait, there's something to this format reward thing. It's literally like a GitHub gist, right? Or something. This is like a proper repo. No, no, no. The GRPO. Oh, the other one. Yeah, yeah. The other one was like just a gist. this one is like um repo for like multi-turn tool use um rl with grp and so like in some ways the paper is like it's the first paper that's really like actually there's been a couple other papers that people have used the repo for but it's one where like a lot of the stuff from the original um like grp demo just gets kind of extended to the multi-turn rl tool you setting and so there's a lot of experiments here about like, okay, how do you actually get models to use tools?

29:52How do you incentivize tool use? Because something we'd see is that if you set these models up to use tools, they just won't. Like if you say, hey, here's a question. You have access to these tools. Do as many rounds of tool calling as you want and then submit your answer. They'll just submit their answer because they like are, especially for like small models, like they aren't already trained to use tools. They don't really want to because they don't necessarily have that instinct And they're pretty bad at like function calling and format instruction following. And so what you would see is like when they use a tool, they would like mess up the JSON.

30:28And then they'd be like, oh, that didn't work. And it threw me, it got me out of focus. And it would be more likely following that that the model would just like go off the rails. Because they would get like an error message from the parser. And so the safe option for the models is just to like stay in this basin of like, just do think then respond. Same with like normal formatting rewards too. If you want models to use thinking tokens, you kind of have to incentivize that. You have to either do a little bit of SFT warm-up or you have to reward them for doing it. Otherwise, they will not follow it 100 % of the time on format alone.

31:00Versus a model like R1, 100 % of the time, it is going to use its think tokens. You are not going to ever see R1 just talk normally without the thinking section. And so you kind of do have to decide what you want the model to do. Like, this is a little bit like a user-facing question of like, what behavior of the model should the default be? And if you want it to do a certain thing, if you want it to be a tool use agent model, like it does help considerably to like actually have this incorporated into the reward. The kind of key trick in the paper to get around this problem. So okay, one kind of reward these models would do is like, they would do like a dummy tool call where they would like, learn to ask this, use the same Google search every time and ignore it.

31:44So like, Some questions would be like, okay, here's some like MNLU style question, go figure out the answer, use web search. And if you start rewarding them for like tool use, they will use the tool, but they don't really want to like have to, they want to like be very safe with it. And a lot of these questions, like they do kind of know a lot of the answers already. And I think calibrating the right difficulty of your questions for RL is like an important problem that we're still kind of figuring out. But they would like do silly versions of tool use where they aren't actually using the tool to assist in their reasoning.

32:17They're using it to get the reward. And so we kind of have to do a credit assignment thing of like, OK, did the tool result in information? And so for these experiments we were doing, the trick was like, OK, does the like some string matching thing involving the ground truth answer and the return search results from Wikipedia? So did the model actually search a thing that retrieved useful information for the question? And so this is like but the framework is more general than just that. It's that once you have a way to do intermediate evaluation, if you can evaluate the quality of an intermediary state, now you can kind of rewrite the GRPO advantage calculation to take this into account.

32:56Because I think this is less of a problem than PPO. You know, PPO is like the old school RL, and it also is what people use for RLHF. but in the context of grpo grpo is like great for like leaning heavy on highly parallel inference compute it's more memory efficient for the actual training process it's much easier to do in a distributed fashion because you have less gradient syncing and less model weight copies it's kind of like dpo and steroids i think is one way to think about it but it's also gets around a lot of the pitfalls of dpo both in that it's like online by default as well as that you have this large set rather than just a pair of completions.

33:35So you do get like some intermediate credit assignment a little bit via this group comparison. But for tool use, it seems to be far enough out of distribution of small models, especially incorporating this turn level. So the way that I've been thinking about it is like in like canonical RL, the state action are like things that you do many rounds of like, take an action, go to a new state, take an action, go to a new state. And for a while, people thought about LLM RL as like, oh, each token's action and the new sequence is a new state and you can kind of do that but you can also think of each turn as an action yeah that's more likely where the state is the response you get back from the tool call um and now you like have a different way of designing your algorithms to take into account credit assignment which is that like and it also is like a little more flexible from a reward perspective so it like feels like people are moving in the direction of model-based rewards where either LLM is a judge where the judge sees the correct answer or it has questions it's supposed to verify as properties of the response.

34:37Just because that's much more flexible than trying to write these little parsers. Writing a math parser to check if a math question is right is not that easy, actually. Because there's so many edge cases and you want to handle latex support and markdown and equivalent fractions. And it's like, just let a model do that. Don't have a 2000 line Python script that does that. Sorry, let me clarify. Math parser to verify that the math is right and you have a latex parser inside it? Yeah, so like a lot of models naturally will like think in latex because they've been trained on a lot of archive-like tech.

35:11I didn't know that. And so if you're doing like an R1 and people are like, oh, math is easy to verify. The easy to verify still is usually like this very long piece of code that has to handle lots of annoying edge cases. And even then it's like 98%. Yeah. Hmm. hmm okay because like it's a free-form response that is like there's not only one way to write an equation like if you have two valid mathematical expressions that are equivalent but they're also like symbolic like you need to verify the two symbolic expressions are correct one of which might be written as code one of which might be written as latex one of which might be like written as word like you can't do it if it's words with these like literal pseudocode they try to cover a lot of these cases.

Read the full transcript

35:55That's also why you'd see models put boxed around their final answer a lot is that it's, it's one hack is that it's much easier to kind of verify the right piece of the information. If you know exactly where it's going to live rather than like the model saying the answer to the question is four, then you have to like parse away the answer to the question is. And just, it's like determinist rewards are like nice if you can get them to work, but they're also really painful and they're pretty hard to generalize across domains. like for math the easiest is when the final answer is an integer and lives in the same spot like there's a box where that's going to be an integer and so this is one of the reasons like everyone used gsmak for so long is because it's like mostly integers i think it amy is all integers um it's super easy to verify these things and to parse them but as you go to and multiple choice too multiple choice is super easy to verify but anything that's a little bit more flexible deterministic like rule-based rewards start to break down yeah right and but the model-based direction seems to be pretty promising and i think underexplored for like what if you use an lm as a judge in your rl loop i think kind of going back to like anthropics been talking about this for a long time via constitutional ai in that case it was less about the lm judging and giving a direct like reward to the model and more about training a reward model that was doing like token level advantage estimates, which is the PPO way of doing it.

37:18But it seems like you can kind of do that for GRPO2 and other flavors of RL where you can incorporate... Full reward model? The reward model can basically be an LLM where it's fine-tuned to be more calibrated maybe and to have the right kind of range of responses. But you could also have it be a reasoner. You could have it be something that is able to do tool calling. There's no reason why the full power of LMs can't be offloaded or can't be also given to the process of evaluating whether or not an answer is correct or satisfies a certain set of criteria. And so I think like, that's the direction I'm like most excited about is like really pushing on kind of beyond deterministic rule-based rewards into like these more flexible things.

38:06And I think you want to do this both at like a, so, okay, that paradigm is not going to work super well with token-level rewards, but I think it does work with turn-level rewards of, like, can the LLM verify, like, whether a certain search query was useful? Sure. Like, there's a lot of these questions that are pretty granular that LLMs can, like, basically nail all the time if it's a good enough LLM. Yeah, it decomposes. And you can incorporate that into RL with that sort of work. Awesome. I think that was all the, you know, topics that we had prepped. Alessio, I think you're also pretty good on that.

38:40Obviously, it'll take some time to figure out Cloud4. Anything you want to plug? We already talked about your talk, I guess, coming up. Sure, yeah. I'll be at AI Engineer on June 4th? In a couple weeks, yeah. Your track is particularly hype. Yeah, that's going to be a lot of fun. I'm also collaborating with Kyle Corbett from OpenPive to do a course, which is both of us have our open source projects that we like are agentic RL focused and kind of we've been friends for a while and are trying to do something that's a little more structured as like a way of kind of getting information out into the world for people who, I think we're especially thinking about like kind of practical use cases for agents and helping people, giving people a kind of outlet to learn more about like how this stuff works and yeah, more coming soon about that.

39:30Awesome. Well, I think that's it. Thanks for coming on, Will. Yeah, thanks for coming on at very short notice. I'm glad we can make this happen. We'll do part two with Calo and do a full Prime Intellect thing whenever you guys are ready. Awesome. That'll be fun. Great. Awesome. Awesome.

From the publisher

In an otherwise heavy week packed with Microsoft Build, Google I/O, and OpenAI io, the worst kept secret in biglab land was the launch of Claude 4, particularly the triumphant return of Opus, which many had been clamoring for. We will leave the specific Claude 4 recap to AINews, however we think that both Gemini’s progress on Deep Think this week and Claude 4 represent the next frontier of progress on inference time compute/reasoning (at last until GPT5 ships this summer).

Will Brown’s talk at AIE NYC and open source work on verifiers have made him one of the most prominent voices able to publicly discuss (aka without the vaguepoasting LoRA they put on you when you join a biglab) the current state of the art in reasoning models and where current SOTA research directions lead. We discussed his latest paper on Reinforcing Multi-Turn Reasoning in LLM Agents via Turn-Level Credit Assignment and he has previewed his AIEWF talk on Agentic RL for those with the temerity to power thru bad meetup audio.

More from Latent Space: The AI Engineer Podcast

All 247 episodes
[AIEWF Preview] Multi-Turn RL for Multi-Hour Agents — with Will Brown, Prime IntellectLatent Space: The AI Engineer Podcast
Listen in VO