[State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAI

31 Dec 2025 · 20 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Latent Space - Episode on Post-Training with Josh McGrath

Podcast Details

  • Title: Latent Space: The AI Engineer Podcast
  • Episode Title: [State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency
  • Guest: Josh McGrath, OpenAI
  • Date Recorded: NeurIPS 2025
  • Focus: Discussion on advancements in AI models, particularly regarding post-training methodologies and developments at OpenAI.

Episode Overview Josh McGrath discusses the evolution of OpenAI's models from GPT-4.1 to GPT-5.1, emphasizing the shift in focus from optimization methods to data quality and token efficiency. The conversation touches upon the challenges faced in reinforcement learning (RL) infrastructure, the impact of models like Codex on design workflows, and the importance of personality toggles in user interaction.

Key Concepts and Discussions

Josh's Journey

  • Transitioned from pre-training data curation to post-training researcher at OpenAI, focusing on more significant behavioral changes in AI models rather than marginal gains in compute efficiency.

RL Infrastructure Challenges

  • Post-training involves more complexity due to:
  • Increased number of tasks.
  • Varied grading setups.
  • Dependence on external partners.
  • He describes the night-time debugging of runs as a constant adaptation to unfamiliar code.

Codex's Influence

  • Codex transformed workflow:
  • 40-minute design sessions now lead to 15-minute agent sprints.
  • Highlights a feeling of being "trapped" by the efficiency of Codex, needing to manage downtime effectively when waiting for agents to complete tasks.

RLHF vs. RLVR

  • Both reinforcement learning methods (RLHF and RLVR) are policy gradient techniques but differ in input data quality.
  • GRPO from DeepSeek Math is noted as a shift toward utilizing more trustworthy reward signals.

Token Efficiency Revolution

  • Importance of focusing on token efficiency over wall-clock time.
  • GPT-5.1 made significant improvements in evaluation metrics while reducing token usage.

Personality in Models

  • User preferences lean towards distinct personality types in AI models:
  • Anton: Efficient, less emotional.
  • Clippy: Friendly, helpful demeanor.
  • Custom instructions allow users to configure model personalities to their liking.

Long Context and Graph Walks

  • Future advancements in long context capabilities.
  • The potential importance of agents and graph walks surpassing raw context length.

Education System Gaps

  • The current education system isn't producing enough individuals skilled in both distributed systems and machine learning research.

Vision for 2026 and Beyond

  • There’s an evolving landscape where both pre-training and post-training are pivotal.
  • The continual shift in the bottleneck for innovation requires emotional stability to adapt.

Episode Chapters

  1. 00:00:00 - Introduction to Josh McGrath and his role at OpenAI.
  2. 00:04:37 - Discussion on the Shopping Model and its functionalities.
  3. 00:07:11 - Personality dynamics in AI models (Anton vs. Clippy).
  4. 00:08:26 - Evolution from PPO to DPO debates and data quality focus.
  5. 00:13:12 - Challenges in post-training RL infrastructure.
  6. 00:17:29 - Insights on token efficiency and its implications.
  7. 00:21:23 - Addressing the need for professionals skilled in both ML and systems work.
  8. 00:24:50 - The ongoing relevance of pre-training techniques.

Key Takeaways

  • The shift towards data quality over mere optimization techniques marks a significant evolution in AI model development.
  • Token efficiency is becoming a critical factor in enhancing model performance and user interactions.
  • The differentiation of AI personality types is an essential aspect of AI user experience.
  • Future advancements will require a balanced skill set in machine learning and distributed systems, addressing current educational shortcomings.

Conclusion Josh McGrath provides valuable insights into the state of AI at OpenAI, articulating the complexities of post-training methodologies, the significance of user experience in AI interactions, and the challenges that lie ahead in both technology and workforce development. The discussion serves as a reminder of the fluid nature of AI progress and its implications on future developments.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Journey to Post-Training Research

0:46 to 1:39

Josh discusses his transition from pre-training to post-training research at OpenAI.

“No, we still are releasing non-thinking models.”

Challenges of Reinforcement Learning

1:40 to 2:27

Exploration of the complexities and challenges involved in reinforcement learning compared to pre-training.

“It's a different kind of data and engineering discipline, too.”

Collaboration and Code Ownership

2:28 to 3:44

Josh shares insights on working with external partners and the impact of code ownership on projects.

“Honestly, I don't think I'll comment too much on how many external partners.”

The Role of Codex in Workflows

3:45 to 4:23

Josh describes how Codex has transformed his workflow and the challenges it introduces.

“But then, like, what do I do during those 15 minutes after?”

Innovations in Shopping Models

4:24 to 5:48

Discussion about the new shopping model released during Black Friday and its features.

“So I think I'm still getting used to that, honestly.”

Model Preferences and User Personality

5:49 to 7:11

Exploration of user preferences for model personalities and how customization is handled.

“I think like there's no reason that we couldn't do it in the same model eventually.”

The Anton vs. Clippy Divide

7:12 to 8:06

Josh compares two approaches in AI assistant personalities and user experiences.

“People really responding to personality?”

State of Post-Training Research

8:07 to 8:43

An overview of the current state of post-training methods and evolving discussions in the AI community.

“So it sounds like you also come down on the side of using it.”

Signal Quality in AI Training

8:44 to 10:34

Discussion on the quality of signals in reinforcement learning and its effects on optimization.

“And since then, we've moved on to RLVR and I think a lot of like agents specific RL training.”

The Importance of Data in AI

10:35 to 11:53

Insights on the significance of data quality and narrative in AI research papers.

“Any other discussions that maybe having in Europe or sort of run about this time on post-training debates?”
Show all 20 chapters

Innovations from Chinese AI Models

11:54 to 12:39

Exploration of how innovations from Chinese models impact the industry and AI developments.

“we're actually having a lot of conversations about with other folks as well.”

Long Horizon Autonomy in AI

12:40 to 14:00

Discussion on the challenges and considerations for long-term autonomy in AI systems.

“I mean, like, yeah, as you said, it came out in the deep seek math paper.”

Token Efficiency and Task Performance

14:00 to 15:02

Learn about the importance of token efficiency in AI tasks and its impact on performance.

“But if you look at a 2D plot of how many tokens it takes for us to get that, it went way down.”

Merging Routers and Thought Processes

15:02 to 17:30

Explore the concept of merging routing processes in GPT-5 for improved efficiency.

“Like at some point you do kind of need to merge them or else you're just going to get these weird bumps where sometimes the router at the top decides something and it's wrong.”

Context Utilization Challenges

17:30 to 19:08

Discuss the challenges and strategies related to utilizing context windows effectively in AI models.

“Talking about long context as well, there is some discussion about, I guess, context rot or like the utilization of the context.”

Future of Context Windows in AI

19:08 to 20:48

Consider the potential for expanding context windows and the implications for AI research.

“And it was 100 ,000 documents totaling about 8 billion tokens.”

Skills Needed for AI Development

20:48 to 22:34

Identify the essential skills needed for individuals pursuing careers in AI development.

“well, the systems matter more than the models.”

Post-Training Innovations and Team Dynamics

22:34 to 24:37

Learn about the innovative work happening in AI post-training and team collaborations.

“If we were to throw codecs at it, obviously we can't do codecs at everything.”

The Ongoing Debate: Pre-Training vs. Post-Training

24:37 to 26:00

Engage with the heated discussion about the relevance of pre-training in the current AI landscape.

“It's a really fun time on post training right now.”

Reflections on AI Progress and Future Speculations

26:00 to 27:29

Reflect on the progress of AI technologies and speculate on future trends and developments.

“It's going to be that many times and I think having some emotional stabilizing to it is probably going to be good for everyone's sanity.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:12Well, here is Josh from OpenAI. Welcome. How else do you introduce yourself? Yeah, I work on a bunch of the thinking models at OpenAI. And recently, I've been sort of focused on doing search-related stuff. But yeah, just a post-training researcher at OpenAI. Yeah, and you were on with us for GPT 4.1. We were talking with Michelle, who's on maternity leave. I didn't know that. And now we're at 5.1. It's been a whole generation. Yeah, it's been wild. And 4.1 was a non-thinking model. And then since then, we sort of switched into doing... Was that your last? Was your last? No, we still are releasing non-thinking models.

0:50But that one was the one that we did that was like API-specific non-thinking. So, you know, focus has shifted a little. Yeah. How'd you get into post-training? So, previously, before OpenAI, I was doing like pre-training data curation stuff. And I think what I was seeing from like the news and looking at papers is like, oh, it seems like a lot of... not pre-training dead but i was like oh there's gonna be so much interesting stuff in post-training and at that point i was like i really want to like make some contributions there and i mean it's not even necessarily that like pre-training was dead but it was definitely changing and like you know do i want to make compute efficiency wins of like three percent or do i want to like change the behavior by 40 and honestly just seemed more exciting to go to post-training and many late nights later.

1:39That's definitely true. It's a different kind of data and engineering discipline, too. It's very strange, like, the kind of work that you need in, especially RL, like, scaling it. Yeah, definitely. I think, like, for example, the number of moving parts in an RL run is just a lot higher. Like, in some ways... Do you need order of magnitude? I don't know if I can do order of magnitude, but if you think about, like, pre-training, you know, you're moving tokens to many machines and then you're getting, like, a basic a scaler from them and then you're back propping. Yeah. The issue with RL is like you're doing tasks and each task could have like a different grading setup and each one of those different grading setups, that's like more infrastructure.

2:23And so, you know, when I'm staying up late trying to figure out what's going on with a run, it could be in way more things than there is in a pre-training run generally. Yeah. And does it matter if you own the code of the task or is it an outsourced third party person or, you know, my sense of it and the external sense of it, obviously I don't see it up close, is that you work a lot with external partners and I'm sure also some internal stuff, but which is better? Honestly, I don't think I'll comment too much on how many external partners. There are some and there's some internal. Yeah, there's, we do like, but I think...

3:00The technical trade-off of like, well, shit, like, I don't own this code. So, well, when it comes to I don't own this code, actually, like, when you know when I'm babysitting a run or something it doesn't really matter if it's like internal external whatever like do I understand the system that's going underneath and I think you end up having to like jump into a lot more code that you're like I actually don't know what this does because like I'll be watching the you know I work on my pieces of a run and then there's also you know other people working on it and like do I understand what their code is doing so that way at like 1230 in the morning when I'm like, something looks wrong and I'm like looking at this code.

3:40Can I like get context fast enough to understand if you're wrong? Do you use Codex at it? Oh, I use Codex so much. It's really changed how I work. I feel like there's a degree to which like sometimes I feel trapped by Codex because if I spend like, you know, 30, 40 minutes writing something that looks like a design doc or something, Codex can do more work than I can do in a few hours in like 15 minutes. But then, like, what do I do during those 15 minutes after? And, like, it's actually just, like, really changed how the flow of my day goes because I have to somehow now manage these, like, 40-minute sessions with, like, 15 minutes where, like, I could do something, but it's actually not nearly as effective as, like, this new flow to the day.

4:26So I think I'm still getting used to that, honestly. Yeah, yeah. I think it should be interesting for, like, also just code-based understanding when you're encountering unfamiliar code. Absolutely. So you, briefly, before we started, talked a little bit about the shopping model, which is like the latest, hottest thing. And obviously, we're just recording this right after Black Friday, Cyber Monday. First of all, any interesting findings from basically releasing shopping in JGBT right into that period? Okay, well, I think the first thing is, I don't know why I would stay in a meeting in August or so.

4:58Like, oh, hey, Black Friday is coming up. Maybe we could do a release by them. In hindsight, like, wait, why would I say something like that? You're like, yes, now you own it. Yeah. Exactly. I guess the most interesting thing to me is the new interruptibility and like the sort of qualitative experience of using it. And the same thing happens with Codex, right? Like you write a prompt and you can like press escape and say like, oh, I messed something up. And we actually did the same thing in the shopping model. So it shows you its chain of thought with like what products it's looking at. and you can write it new messages saying like, oh, you know, I actually wanted...

5:34Didn't use this. Yeah, like I wanted USB-C on this or whatever it is. And like, I think that's a really new, interesting like interaction paradigm that we have in a couple of our different services. And I'm excited to see how people use it and if they enjoy it. Yeah. Why did it have to be its own model and not just like a new tool? Stay tuned. I think like there's no reason that we couldn't do it in the same model eventually. But I think, you know, if we want to try out new things, sometimes it makes sense to make a new model. And I think it just made sense to this time say, can we do a deep research style model, but like for shopping where it's going to look really hard all across the internet for different things.

6:13You know, I think if you look at like deep research, the original one and GPT-5 thinking on like high reasoning today, I think you'll see that like eventually the models all sort of converge in their capabilities. Yeah. Would you say that this is a discussion that also a little spicy that I've kicked off in the community, there's still maybe 30 % of the community is still using deep research. A lot of them have moved over to just using five thinking as deep research. Is that the spiritual successor? Are they direct replacements? Are there things that we lose in the original deep research model if we do that?

6:45I mean, I think if you look at our published evals, they look like basically on par if not better. So like, I mean, that's personally what I do. I use like thinking on high versus using the deep research model. But I think as we've learned over the past few months, sometimes people prefer the quirks of one model over another. And so people like the deep research model, more power to them. People like 4.0? Anything special in the 4.0 post-trading? People really responding to personality? Is that a differentiator that people really care about? It's a part of your job to care about personality? Yeah, definitely people care quite a bit about personality.

7:26I think, like, over the past few months, we've been working a lot on giving users more choice over what personality they want. Right, which is the toggles. Yeah, yeah. So, no, we have those toggles. What's your favorite toggle? Honestly, custom instruction for, like, I want, I personally want my model to, like, be a tool. And so, like, I don't necessarily, like, want the warmth or anything. I just want some answers because I'm, you know, mostly using it at work. Yeah, so I call this the Anton versus Clippy divide. So Anton is the Silicon Valley HBO machine. It's the owner that does work. and doesn't try to be helpful or friendly or anything.

8:00It tries to be helpful, but doesn't try to be cheery. Whereas Clippy tries to be cheery. And I'm like, well, stop smiling at me. I'm having problems. So it sounds like you also come down on the side of using it. Anton, yeah. I think a lot of developers want Anton. George is like, it just quietly does its work. And when it's done, it shuts up. Yeah, yeah. Well, I think we're doing a lot of work to provide both people, Antons and Clippys, and I hope they all like it. Yeah. So just generally, I was thinking about like, well, what can we update people on post-training? You know, what do we know today in Neurus 2025 that we didn't know in Neurus 2024?

8:38I would say like a lot of people at the time, there's still like this whole PPO versus DPO discussion. That was there. That was a whole era. Yeah. And since then, we've moved on to RLVR and I think a lot of like agents specific RL training. I guess like, am I missing any large chunks of the post-training debates that are going on? Yeah, I mean, so not necessarily debates internal, but like my read personally from like looking at different papers that are coming out. When you look at like an RLVR paper or like a RLHF paper, they read more like an optimization paper. And to me, like the sort of interesting thing that's going on is we have this like spectrum of how high quality a signal is.

9:23So like really at the end of the day, like RLHF, RLVR, they're both policy gradient methods. But what's different is just like the input data. And it's always interesting to me that we call RLHF non-verifiable because we've trained this model to be good at like predicting human feedback. So in some sense, that's like verification. but obviously it's human preference rather than truth yeah yeah but like if the if like your value of truth is like does the user like this more like there's there's something strained that i think we haven't like looked at that axis of okay well how like sort of clean is this signal how much do i trust it and like i totally agree that you know you don't necessarily trust the rlhf signal as much as like is this the solution to this polynomial but i think there's a whole spectrum of like how high quality is the signal?

10:12What's going to happen when I like do a lot of optimization against it? And that's very different than I think worrying about like the variance of different gradients, which I think is what you end up seeing in a lot of the papers that are currently coming out. Rather than being like very data centric, they're pretty optimization centric, even though I think the innovation really is where the data is coming from. Yeah. And before I want to go broad before I go deep. Yeah. Any other discussions that maybe having in Europe or sort of run about this time on post-training debates? Like, what are, you meet your peer at Anthropic and DeepMind.

10:46What are you talking about? Well, Anthropic and DeepMind, we're all saying I'm working on stuff and things, you know? I think it gets more so talking a lot more broadly with my friends there. Or we're just talking about, man, the infra is so hard to keep up. We're not necessarily talking too much about methods directly. Because on one level, it kind of doesn't matter. Yeah. And I think also, like, there's something that's very different about academic work, where, like, what really matters is how narrativizable it is. And I think that's one of the reasons you see, like, a lot of optimization papers come out, is a lot of the data work, there's a less clear narrative around it.

11:28I think the data and the scaling is actually more important than the specific. Yeah. But it doesn't have, like, necessarily the same narrative that you get out of, like, some of the papers that you see here. And so there becomes more of a, given a specific vertical, how do I understand that? And I wish there was actually more papers on it here, but I think it can sometimes be harder to wrap up into a clean story. Yeah, that's also something that we're actually having a lot of conversations about with other folks as well. What's next, right? Where do you go from here now that we have some kind of roadmap?

12:03I think what's interesting also for me is, I guess the innovations that are exposed by the Chinese models are maybe copies or like discussions of what's going on in the labs. I think obviously GRPO, you mentioned like a lot of these RL optimizations, they come out, they present themselves as optimizations. GRPO came out in the DeepSeek math paper, which when it came out, I read it and I was like, OK, this is kind of cool. it's like a little bit cheaper, but like it does seem to have a more broad impact, I think, on the industry as a whole than was initially appreciated. I just want to I don't feel like we've processed that enough.

12:41Yeah, definitely. I mean, like, yeah, as you said, it came out in the deep seek math paper. And like, it's an interesting optimization method. But it's like, the more interesting thing that they have a new reward signal that they sort of like, that we can really, really trust. Like when you know, you find the answer to a math problem, it's a lot less debatable than like oh well is this thing that the human preferred actually what we want to do yeah like you want to be right at math yeah and so i think in some ways that's underappreciated in i would say what's getting published yeah yeah let's talk about i guess long horizon yeah what do people consider in terms of like very long horizon like we're talking like 30 hours you know more than more than a day of autonomy does this is it just more of the same or is there anything like sort of qualitatively different?

13:27Okay, so first off, what I would first say is I tend to think more in terms of like actual number of tokens than time because I think... Yeah, the human in the loop can take a while. Yeah, well, and also like it gives you a different measure to optimize against, right? Like as I was saying earlier with when I use Codex, it does something that would take me much longer. You know, it would take me like four hours in, you know, 10 minutes. What we can actually push on there is token efficiency. Yeah, that is a huge, huge research area. Yeah, and so you can see from 5 to 5.1, our overall evals, we bumped some.

14:03But if you look at a 2D plot of how many tokens it takes for us to get that, it went way down. And so I think that's like a... Did you guys hear when you had that? I know it's such a great chart. Dude, I live by those charts. That was your chart? Okay. Not necessarily that, but that shape of chart. I think that's something that we think about a lot. just because it contributes so much to your experience. Like, how long does it take to do this task? Yeah. And I think the other thing is, as you're pushing that token efficiency, it changes, you know, how many tool calls can I make? And, like, how many different things can the agent do in a reasonable number of tokens that we can actually serve?

14:43Yeah. And so I personally think in terms of tokens, yeah. I think the interesting thing, or the hard-to-understand thing from the outside is having a specific router in GPT-5, But then also basically having an implicit router in terms of the thinking spending thing, that conflates a little bit, right? Like at some point you do kind of need to merge them or else you're just going to get these weird bumps where sometimes the router at the top decides something and it's wrong. And actually, if you just handed it to GPT-5, you would have figured it out. Yeah. And I think, you know, we'll figure out the correct abstractions over time.

15:18I think like there's a... Is the intention still to merge? Because that's what it was said in the paper. Yeah, I think like eventually, you know, we'll have AGI and like you're not going to have to worry too much about how hard to think directly. It'll just, you know, we'll have one tool that you always go to and it knows how long to think for and things like that. I think that the abstractions in the way that we drive these things today, it'll change. And like, you know, I think even the amount that we've changed from, you know, having a non-thinking model to you can choose between two. And like, you know, now we can sort of route and how hard do you want to think?

15:50Like, we're adding lots of knobs, and, you know, eventually it'll probably simplify. Yeah. Another super interesting knob that everyone is doing is context compaction or memory compaction. What's going on there? Nothing to share at the moment. Let me share. Clearly an important feature, clearly inspired by codex usage as well, obviously. But I think, like, from the engineer's point of view, it feels like I used to do that as part of my harness, and now it's the models doing it for me. and I don't know how to think about that like in terms of, I guess, I'm used to having more control and now I have less.

16:26Yeah, is there a specific like? There's a specific question. I'm just getting like feedback on like, well, is this a trend that like we need, where you, it's basically a permanent fact of life from here on out. Oh, I see. You know, I don't know. I worked on long context. That was why I was on last was for 4.1 where we, you know, I think 10X the effective context window for 4.1. And so there will always be some dance of like, well, if we want to push as much as what we can do, not only should we increase the length of the context window, but we should also have strategies for keeping that context window available for as long as possible.

17:02I'm guessing that both things will sort of happen just because we want to put as much power into the models as possible. Yeah. Yeah, I think we're still in a period where we should all be expecting changes in the interfaces that all of the models give to us. that way we can improve the models. Because if we lock the interface, I think what would be sad from my perspective is if we lock the interface, if we discover something new about models, we might sort of trap that improvement under an interface that needs to change. Got it. Talking about long context as well, there is some discussion about, I guess, context rot or like the utilization of the context.

17:37Even if you gave us like a million token context, probably wouldn't use all of it. What's the recommendation there? Where are things going? Are we going to have, I guess, perfect context by next year? Is that an impossible dream? I don't know. No, it's not an impossible dream. I think I'll give a shout out to some of the evals that we did for 4.1 called GraphLocs. I love GraphLocs. We covered this in the podcast. Yeah, we did. I think if you look over time, all of those evals are still climbing. And I think one of the interesting things about that is you have to do complicated transformations across the entire context window.

18:13Like that's sort of the issue with those heat map plots of the those different. Need a little heat stack. Yeah. But the problem is if you only have to sample from one point in the context window, it's like sort of easy. Whereas with those graph walks problems, you're having to do multiple transformations across the entire context window. And so I think keep watching those. I think they've been climbing. They'll continue to climb. I would say that that's definitely like a temporary issue that we are climbing on over time. Yeah. So and then like, is 10 million tokens realistic? Is 100 million? Like, is there a natural end or there's no end and we just are going as far as the eye can see?

18:52Oh, gosh, I don't know. Like, what do you think? Yeah. I feel like, okay, there are use cases that require billions. And there are use cases that require many, many billions, maybe trillions. Yeah. Out of curiosity, like, what would be billions of tokens? We just had a context engineering discussion about like a rag code base over support issues for a company. And it was 100 ,000 documents totaling about 8 billion tokens. You can't stick that in a context window for now. That's fair. I guess that, so I would still say like, I don't know, but I think I've been like really surprised. It reminds me of when I was doing like more information retrieval stuff and like BM25 and these like very simple like Ngram indexes were like just super hard to beat.

19:32I think the agents with grep are like, they feel really similar to me where it's like, it's just unreadingly effective. So then I will not use your 10 million token context window, even if you gave it. Maybe, but like, what if we're using that context window in service of like some larger goal that just has a lot of sub search calls? Which is why I'm saying like, I just don't know. And I think that's what makes it so exciting. Yeah, yeah. I would say also like the other modalities like video would eat up a lot. And then obviously the hard sciences have proteins and all that, which a lot of information just encoded in physics.

20:13So I mean, yeah, I mixed feelings about it just because I'm like, well, this will never scale, not with like full attention. And we probably just need to invest in systems anyway, which means we're good with what we have. I mean, like, get your graph walks up. But, like, I don't know if we need, like, 10, 100x, when actually maybe we need to figure out ways to 1 ,000, 1 millionx. Yeah. Right? Like, these are just different slopes. I mean, I'm glad that you're happy with the current context windows. I think my dream would be to push it and see what happens anyway. The engineer's incentive is always to say, well, the systems matter more than the models.

20:53And the researcher's incentive is to say, well, screw your systems or we'll just put the models. Oh, no, it's so differently. Yeah, I think that's one of the most like sort of beautiful things about post training and open AI is everyone. Co-design. Yeah, it's also co-designed. Like, you know, I spend a lot of time just doing our system stuff. And I also do lots of stuff like where I'm making graph walks and I'm like doing a lot more like things on the learning side. And I think it's a great culture to have a place where people just move seamlessly between the two. Yeah. What are you guys hiring for?

21:26Presumably you're hiring. What are you guys hiring for that is hard to hire? What is the skill set that is like, we really need this, can't find it. Please, everyone, go skill up on this. As my definitely personal opinion here, I think we're still having trouble, not at OpenAI, but I think as a whole, producing lots of people that want to do lots of both systems work and ML work. And I think if you're trying to push the frontier, you don't know which place is currently bottlenecking the frontier and it changes all the time. I mean, even within one project, it might change multiple times where the current bottleneck is.

22:02But I think the education system we have right now isn't really optimized for that. So like I personally, I studied math and then I was very, very lucky to have some like great mentors after school that like taught me to be a good software engineer. But it seems like if we're going to be in this place for a while, and I think we will be, we should probably be producing more students that are great at doing both distributed systems and a lot of core engineering, as well as the statistics and other things that are required to be a good machine learning researcher. If we were to throw codecs at it, obviously we can't do codecs at everything.

22:37That's why it's still, let's say, which will progress faster? Which is more solvable by LLM? That's a spicy question. you can't say they're both equally hard I don't know, maybe they are they're differently hard one is more hill climbable than the other which is it, because then we can go do it I think one thing that's slightly simpler about some of the ML research ML research is also distributed systems to be clear, but some of the things that I would say get traditionally called ML research are things that you can treat a bit more of as a black box whereas like you know the the environment to train on you know building these these different systems is actually just like a complicated engineering problem and so theoretically i would say that they're like probably roughly equal um but i think that there's some there's some amount of effort i feel like to making the the environments for yeah yeah but let's say they They require GPUs in themselves as well.

23:43Yeah, I guess they both would, but yeah, that would be my guess. But I don't have my confidence in it. So a lot of people are building this like AI scientists, right? They're automating research. You guys have your own benchmark on TaperBench. And that's the one area that, for example, at Cognition, we've just decided to not do because it's so hard. Okay, any other people on the post-training team England and Shada have done interesting work this year. They should get more attention, but they're not getting credit. Well, okay, for sure, everyone on the shopping team that I was just working with.

24:17So like Andrew Hoyal, Manuka Strada, John Hallman, all great people. Yeah, Issa Fulford, obviously the manager for it. And she was the original deep research person. There was like three of them. Yeah, yeah. And so definitely that part of the team. But I mean, everyone is so great. I think it's hard to give out a list. It's a really fun time on post training right now. It's exciting every day. Yeah, it feels like we're all enjoying our Diet Cokes together in the office late at night. Yeah. Oh, I did want to squeeze this in before we end. Nobody actually serious is saying that pre-training is dead.

24:53It's just a meme. There's a lot of work going on in pre-training. And in fact, actually, a lot of my researcher friends are saying too much money is going to post training. That's also spicy. I don't know. One of the charts I hold in memory from this year is the Grok 4 chart. I don't know if you've seen it, but it's basically saying, well, we scaled pre-training to here and about this level of compute, and now we're spending the same level of compute on post-training as well. That's very controversial, I guess, to me, because we're all used to post-training taker taking orders of magnitude less, data, compute, whatever, and obviously we're scaling that up now.

25:28Do we get to a point where they're equal? I don't know, but it's a topic for conversation i think how much do we invest in this versus more like different free trading yeah yeah yeah so first off neither neither one of those is dead i think it's really interesting to sort of be living through something that i you know all my other like historic or technological revolutions are things that i read about in in history books and like this one's live as it's happened yeah you don't know the end yet yeah and so there's this almost like fog of war where I'm like oh did people think that like we got like the steam like uh the steam engine and they would have you know the factories I don't know if you know this but like the factories they used to be like very linear because you had to drive like one motor across an entire room and it made it so when electricity got developed they just tried to do the same thing and they're like ah this isn't all that useful and it took I think like a couple of decades before they realized wait if we have electricity we can move the little like stations in what's whatever is most ergonomic and then you know manufacturing was transformed by electricity and i think like it really gives me no confidence in being like oh this thing is dead yeah our timelines are so short yeah but usually the way like good ideas get experimented and funded and propagated actually that's there's still a human timeline it's not on ai timeline yeah yeah and so i think like things will maybe be like dormant but it'll be spiky like there'll be all some you know some whoop yeah yeah and then we'll all feel different it's like we're what's the meme?

Read the full transcript

26:59It's so over, we're so back. It's going to be that many times and I think having some emotional stabilizing to it is probably going to be good for everyone's sanity. More sanity. Well, thank you so much for joining. Thanks for all the great post training this year. Thank you. Continue giving feedback. I love to hear what you think. Awesome.

27:27Thank you.

From the publisher

From pre-training data curation to shipping GPT-4o, o1, o3, and now GPT-5 thinking and the shopping model, Josh McGrath has lived through the full arc of OpenAI's post-training evolution—from the PPO vs DPO debates of 2023 to today's RLVR era, where the real innovation isn't optimization methods but data quality, signal trust, and token efficiency. We sat down with Josh at NeurIPS 2025 to dig into the state of post-training heading into 2026: why RLHF and RLVR are both just policy gradient methods (the difference is the input data, not the math), how GRPO from DeepSeek Math was underappreciated as a shift toward more trustworthy reward signals (math answers you can verify vs. human preference you can't), why token efficiency matters more than wall-clock time (GPT-5 to 5.1 bumped evals and slashed tokens), how Codex has changed his workflow so much he feels "trapped" by 40-minute design sessions followed by 15-minute agent sprints, the infrastructure chaos of scaling RL ("way more moving parts than pre-training"), why long context will keep climbing but agents + graph walks might matter more than 10M-token windows, the shopping model as a test bed for interruptability and chain-of-thought transparency, why personality toggles (Anton vs Clippy) are a real differentiator users care about, and his thesis that the education system isn't producing enough people who can do both distributed systems and ML research—the exact skill set required to push the frontier when the bottleneck moves every few weeks.

We discuss:

Josh's path: pre-training data curation → post-training researcher at OpenAI, shipping GPT-4o, o1, o3, GPT-5 thinking, and the shopping model

Why he switched from pre-training to post-training: "Do I want to make 3% compute efficiency wins, or change behavior by 40%?"

The RL infrastructure challenge: way more moving parts than pre-training (tasks, grading setups, external partners), and why babysitting runs at 12:30am means jumping into unfamiliar code constantly

How Codex has changed his workflow: 40-minute design sessions compressed into 15-minute agent sprints, and the strange "trapped" feeling of waiting for the agent to finish

The RLHF vs RLVR debate: both are policy gradient methods, the real difference is data quality and signal trust (human preference vs. verifiable correctness)

Why GRPO (from DeepSeek Math) was underappreciated: not just an optimization trick, but a shift toward reward signals you can actually trust (math answers over human vibes)

The token efficiency revolution: GPT-5 to 5.1 bumped evals and slashed tokens, and why thinking in tokens (not wall-clock time) unlocks better tool-calling and agent workflows

Personality toggles: Anton (tool, no warmth) vs Clippy (friendly, helpful), and why Josh uses custom instructions to make his model "just a tool"

The router problem: having a router at the top (GPT-5 thinking vs non-thinking) and an implicit router (thinking effort slider) creates weird bumps, and why the abstractions will eventually merge

Long context: climbing Graph Blocks evals, the dream of 10M+ token windows, and why agents + graph walks might matter more than raw context length

Why the education system isn't producing enough people who can do both distributed systems and ML research, and why that's the bottleneck for frontier labs

The 2026 vision: neither pre-training nor post-training is dead, we're in the fog of war, and the bottleneck will keep moving (so emotional stability helps)

—

Josh McGrath

OpenAI: https://openai.com

https://x.com/j_mcgraph

Chapters

00:00:00 Introduction: Josh McGrath on Post-Training at OpenAI
00:04:37 The Shopping Model: Black Friday Launch and Interruptability
00:07:11 Model Personality and the Anton vs Clippy Divide
00:08:26 Beyond PPO vs DPO: The Data Quality Spectrum in RL
00:01:40 Infrastructure Challenges: Why Post-Training RL is Harder Than Pre-Training
00:13:12 Token Efficiency: The 2D Plot That Matters Most
00:03:45 Codex Max and the Flow Problem: 40 Minutes of Planning, 15 Minutes of Waiting
00:17:29 Long Context and Graph Blocks: Climbing Toward Perfect Context
00:21:23 The ML-Systems Hybrid: What's Hard to Hire For
00:24:50 Pre-Training Isn't Dead: Living Through Technological Revolution

More from Latent Space: The AI Engineer Podcast

All 247 episodes
[State of Post-Training] From GPT-4.1 to 5.1: RLVR, Agent & Token Efficiency — Josh McGrath, OpenAILatent Space: The AI Engineer Podcast
Listen in VO