OpenAI's Yann Dubois: Why AI Progress Suddenly Feels Real

21 May 2026 · 1 h 14 min · 31 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

OpenAI’s Yann Dubois (Jan Dubois) explains why recent AI progress feels like a sudden step change, focusing on GPT 5.5, reliability for agentic models, and post-training reinforcement learning shifting from “verifiable” math/coding competitions to messy real-world tasks.

Guest backgrounds

Jan Dubois co-leads OpenAI’s Post-training Frontiers team. Before OpenAI, he was at Stanford, co-authoring Stanford Alpaca. He previously worked on NLP for under-resourced languages (e.g., Khmer, Bahasa, Thai, Vietnamese) and studied biomedical engineering in Switzerland.

Key claims

  • Reliability threshold: by around December (at OpenAI), models became trustworthy enough that users can rely on them for substantial work.
  • “Step function” feeling comes from reliability + faster internal tooling/coding acceleration + RL moving to real-world utility.
  • Reliability for agentic systems means reducing the probability of being wrong as the model runs longer.
  • GPT 5.5 improvements include efficiency (about 2x faster on many tasks) and better integration of vertical and horizontal gains.

Notable examples

  • RL rewards previously relied on ground truth (math/coding competitions); now tools are reused for real-world coding.
  • GPT 5.5 is described as strong in agent decoding, computer use, knowledge work, and early scientific research.
  • “Thinking longer” (GPT 5.5 Pro) increases correctness probability but has diminishing returns; efficiency curves shift left while Pro extends them.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Reliability Threshold in AI Tools

0:00 to 0:35

Learn about the recent advancements in AI reliability and how they enhance utility.

“You need to reach this level of reliability to really make any of these AI tools very useful.”

Unpacking GPT 5.5's Impact

1:22 to 4:14

Explore how GPT 5.5 represents a leap in AI capabilities and its implications for real-world applications.

“Please enjoy this fantastic conversation with Yann Dubois.”

Challenges and Achievements in AI Development

4:14 to 9:06

Discussion on the challenges faced during the development of GPT 5.5 and the team's collaborative efforts.

“So we're going to unpack a lot of this, particularly on the RL side.”

The Role of the Post-Training Frontiers Team

9:06 to 12:30

Insights into the functions of the Post-Training Frontiers team at OpenAI and their contributions to model improvements.

“One thing which I will say because you asked also about one of the things that we are really proud about for this model I would say two things.”

Yann's Journey to OpenAI

12:30 to 14:01

Yann Dubois shares his background and experiences leading to his role at OpenAI.

“Oh, it's a long story, but I'll try to keep it really short.”

Yann Dubois' Background and Early Career

14:01 to 15:10

Learn about Yann Dubois' academic journey and work history in AI.

“with Thai, Vietnamese, and all these different languages.”

Behind the GPT-5 Video Announcement

15:10 to 15:43

Explore the behind-the-scenes experience of the GPT-5 announcement video.

“So I was a little bit stressed that it wouldn't work.”

Advancements in Reasoning Capabilities

15:43 to 18:17

Understand how reasoning in AI has evolved and its significance.

“So we started effectively talking about reasoning.”

The Role of Test-Time Compute in AI Performance

18:17 to 21:16

Discover how the amount of compute affects AI reasoning and accuracy.

“And that's why also now current evals look much more realistic.”

Reinforcement Learning and AI Reasoning

21:16 to 23:28

Learn about the impact of reinforcement learning on AI's reasoning paths.

“It should just think for as long as it can.”
Show all 31 chapters

The Importance of Pre-Training in AI

23:28 to 27:03

Examine the role of pre-training in enhancing AI model capabilities.

“So let's talk about how the different components of modern AI systems work.”

Future Frontiers in AI Data and Learning

27:03 to 28:00

Discuss the future of data sources and their implications for AI progress.

“or the current frontier for data, multimodal data?”

The Role of Embodied AI in Understanding the World

28:00 to 29:30

Explore how embodied AI can enhance our models by interacting with the real world.

“and embodied AI, you will learn a lot about the world and you will kind of improve general intelligence and usefulness to users by learning how the world interacts with itself.”

World Models: Balancing Simulation and Reality

29:30 to 31:04

Discuss the importance and challenges of using world models in AI.

“as a quick detour, that leads us to the concept of world models.”

Understanding Mid-Training in AI Development

31:04 to 32:42

Learn about the mid-training stage and its significance in AI model development.

“All right, so going back to pre-training, mid-training, post-training, let's talk about mid-training.”

Post-Training: Enhancing AI Utility

32:42 to 34:19

Delve into the post-training phase to see how AI models become useful for users.

“Post-training, let's start at a high level by defining what that is.”

Reinforcement Learning and Its Challenges

34:19 to 35:58

Examine the stages of reinforcement learning and its difficulties in AI training.

“So this is what we call behavior cloning.”

Capabilities and Scaling in AI Models

35:58 to 37:50

Analyze how reinforcement learning impacts capabilities in AI models over time.

“And then once it's already at a pretty good level, they just do this reinforcement to go beyond what we currently have.”

The Complexity of Scaling Reinforcement Learning

37:50 to 39:09

Discuss the complications and considerations of scaling reinforcement learning.

“And like now when you look at reinforcement learning from models like Kimi or from DeepSeek models, it seems that they are closer to 1 million data points.”

Current Frontiers in Reinforcement Learning Techniques

39:09 to 42:01

Explore the latest techniques in reinforcement learning, focusing on GRPO and its effectiveness.

“was precisely that, that it was hard to make work.”

Exploring the Frontier of Reinforcement Learning

42:01 to 45:09

Learn about the current techniques and methods in reinforcement learning and why simplicity often prevails.

“What's the current frontier of reinforcement learning?”

From Craft to Science in AI Development

45:10 to 48:20

Discover how AI development transitions from experimentation to a more scientific approach through iterations.

“So still in reinforcement learning and circling back to some of the things you said at the beginning.”

Generalization in Machine Learning Models

48:21 to 51:24

Understand how models generalize from one domain to another and the importance of horizontal capabilities.

“So speaking of generalization, So there's been that clear evolution from math and coding success to now starting to cover different areas.”

The Hallucination Problem in AI

51:25 to 56:00

Examine the causes of hallucination in AI models and how reinforcement learning might mitigate this issue.

“the model that is trained on one particular data set and this is what i was alluding to before is at least my mental model is that generalization happens in terms of capability.”

The Trade-offs in Model Generalization

56:00 to 58:36

Learn about the complexities of optimizing AI models across different domains.

“But if you have good reinforcement in pipeline, that shouldn't happen too often.”

Challenges in AI Model Evaluation

58:36 to 1:02:40

Discover the difficulties involved in evaluating AI models and the importance of metrics.

“The first one is most of the people working on these models are pretty good at coding and they really care about coding because that's what they use as the each day kind of drivers.”

The Exciting Future of Continual Learning

1:02:40 to 1:08:06

Understand the potential and current limitations of continual learning in AI.

“And yeah, the tide is shifting, but not fast enough.”

Harnessing AI Models for Specific Goals

1:08:06 to 1:10:02

Explore the role of harnesses in improving AI model capabilities and their future.

“I actually don't quite know, to be completely honest with you.”

The Importance of Domain-Specific Harnesses

1:10:02 to 1:11:08

Learn why tailored solutions are crucial for improving AI reliability.

“they want to go from this like 80 % maybe reliability to maybe the like 85%.”

The Last Mile Challenge in AI Applications

1:11:08 to 1:13:19

Explore the significance of addressing the last mile to unlock AI's full potential.

“or could already feel that in every single domain.”

Closing Thoughts on AI and Startups

1:13:19 to 1:13:32

Hear optimistic perspectives on the future of AI in the startup landscape.

“But yeah, that's not what we're doing now.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00You need to reach this level of reliability to really make any of these AI tools very useful. And I think we just crossed that probably December last year, at least at OpenAI. Now we can trust these models to do a lot of the work that we are doing. The last few months have been pretty wild. We moved from like competitions to usefulness to users. And that's what we are feeling right now. I think most of the time, the bar lake is the last mile. there will always be a lot of space left for this last mile in different verticals. And I would highly encourage people to continue working on that. Hi, I'm Matt Turk.

0:35Welcome to the Matt Podcast. My guest today is Jan Dubois, who co-leads the post-training Frontiers team at OpenAI. The recent release of GBT 5.5 was yet another major milestone in AI. And Jan's team helped build it alongside OpenAI's prior top reasoning models, including O3 and GBT 5 thinking. Before OpenAI, Jan was at Stanford, where he co-authored Stanford Alpaca, the landmark project that kicked off much of the modern post-training research community. In this conversation, we go deep on what's actually new in GPT 5.5, why reinforcement learning is moving from math and coding competitions into messy real-world work, why AI progress can feel like a sudden step function, and why continual learning remains one of the big unsolved problems in AI three years after JetGPT.

1:22Please enjoy this fantastic conversation with Yann Dubois. Hey, Yann, welcome. Hi, Matt. Thanks for having me. It's been another wild last few weeks in the world of Frontier AI with the release of GPT 5.5, of Cloud Mythos Preview. So it feels like we have unlocked yet another step function in progress, particularly in cybersecurity, agenda coding. What's the best way to think about this from your perspective? Are things accelerating? What is happening? Yeah, the last few months have been pretty wild. Internally, we also really feel it. And I think anyone who's working with, anyone who's coding basically is really feeling it right now.

2:09I think that's really because of three reasons. The first one is, even though in my mind, the progress is actually pretty continuous, you need to reach this level of reliability to really make any of these AI tools very useful. And I think we just crossed that probably December last year, at least at OpenAI. That's where I thought we really crossed that threshold where now we can trust these models to do a lot of the work that we are doing. So it feels like a step function, even though I think actually in terms of capability, it's pretty continuous. So that's the first thing. The second reason is once you start having models that are really good, you accelerate yourself, especially in terms of coding, given that we all code internally.

2:56You accelerate yourself both for having these models, like train the other models, but also like build like the tooling that we need as researchers to like do our job. And all this acceleration, I think, means that we saw these last few months going faster and faster. The third thing that I think we are feeling is all of last year, we really built these reasoning models. And we really started pushing a lot on reinforcement learning. And initially, when we had like O1, O1 Preview, even O3, these models were still optimized for what we call verifiable rewards. Things where we actually have access to ground truth and it's easy to test whether you're correct or not.

3:39that is, for example, the case in like math questions or like coding competitions. And what I think we are realizing now is that we were able to take many of the tools that we built for these like verifiable reward cases and we were able to use them more generally for reinforcement on like real use cases. And I think that's like really why we're feeling that right now in like just real world coding rather than like competition. So we moved from like competitions to usefulness to users. And that's what we are feeling right now. Okay, fascinating. So we're going to unpack a lot of this, particularly on the RL side.

4:19For the first thing that you mentioned, reliability, is that engineering? Is that models? What makes a model reliable in the way you meant it? It's a little bit of everything, but in general, given that these are agentic models, the longer, if you just think about it as like every two minutes, there's like a certain probability that they're wrong. The longer that they run, the higher the probability that the final answer is going to be wrong. So it's just something inherent in agentic models. And what we've been pushing a lot on is making sure that the model, we decrease this probability of being wrong every two minutes.

4:55So purely from a model point of view, of course, there's a lot of reliability that is also happening on the applied side. And the team at OpenAI has been doing an amazing job on that. But I'm even talking only about reliability of our models and like making sure that like basically we decrease the probability of being wrong great so 5.5 formerly known as spud was uh as mentioned a big deal uh is a big deal and i'm just curious from the inside what was what are you guys the most proud of what did you find the most challenging give us some some some color on like how you all uh felt uh you know releasing this we're all really excited about 5.5 to be honest it is one of these models where everyone in the company was extremely involved uh in building um and i think that we really feel it now uh that's like we got a lot of attention because of the 5.5 and it's uh it seems like all the stars were aligned that doesn't always happen um and i was just like a great model for this.

5:57So we did feel it. It's kind of funny because in general with every model that is looking really good early on, we have a model, we all get really excited about it. And then there's like tons of doubts that start coming up because it's like, oh, like everyone is so high, is like hyping this thing internally, but actually it's like bad at all these other things. And then there's another wave where like people start under hyping it and it kind of goes through waves and it depends like when we actually ship it how like people feel about it internally but that's true of like most models that we have um so 5.5 was not that different in this case but definitely maybe had like a higher amplitude of the wave so people were very excited then very not as excited and and we shipped it and people were happy externally how long does that process take like you know you uh including the waves of going up and down and of excitement?

6:50I guess it depends on the release and the importance of its release. But is that a few weeks, a few months? It really depends. So I can't talk exactly about what went into 5.5, but it kind of depends which part of the pipeline is training parts of the model. So we really have different sub-teams, including pre-training, and you have the mid-training stage, and you have some post-training. And usually the closer you get to products, like portioning being the last one, the faster the iteration cycle is. And if you're more upstream, the slower the iteration cycle is. So it could go from, let's say, from months to days.

7:335.5 was particularly good on agent decoding, computer use, knowledge work, and early scientific research. How does that work internally? Do different people focus on those different parts? How do you get to that result? Yeah, we definitely have different teams that are working on specific use cases and are pushing on these use cases. My team specifically is actually the one that is kind of taking all these vertical improvements and try to put them together in the final model. You could see it as a team that is doing both kind of the smoothing function. So you have all these improvements, but you need to make sure that the model doesn't feel too spiky, doesn't feel differently on different verticals.

8:14And also you need to have some teams that are working, and that's basically what my team is doing, on all the horizontal improvements. So there are many things like instruction following, function calling, or like thinking about how much should a model think for on different problems. Those are very horizontal and that kind of impacts all these use cases. So we have both these more vertical teams and these more horizontal ones. And both are very important to improve on the model. and the good thing is that these things can kind of be improved orthogonally. So you might have like multiple different teams that are working on certain verticals and maybe for one model, there's only half of these teams that made integrations basically in the last run and like improved the model on these capabilities and maybe for the next model, it'll be the other half.

9:03So that's kind of at a high level how it works. One thing which I will say because you asked also about one of the things that we are really proud about for this model I would say two things. Number one is the efficiency of the model. We really, really improved the efficiency of the model. And like we, most of the tasks can be basically performed, I would say like 2x faster now with this model. So that's great. And the other one that I already mentioned before, but it's kind of this alignment of the company and making sure that like everyone is working towards the same goal. And that really takes the entire company working towards like this north star of building one good model in like specific timelines.

9:47So very, very proud of how that happened. Great. And then speaking of efficiency, how do you optimize for that? We're talking about efficiency per token. Are we also talking about latency in serving the model? What part is AI research versus engineering? So that's what I mean when I say it's the entire company, is that it really comes from everywhere. it has to come from like inference optimizations. It has to come from the model being more efficient in its thinking time. So you have basically every token that you think for. Basically the usual plot that you should be looking at is X axis, the number of tokens that you think for and Y axis, the performance.

10:27So this is these test time scaling curves that we look at. And research basically tries to move this curve to the left. So think less to be the same level or more correct. And then inference also deals with this x-axis, but switches it from number of tokens to actual latency. And the final thing that people care about is latency on x-axis, performance on y-axis. And this is where everything comes together. And this is really what happened with 5.5. So yeah, that's why I was saying I'm really proud of the company for this one. Okay, great. Let's talk about you for a minute. So you are in the Post Trading Frontiers team.

11:08So that team you described as horizontal. What does the team do in general? Yeah, I would say there's three things that we do. So in a broad sense, we are the Post Trading org. And my team is the Post Trading Frontiers one. So there are three things that my team does. Number one is we kind of decide what goes into the final run. So as we talked before, there's like money verticals. and someone needs to decide what can go in, what cannot, and also provide the science experiments for people to iterate on something that is going to be representative of the final run. So this is the first thing that my team does.

11:45The second thing that my team does is bringing everything together and actually doing the big run. So this has, as you might imagine, we train on a good amount of GPUs, so there's a lot of infra work that is needed, but also there's a lot of ML work that is needed by putting everything together and making sure things work well together. And then the third thing that my team does is horizontal improvements to the models. Basically, there are some things that these vertical teams will not usually look too much at. For example, the thinking time, as I said before. So how much should the model think for on certain answers?

12:17Or instruction following, function calling, things like memory, and general improvements to the model that are really across the stack. So that's what the Pushing Frontiers team does. and I'm leading that team. Okay, great. And what was your journey to OpenAI? Oh, it's a long story, but I'll try to keep it really short. Basically, I did my undergrad in biomedical engineering in Switzerland. I'm from Switzerland. And then I went on an exchange in Canada and I learned about what to VEC. So I don't know if you heard about this algorithm, but it basically takes words, which is like something discreet and puts it in a vector space.

13:01So puts it basically in a way to think about it as a plane where if words that are more similar to one another will be closer to one another. So it brings these like discrete words into like some continuous space that is semantically meaningful. And I was absolutely blown away by that algorithm. And that's when I decided that I wanted to work on natural language crossing and just like understanding language. At that time, I was very wrong, but I thought that English NLP was basically solved. Well, like close to being solved. That was in 2017. So that was right when Transformers started. I was actually right before Transformers.

13:36So I was very wrong, but I decided that I wanted to work on under-researched languages. And basically, I wanted to improve NLP on languages where we don't have that much data. So I went to work for Grab in Singapore and I was basically building the natural language processing pipeline for them, working with Khmer, with Bahasa, with Thai, Vietnamese, and all these different languages. And then I'm skipping a little bit. I did more academic type of work in different countries and I ended up at Stanford, did my PhD there. And after this, I had a small stint into startups and then I went to open it.

14:18Yes. And I remember seeing on your blog or your page a note for quant firms to not reach out to you because you were not interested in hedge fund work. Yeah, I always think it's very important for me to think about the positive impact that I'm having in the world, or at least that I'm trying to have. So that's why this note is there. Yes. And as we were saying just before we started recording, people may have seen you in the GPT-5 video announcement. And you did this very funny demonstration of an app that was built on the fly to teach your partner how to speak French. So people should go check that out.

15:07Exactly. That was a fun one. That was a fun one. So GPT-5 was not that reliable. So I was a little bit stressed that it wouldn't work. But it ended up working. So this was truly live and presumably very rehearsed, but truly live. Actually, right before we did that, like the last rehearsal, it did not work. So I got slightly stressed about that. But yeah, seems like live ended up working well. Yeah, no pressure. But yeah, that landed perfectly. Okay, very cool. All right, so let's unpack some of the things we alluded to in the intro. So we started effectively talking about reasoning. And I'm curious what reasoning means in 2026 that's any different from, you know, a conversation we could have had about 01 or 03.

16:04in particular one of the claims of 5.5 and also my experience as a user is that it's particularly good with messy data which seems to imply that it needs to reason through ambiguity more. What has changed? What I would say is that O1 and O1 Preview were really breakthroughs in the research community about having a model that can think and the longer they thought for, the higher likelihood they would be of being correct. So that was really a breakthrough. But initially, and if you look at old blog posts, you would mostly see math evals and also maybe coding competitions, but things that are really easy to test whether you're correct or whether you're not.

16:54And it also gives you some suggestion about how we were training some of these models. and how I see maybe all of last year and especially the end of last year and the beginning of this year is that we were able to take these algorithms that work with verified rewards like things where we can say you're correct or you're not to the messy real world and really optimize for the utility that we provide to users and like making them more productive. So I think that's what really changed. Okay, so it's the post-training reinforcement learning part largely? Yeah, I would say that's I mean, there's also another big part of it.

17:36Number one, basically the first thing is that, of course, when you develop a new method, the method is kind of fragile and it's not that reliable and it's hard to basically productionize. So this part also improved a lot. But then it's also really basically we had a tool that we could start optimizing for different things. And initially, when we were developing this tool, we were making a lot of simplifying assumptions in the real world, basically. And now we are removing these simplifying assumptions. And at least in post-shrink, we are able to optimize really user utility and make sure that these models are useful and the tasks that we're looking at are useful.

18:17And that's why also now current evals look much more realistic. I mean, if you think about GDP Pival or even if you look at like 3BENCH Pro or 3BENCH, these look way more realistic than, let's say, some code force or like coding competitions that we were looking at with O1. And still on the topic of reasoning, what's ultimately the difference between 5.5 thinking versus 5.5 Pro? Is that just more test time compute, more tokens and more time invested in solving a problem? Yes, basically, it's just a question of how much test-time compute we pour into the model, or we pour into this entire system that we're shipping.

19:02So we've seen again and again, the longer the model thinks for, the better answers we will get. The problem is that these curves that we're talking about are definitely not linear, and there's some plateauing effect, and they kind of look logarithmic in some sense, or depending on which evals. So you can pull two times more compute and actually only get small performance gains. I personally don't use Pro that much because I really don't like weight. I'm pretty impatient. So I don't like waiting for that long. And I know that the probability of being correct definitely improves, but it doesn't improve enough for me to use it.

19:49But there are some people who use Pro and who really love it, especially actually for academic research. And I know especially a lot of mathematicians who are using it. And that's because they kind of just have this in the background that is running for maybe one hour, two hours, and they don't really need to iterate really quickly with the model. And Pro is really good for that. I'd love to reconcile this with what you were mentioning about efficiency earlier per token. So is the idea that you would be able to think longer but also be more efficient, therefore solve the task better? How do the time aspect and the efficiency aspect interact?

20:32Yes. So if you go back to the plot that I was talking about, or I was thinking about, where on the x-axis we have latency and y-axis we have performance, we're basically moving this curve when we say that we improve efficiency more and more to the left. So we're becoming more efficient or we spend less time to achieve the same performance. But what Pro does is that it extends this curve. So it says like, I'm going to think for much longer, but I will have a higher likelihood of being correct. But every iteration of the Pro model also moves to the left. So it also becomes more and more efficient.

21:04The important part is there will always be tasks where you just want to maximize the probability of correctness and you don't really care about latency. For example, if I start a job before going to sleep, I mean, the model has like eight hours. It should just think for as long as it can. And this is what it probably gives you. And in layman's term, what does that mean practically? or how does that work practically? If the model goes in the wrong direction, then it would interrupt itself earlier. Is that one of the axes? So for the efficient, okay, so there's two things. Are you asking for the efficiency?

21:46What does it mean? Yeah, for the efficiency. Yes, largely for the efficiency. I'm just curious how reasoning gets more powerful. Yes, that's a good question. Let me give you maybe a metaphor from like humans. If you have someone who's an expert in a certain domain, and you compare them to some undergrad that is starting in that domain, the undergrad doing that task might take one day, two days, and we'll have to think through a lot of the possibilities and investigate because it never did a certain problem. while someone who's an expert in that field will usually just like know what direction to take.

22:30And it will not spend the time on like investigating 10 different directions because it knows that there's like one that is more likely to be correct. So this is a type of efficiency that we're talking about. It's basically models where we optimized more on like real world problems. And as a result, it was kind of trained to figure out with a higher likelihood which paths of reasoning are more likely to be correct. So this is the part on efficiency. There's also what you suggested is that part of it is the model knowing when it's going down the wrong path. But this is also something that the model can be trained for with reinforcement learning.

23:10It's like knowing, okay, that seems like not a great path. Let me backtrack and let me go and test something else. And if you train the model less, it might realize it's in the wrong path much later. Okay. All right. So it seems like a lot of this goes back to reinforcement learning and post-training. So let's talk about how the different components of modern AI systems work. So let's talk about pre-training, mid-training, and then post-training and spend more time on post-training since it's so important. Starting with pre-training first at a high level and realizing that you may or may not be able to talk about how the things are done, what happened in the context of 5.5 specifically.

23:56You know, big narrative of last year was that pre-training was hitting a wall and was not going to yield much progress. That seems to not be the case at all in 2026. Can you walk us through some ideas for what is happening in pre-training and why it's progressing now in a way that people hadn't predicted last year? For pre-training, I can't talk in a lot of details about what is happening internally. Besides that, the team has been really doing a lot of good work and our models are really getting better and better. One thing that I do want to highlight when we're talking, for example, with efficiency, if you have larger models, the amount of thinking time, so the amount of tokens they will think for, will usually decrease.

24:49And the way that you can think about it is that metaphorically, the model already thinks through its weights when it generates a certain token. So you can decrease the number of tokens that it needs to generate for thinking by kind of increasing the size of the model that you're training. So oftentimes if you just increase the model size, if you basically train, pre-train larger models, you will get better efficiency. And the good thing with larger models is that they can be paralyzed better at inference time. So even though you might think, okay, you actually generated fewer tokens, but by a larger model, so you actually might decrease the overall efficiency of the system.

25:40This is not true because the larger the model is, the more chances you have to actually optimize PC for inference on GPUs. So you will be able to make the overall system more efficient. So that's one thing I wanted to say with larger models that are actually giving you a lot of efficiency. Otherwise, in terms of pre-training, I think it's very interesting. I actually also thought maybe two years ago that pre-training was kind of hitting a wall. And when we see, for example, if we talk just about Entropic, I mean, Mythos seems like clearly just a much bigger model when you look at the cost. The cost of the model, usually that's how you know, by the way, if it's a bigger model, you just look at the cost per token.

26:30And clearly they are getting very good performance just by increasing the size of the model. So I think the field was a very, at least part of the field was surprised about that. There were a lot of conversations about hitting data walls, and it seems like we did not quite hit it. So the larger the model is, the more data it needs to ingest to be trained. And it seems like different companies kind of found different ways to overcome the fact that we don't have that much data on the internet. Is the next frontier, or the current frontier for data, multimodal data? Is it synthetic data? I think synthetic data can probably work well in a data-limited regime.

27:19I think multimodal is an interesting one. I definitely can't talk about what we do internally, but I used to work on multimodal representation learning back in the days, and I always thought that it would really help kind of your reasoning abilities if you have a lot of multimodal data. And I still think this, but for example, like if you look at entropic models, they tend to not be that good on multimodal and they are still really smart. So it seems that it's not as necessary as at least I would have thought in the past. I still believe that once we go to embodied agents and embodied AI, you will learn a lot about the world and you will kind of improve general intelligence and usefulness to users by learning how the world interacts with itself.

28:13But at least looking, for example, on entropic models, it seems like they don't need that much multimodal data to have strong models. And by embodied intelligence, you mean potentially robotics. And so if you use a video that shows how gravity works and how a robot evolves in space, then presumably that would be more useful. Is that the thought? Yes. The intuition that I think many people had, and I definitely felt for a long time, is that it's hard to understand the world only through text. And there will be... It's hard to understand what physics is without really seeing what... For example, you can't understand gravity without really seeing things falling.

28:57and when you look at our models I mean they kind of understand gravity without having seen that but it still seems not obvious like it still seems like they would get it more and like they are still kind of missing some common sense aspects so I do feel like we will improve the common sense of our model by having them interact in the real world but we're still pretty far from that I think and by we I mean just generally the academic community and the AI community seems pretty far from that. Yeah, and while we're on the topic, as a quick detour, that leads us to the concept of world models. So taking your open AI hat off, are you bullish on world models?

Read the full transcript

29:42World models in the sense that, yes, you can try to replicate or simulate things, like basically work in an environment that is simulated yes the problem is simulations are always going to be really hard and not going to be truthful so I think there will always need to be a certain little bit of training that will need to happen in the real world to make sure that the model realizes these mismatches between the simulated world and the real world and I think we as a field have a tendency of optimizing something that is simulated or not quite realistic past the point where this is useful. So that's something that I think we should always be careful with is we spend a lot of time and effort on optimizing something simulated or not quite realistic, and it's great at the beginning, but at some point once you start optimizing too much for something, it's not representative of the real world, and people continue doing that just because that's what they've been doing for a long time.

30:49So I just think people need to realize when to stop that. I don't work with this type of synthetic environment as much, just because I don't work on embodied AI. So I don't know if we heard that yet. Okay, great. All right, so going back to pre-training, mid-training, post-training, let's talk about mid-training. It's maybe something that people have heard about a bit less. The term comes up a bit less. What is it and why is it important? Mid-training is just this idea of something that's between pre-training, as you might realize from the name, and kind of the post-training part of the pipeline.

31:29And really, the idea is if you have high-quality data that is more representative of what you really want in your final model, you should overtrain on that data. so taking a step back here pre-training, what is it? pre-training, it's basically trying to learn everything from the world by learning everything from internet at a high level the problem is that most things on internet are not really useful if you think for example about Wikipedia or like GitHub which is like coding data it just seems like there's way more information in there than some random forums yeah some random farms that maybe not have that much information.

32:17For example, ads. There's also lots of ads on the internet. You probably don't want to train too much on that. But in pre-training, we train on everything. And in mid-training, we basically overweight this type of high-quality data that we think is more useful for training the final model. And this is something I can't talk about what is happening at OpenAI, but it's something that is happening definitely in all the academic community right now and in all the open-source models have this stage of mid-training. Great. Post-training, let's start at a high level by defining what that is. So there's reinforcement learning, but that's not the only part of post-training.

32:51What else is there? It kind of depends how you define the term and where you put the boundaries. In my mind, post-training, including, I'll take it from a very broad sense, which includes all the reinforcement learning and the training for our reasoning models. It's just the idea of having something that knows everything about the world to making something that is useful to people um so pre-training i think about it or the metaphor that i like giving is uh you go in the library and you have a lot of books about everything and in theory you can find all the information that you want in the library but it's much more useful to talk to an expert who has learned these books and that you can ask questions to and they can answer they can answer and like they can understand what you're actually looking for.

33:40So this is kind of the goal of post-training at a very high level, is making something that is useful to users and is easier to interact with. So there are multiple stages. I'll talk mostly, well, I'll talk only about things that are happening outside of OpenAI and kind of the usual stages. There's usually some SFT that is happening. Which is supervised fine-tuning. Supervised fine-tuning, yes. Supervised fine-tuning. And that's actually what, early on, most of the models that we're portraying were only doing supervised fine-tuning. The idea is that if you have humans that can give you the desired final answer, so if you have humans that can give you the gold answer, you can basically clone the behavior of the human.

34:34So this is what we call behavior cloning. The problem with this is that you will never get better than what your ground truth gives you. And humans are actually pretty limited in many sense. So you will never overcome the human labelers that you're working with. The reinforcement learning stage goes from behavior cloning to really optimizing rewards. So the idea is, I don't know what the ground truth is. I don't know what the perfect answer is. But here's how I would say whether the answer is correct or not. And here are the things that I want in the answer. And what you do is you start optimizing, you start having a model that tries to get more reward, basically optimize more this reward function.

35:23That's how we call it. And it goes beyond what you currently have, what humans can do, or at least the humans that you're working with can do. So this, I would say, is the two big stages. Then in reinforcement learning, that depends on which models are being trained. At least in the open source community, it seems that there are different ways of doing that. Reinforcement learning when you have very fireball rewards. So reinforcement learning where it's really easy to say whether something is correct or not, and you can really kind of have a binary reward for this. And that goes back to how we talked about a one and a one preview in the past.

36:02and then you have reinforcement learning without verifiable rewards where maybe I could do pairwise comparisons I can say this answer is better than this other one but I don't really know I cannot quite say this is the perfect answer so of course it's a continuum and there's everything in between but I would say these are the three high level things to think about when you think about post training in general and how people are usually doing it in the open source world is that they take SFT, they clone the behavior that you can collect online or from humans. And then once it's already at a pretty good level, they just do this reinforcement to go beyond what we currently have.

36:47Because if you just started from reinforcement learning, it would be very inefficient. Because the problem with reinforcement learning is that you have to stumble across the right answer, basically. Because how reinforcement learning works is you sample many times, essentially, from the model that you're training. And you say, this one is correct. This one is not. And you say, do more of the one that is correct. So you have to stumble across the right solution. So you're much better off first getting as much, as close as possible to the best you can do. And this is this behavior cloning and then doing reinforcement.

37:20Does reinforcement learning create new capabilities or does it make the model better at existing capabilities? It's really hard to say because pre-training, what is trained on all of the internet, arguably already has all capabilities in it. So it would be even hard to answer this question scientifically because arguably everything is already there. What I would say is that if you look at models that we were training or that we were post-training like two years ago in the open source world, for example, I worked on one of them, Alpaca, where we used 50 ,000 examples for SFT. And like now when you look at reinforcement learning from models like Kimi or from DeepSeek models, it seems that they are closer to 1 million data points.

38:14So definitely people scaled up a lot the reinforcement learning stage. and from this it seems that they've learned like new capability like this reasoning aspect, this fact that you can check your answer and try to improve it. So you can really think for longer to get a more correct answer. So all this to say that arguably everything is already in pre-training but we were definitely able in the last one year and a half even in the open source world to have more capabilities after reinforcement than we used to before. I heard several times that reinforcement learning is pretty finicky and hard to scale.

38:59And part of the reason why we, as an industry, didn't do reinforcement learning as part of the initial kind of LLM sort of progress curve was precisely that, that it was hard to make work. What is hard about scaling RL? Is that a question of data sets, knowing where the rewards are, or something else? I would say most people who did not work in reinforcement learning in the academic and research community up to two years ago probably thought reinforcement learning just doesn't work and is too finicky to work with. I used to be that type of person. And actually, when I saw ChatGPT come out, they had this blog.

39:38I was not at OpenAI at the time. I saw this blog that says that they use reinforcement learning. And my first thought was, I can do the same without reinforcement learning. Because this is just an overcomplicated method. And this is actually the product that we started working on with Alpaca, was exactly, let's try to reproduce that only using SFT, just by doing this behavior cloning. Yeah, and like, for example, Yann Lequin famously gives this metaphor of, oh, the reinforcement learning is just like the cherry on the top. So I think that was really the intuition that most people had. it seems that after crossing a certain scale of models that know basically everything about the world and what we call like good priors about the world it seems that reinforcement learning just started to work and this is not only with LMS robotics seems to have or it seems to be entering the same stage where they're realizing that actually it used to be very finicky but now that we use models that like know already everything about the world it actually learns pretty well.

40:37Now to answer your question about what is still complicated with reinforcement learning, one is an infra aspect. So just like systems in general, reinforcement learning, you have at a very high level, basically to sample, as I said before, many answers and say what is correct and what is not. And this sampling is just very expensive and you have to do it at scale. The other issue that also in the open source world people are seeing right now is that when we are training more agentic systems, you only know whether you're correct at the end of your very long rollout. So you get very little information per token of whether you were correct or not.

41:30And it's hard to say, it's hard to basically do attribution. It's hard to say what part of your entire answer was the one that led you to being correct. So that's more of an issue on the machine learning side. The ideal world in machine learning is when I can say exactly like this thing was good, do more of that. And the problem again with these agentic systems and reinforcement agentic systems that you don't really know which part was good or not until you arrive at the end. That's another big issue for reinforcement learning. What's the current frontier of reinforcement learning? It seems like there's a jungle of acronyms like GRPO and other techniques.

42:12What are you using? What are you excited about? What do you think is promising? So I can't talk about what we're using, but for example, in the open source world, GRPO seems to be working very well. And people used to have different methods like PPO and DPO, and people seem to have really converged to this one. The big difference with other methods is that you, again, you do this simple method that I told you about sampling as many answers as possible, and you say which one is correct. So in some way, GRPO is a very simplistic method. And in general, we saw over and over again in machine learning that the simplest method where you can scale up in terms of compute usually is the one that ends up working the best.

43:03And that is kind of what is happening here, at least in the open source world. As you describe some of the challenges, question crossed my mind. You know, you often hear that AI systems are not built, they're grown. How are you characterizing it as well? What part is science versus a craft or trying multiple things? than just keeping what works best in your day-to-day life? Yeah, that's a great question. I think how it usually works is that it starts being craft. People just try out many things and they start building a mental model of what works and what doesn't. And over time, we move from this craft land to more science.

43:47Science is... or like more scientific approach are really the ones that like first end up working. It's hard. It's very rare that you take a really scientific approach and you say like this is the optimal thing to do and you do it and it just works. Like people just there's some sense of alchemy. People just have like a good flair for something and they make it work. And then other people or that person starts trying to improve what we are doing by being very scientific. And I would say this happens over and over in machine learning. So first craft, then science, and both are really important, but it's different stages of the pipeline.

44:35In terms of engineering, this is definitely something that is always necessary. so I would say most researchers have moved to being relatively good at like at least I wouldn't say good engineers but good at working in like complex systems and like figuring out what they need to try out and the systems and the infra that we have has become more and more complicated so definitely the work required changed over time Fascinating. All right. So still in reinforcement learning and circling back to some of the things you said at the beginning. So if I want to make my model better at computer use or genetic coding or whatever domain, then I would spend a particular amount of time doing specifically reinforcement learning for computer use and putting together a data set and then coming up with rewards.

45:35Is that how it works? Like you just pick one problem and you just do reinforcement learning specifically for it? To be clear, I talk more about reinforcement learning because also this is like the part I know the best. And this is what I've worked, like pushing I've worked on for a long time. We talked about mid-training before. Like all these things are also extremely important and you can improve it in different parts of the pipeline. As I said before, the closer you are from the final stage of the model, usually the smaller the scale of the training becomes. So you can iterate fast on that because now you can iterate in terms of days rather than iterate in terms of months.

46:17So usually people start from this fast iteration loop and then they go deeper and they make bigger changes across the entire stack. So this is not to say that only reinforcement learning matters. I'm really not saying that. but it's just like that's why people will start doing changes and then that will permeate and we will go deeper into the stack. So this is how it works. And like in the open source world, it's very much like that too. I think you see way more post-trained models than you see new pre-trained bases. And you see way more like improvements in like the algorithm and that's why we talked about, I mean, GRPO, DPO, BPO, like there are so many XPOs And that's because people can iterate really quickly on this final stage of the pipeline.

47:06And the jagged nature of those models, does that come from this approach of picking this problem and that problem? And therefore, it's going to be excellent that those problems are not as good as other problems? Or is that a more fundamental characteristic of AI models? There's definitely some of that. for sure if you optimize more on specific types of problems, you will be better in that setting. I would say, this is my intuition, is that it's less about the exact problems that you're optimizing on, and it's more about the class of problems that you're optimizing on. So for example, if you are really good at math competitions, your model will probably be pretty good at coding competitions.

47:53So it's not about the domain. It's more about the skills that are necessary and the way to think and this horizontal capabilities that you need for performing these tasks. And that's what I think you're usually seeing when someone is really bad at something, it's actually bad at that in any domain, in any language. So you have to think about this domain and then this generalization of this domain, not necessarily per domain capability. So speaking of generalization, So there's been that clear evolution from math and coding success to now starting to cover different areas. So that's the whole GDP valve thing where like across the economy, different areas are being evaluated in terms of like model performance.

48:41It's sort of the same question. Is that the result of overall model progress or is that deliberate? it. Okay, now we're going to take this part of the economy and build a data set for it and do mid-training and do post-training. How does that progress work from those very specific domains to generalizing to the rest of the world? It's definitely something that we actively push on. I think people are realizing, I mean us and also other companies, that we are moving towards this world where we want to really make products that are useful and like improve like productivity of people um and and help people in the day-to-day life so i think there's a there's a very active move to deciding what are the domains that we should be prioritizing what are um now that we know we have an algorithm that we can apply in different places uh what we are constrained by is more collecting the right data having people who really care about a certain problem work on that problem.

49:49But there are not that many people who can do these things. So you really need to prioritize. So this is, yeah, it's a very active, it's a very active, proactive kind of approach here. And in general, I would say the performance of the model really depends on like the number of people who care about the final output of the model who are looking at that model. So if they start looking more on specific verticals, like these verticals will improve really quickly. But again, we don't have that many of these people that can do these things. But to unpack something that you alluded to, I think a minute ago, do models actually generalize now more, especially from a reinforcement learning perspective?

50:38So making a model very good at domain A or B then is likely to make the model better at C, regardless of the amount of effort you put into developing rewards for domain C? So I think there are different axes of generalization. One, there's an algorithmic generalization. And that's like really, can I use the algorithm that I developed or this black box that I developed for domain A, and can I use it for domain B? and at least again like even talking about the open source wall it's really seems that like people are able to do that they take grpo they apply it in like many different places and it just works so that generalization seems to be relatively good which which is why we're seeing a lot of progress otherwise it would be hard to make progress then there's the generalization of the model that is trained on one particular data set and this is what i was alluding to before is at least my mental model is that generalization happens in terms of capability.

51:41Like if the capability is the same, you will see generalization across domains. Again, like different languages, like coding, like you can optimize for C++ coding, for having a good C++ model with very little training on C++. Partly because this pre-trained model, or very little RL in C++, partly because this pre-trained model has seen all of C++. And so it already kind of like understands the basics of that language. So that type of generalization definitely happens. The generalization that I think is harder are these, when we don't have these like horizontal capabilities. So I'll give you one concrete example.

52:26If my model is very intelligent in terms of being correct on like competitions, I usually take that example because it's like somewhat constrived at like math competitions, like coding competitions. From a human perspective, people that are good at these things are usually just smart. And if they are smart, or like someone might think that at least, that they are just smart. And if they are smart, they can actually do other things too. But that is really not true. And that type of generalization is really not true because many things where we need to have humans working on like expert domains, like the world is very messy.

53:00and these coding competitions and math competitions are extremely well-specified. And you need to have the capability of understanding underspecified tasks, understanding how to deal with the messy world and understanding what are even the resources that you need to answer the question. If you look at the math competition, you usually have everything in the prompt. It's like you have five lines or maybe 15 lines and it's like all the information that you need to answer this question. In the real world, if I'm a consultant, if I work in like finance, I need to go on the internet. I need to like find and extract different information just to understand before doing any of the reasoning, just to be able to do that reasoning.

53:47And this type of like horizontal capability is the thing that doesn't usually, like you generalize if you have that horizontal capability, but in many cases we don't have that horizontal capability. so yeah that's why we hallucinate actually in every domain like when you have hallucination of lms if if a model is really bad at saying that it doesn't know that usually happens in every single domain you won't have like one domain where the model is extremely calibrated about its knowledge and another domain where it's not and as a quick day tour is uh is is hallucination also a reinforcement learning problem where you reward the behavior to say i don't know when it occurs?

54:29John Schulman has a great presentation about that, I think from like one or two years ago, where he was saying that if you do behavior cloning, so this like SFT that we talked about before, you will basically reward and optimize for hallucination because what will happen, or you could optimize for hallucination because what will happen is if your model doesn't know about something, but now you say that the right answer is to say that something. So I'll be very concrete. If the model doesn't know about a paper, and now in an answer that you give, that is given by a ground truth answer given by a human, you say, here's where I got the information, and then you cite that paper.

55:13Like what you're actually optimizing the model to do is citing something that doesn't exist because it doesn't know that that paper exists. And so John Solomon had this great presentation saying like SFT is going to force like a hallucination while in reinforcement learning, given that, as I said, you kind of sample from the model in the first place, extremely unlikely that you sample something that it doesn't know and it's correct. That's like extremely unlikely. So you will never reward that behavior. You will only sample things that it doesn't know and being incorrect and then you will kill that behavior.

55:49So hallucination, at least the intuition that people have is that it can come, for example, from SFT and it can come from this like porturing pipeline. But if you have good reinforcement in pipeline, that shouldn't happen too often. And going back to generalization as well, are there examples where actually getting better at one domain makes the model worse at the rest? A little bit to what you were saying about like some people are very good at math, some people are very good at English. Pretty often they're not the same people. in domains usually not what will happen though is um you will make decisions based on which domain we optimize for and if you optimize for one domain you will be able to optimize less for another one so it's not necessarily that optimizing for one thing will make the other one worse it's just that as a result you can optimize less for the other one because your compute constraint your data constraint, you have like your human bottleneck also in terms of that work.

56:53What does happen is you can have negative kind of generalization, like bad generalization, or negative transfer more for these horizontal aspects of the model. So I'll give you a very concrete example. Explicit instruction following versus implicit instruction following. If I have a model, And this is, we often hear, for example, from OpenAI models that they tend to be really good if you tell them exactly what you want. But as a result, sometimes we hear also that they're like less good if you are not as specific about what you wanted. For example, if I make a typo and I say like change this file and I make a typo in this file, an extremely good model at explicit instruction following will change the wrong file, the one that has a typo.

57:47But humans would probably realize that you made a typo. And as a result, there are cases where this explicit instruction following goes against this implicit instruction following. So you will have cases where basically these horizontal capabilities go against each other. And maybe to close on this whole reinforcement learning conversation. So is your sense that as we progress from being excellent at coding and excellent at math and move to the rest of the economy, do you think that the rest of the economy is a tractable problem? Do you think we can get to the same level of performance ultimately?

58:27Yes, but I was like, yes, we can. I don't think there's anything like really deeply special about these domains where we cannot optimize and where we can get the same with other domains. The bad is for at least two reasons. The first one is most of the people working on these models are pretty good at coding and they really care about coding because that's what they use as the each day kind of drivers. and there's nothing better than the user being also the one who like trains the model because like then they understand the issues it's um it's very hard to really like for me for example it's very hard to really understand like what should we change on the like on like legal uh aspects of the model if i don't understand anything about the legal domain um so that's one thing The other thing that you will often hear about, and I mentioned also briefly about before, is this kind of verifiable rewards.

59:26There are domains where it's easier to say whether something is correct or not. For example, in the case of cyber, like you mentioned that before, that cyber has been improving a lot, cyber capabilities as our models. And this is because in cyber, it's extremely easy to say. Are you correct? Did the cyber issue that you find is a real issue or not? It's very easy to test it. So there are domains where reinforcement learning is just easier to apply. But there's nothing, I would say, in the capacity of the model that is constraining the model to be as good at legal and medical and other domains.

1:00:08So the short answer is we know less about these domains. and definitely there are some domains that are easier to optimize for in reinforcement. Great. Let's talk about evals for a minute. That's a hugely important topic. Maybe to start, why is it so hard to evaluate a model in the first place? Evaluation has been harder and harder as models become better. And that's because the tasks that we ask to the model become more and more general and more and more open-ended. So now I maybe just say, build me a website that does X. Well, before in the past, I would just be like, hey, is there a specific bug in this implementation that you have?

1:01:00And it's much easier to say whether it is a bug because I can extract, I can have a human that says, here are all the bugs that you have, and then you can apply that automatically. while the website one is very hard to know what is the optimal answer because there are many good answers. There are many good ways of building a certain website. This open-ended nature of models really makes evals harder. There's also another issue is that models in specific axes are becoming better than the majority of humans. And so we have fewer and fewer humans that can actually evaluate these models in particular axes.

1:01:37so that's definitely constrained another one to be honest is kind of cultural most people want to improve the model and they they think that the best way to do that is kind of training the model when in reality finding issues and like making sure that we can quantify improvements is just as important if not more important but there's always this like cultural gap that was especially true I would say in the academic world up to like two years ago when evals were always fixed, benchmarks were always fixed, and even data sets were kind of always fixed, maybe let's say four years ago. And there was like a mentality shift of like, okay, data is actually critical.

1:02:18And now there's a lot of people working on data. And I think evals were still not quite there. People don't really fully, everyone knows that it's important, but like people don't really understand like how impactful it could be to work on evals. So actually my first product that opened AI just came in and I was like, I want to work on data and evals because I know that this is the thing that no one is working on. And as a result, I know that's super impactful to work on that. And yeah, the tide is shifting, but not fast enough. And is the pace of progress in model as a judge and AI, evaluating AI, is that moving as fast?

1:02:53Is that a distinct part of research? Or is that fundamentally the same idea or the same techniques? It's really fundamentally the same method. So there's like nothing. Also, most of the things that we do in evals, especially now that we have reinforcement learning, could just be applied nearly exactly as is during training. So that's another reason actually why evals are so complicated is that every time you build an eval, you actually build a way to build training datasets. So now you're going to optimize that training dataset. Well, even if it's not that eval, it's going to be the same type of data.

1:03:27And now you're going to do super well because we have this generalization of capabilities that I was telling you about, you will learn that on that other data set and now you'll become really good at that eval and that eval will become obsolete really quickly. So that's also an issue with evals. But yeah, to come back to your question, the model as a judge, it's really important. And I think it's one probably of the most important things because as we get better models, we have this self-reinforcing loop and we have this capability flywheel where better models become better teachers for other models.

1:04:05And this is really important for training, but then you can also do the same thing for evaluation. So a lot of my team works on that. And I think it's really critical is to work on this model as a judge kind of framework. Okay, fantastic. All right, so as we get towards the end of this conversation, I'd love to zoom out a bit and get your sense for where things might be heading. Obviously, it's incredibly hard to make predictions on AI years out, but let's call it the next 12, 18, maybe 24 months. Is your sense that things are going to continue progressing or are we heading towards something that could feel more like a discontinuity?

1:04:49In terms of progress, as I was saying before, I think it's always continuous. Now, the feeling of discontinuity will happen. It did happen three months ago with coding or four months ago with coding. And I think that will happen now in every other domains. Like most people are not feeling the same way, like kind of the capability of our model and the usefulness of our models, the same way as like coding and like software engineering is feeling right now. So this will definitely permeate, I think, through many other verticals. Now in terms of like capability bump in terms of, let's say, the verticals that we're already looking at, I think it will be more continuous.

1:05:28And there will never be big discontinuities. Most of them are always local discontinuities, but you zoom out and it always just feels pretty smooth. It's not always like this, but that has been the case most of the time. And I can definitely not predict when is the next big discontinuity. What is your sentiment on this general concept of accelerating loops in AI? So whether that's continual learning to make models more current and able to learn faster to this broader concept of AI building AI, like in an increasingly automated way, fact versus fiction. And what are you excited about? I'm extremely excited about continual learning.

1:06:11I think we haven't quite cracked it. I mean, we have like codex memories and that is helpful, but it's definitely not like the end state. I have a friend who always like tells me about again another type of plot that we should be looking at which is x-axis time, y-axis utility that you provide to users and right now or like usefulness basically of the models and right now actually most models at day zero if you just drop them in a company arguably they are more useful than most new employees so they start higher at T0. But then across time, they are mostly constant because they don't really learn kind of company knowledge.

1:06:57They don't really learn to be more efficient over time on doing the things that they are doing while humans learn really quickly. And what is important is kind of this integral or like kind of the area under the curve of these curves. And as a result, I think humans are still more useful in many cases. And that's why what we will need is to make like continual learning is to make this curve now monotonically increasing over time and basically make models more and more useful the longer they work in a certain environment. So I'm extremely excited about it. I'm actually surprised that we're not quite there yet.

1:07:37Three years ago when ChatGPT came out, I remember I was doing a startup with friends and we were thinking about working on continual learning and like personalization and like memories in general. and we're like, ah, OpenAI is going to do that in the next six months. They have all the data, they're going to figure it out and they have all the users and their models are going to learn super quickly from users. And three years later, I don't think we're there yet. And quickly in layman's terms, what is the fundamental difficulty? It's a good question. I actually don't quite know, to be completely honest with you.

1:08:10I don't quite know why it's taking us that long to figure it out. it's this type of domain that I think if we really put enough resources behind it we would figure it out of course especially when we talk about this memory inside of a company there's big questions about permissions and there's a lot of questions about privacy and what you can share and what you cannot across users but for a single user even for a single user we're not quite there And I don't quite know why, at least at the high level that I can talk about. I don't know why. Yeah. What you bring up is, I think, really interesting for AI builders and investors and startups, which is this question of the models getting increasingly smarter within an enterprise.

1:09:08In particular, there's like this whole tension between what the models are able to do and then what a lot of people have built around the model. So, you know, a year or two ago, it was RAG. These days, it's all about harnesses for agents. And a lot of people are wondering whether the models are going to end up eating the harness, whether the harness is just a temporary thing. From your perspective, what do you think happens? Yeah. I think harnesses can really improve the capability of a model right now. I think given that we're seeing this really fast progress in terms of capability, I personally wouldn't push that much on the harness.

1:09:51that unless it's like the harness is something for like very concrete goal that you're trying to achieve right now. So certain companies, like if they are focused on like a specific vertical, they want to go from this like 80 % maybe reliability to maybe the like 85%. And like Harnesses will give them that. And I think that's like very important, but like they will, they need to do it while knowing that they will have to retune that harness in the future. And I think that's totally fine. If you try to have like a general harness that will sustain over time, I don't think that will work. Harnesses for specific domains as a short-term thing that you need to do, I think there will always be so much you can do in harnesses.

1:10:42And if anything, I think everyone should do more of that. if they have a specific problem in mind, because we're leaving so much on the table without a good harness. Arguably, if we just, I think if we froze the models that we have right now and you really worked on the harness and like maybe like we also spend more time like training with like a great harness, I think people would really feel the AGI in every single domain or could already feel that in every single domain. But given that we're not freezing it and we're going to continue training better and better models. I think the harness, we don't really understand what the final harness will be.

1:11:21And it's not, and like it will always change. Some question about applications. So we alluded to your progress in different verticals. And that was, you know, GDB Val in general, but also how to bench telecom, which does complex customer service workflows and then progress against finance agents, automating 88.5 % of internal investment banking modeling tasks and then 51.1 % on Office QA Pro. So bit by bit, you're doing more and more of this. So do you think people should be building applications anymore or is ultimately, as we get closer to AGI, all of this going to be part of the model capabilities?

1:12:09which is there's so much space on pushing, pushing for like external companies or like startups pushing on specific verticals. I think there's so much space for that. The reason why is because a lot of people kind of think about intelligence and quotations and all like kind of like raw capability as being the real bar neck, but I don't think that's true. I think most of the time the bar neck is the last mile. it's like making sure that the model has access to has the right permissions or has also access to the right connectors and things like this and we are going to be very focused on this general aspect and I think there are other companies that should be focused on more of the verticals and providing maximum value of what we currently have.

1:12:59So I think there will always be a lot of space left for this last mile in different verticals. And I would highly encourage people to continue working on that. And maybe one day when we stop making horizontal progress, which I don't think is anytime soon, maybe we will start focusing on that. But yeah, that's not what we're doing now. Okay, well, that feels like a very optimistic note, at least for the startup ecosystem to end up on. Thank you so much, Jan. This was terrific. Really enjoyed it. Thank you so much for spending time with us. Great. Thanks, Matt. Hi, it's Matt Turk again. Thanks for listening to this episode of the Matt Podcast.

1:13:40If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests. Thanks and see you on the next episode.

From the publisher

What actually goes into pushing the boundaries of post-training and reinforcement learning for frontier models like GPT 5.5? In this episode of The MAD Podcast, Matt Turck sits down with Yann Dubois, lead of OpenAI’s Post-Training Frontiers team, to unpack how reinforcement learning has evolved beyond verifiable math and coding competitions to actively optimize for messy, real-world utility. They explore the delicate balance of test-time compute and model efficiency, the growing challenge of building open-ended AI evaluations, and why the ultimate opportunity for startups lies in mastering "last mile" application harnesses rather than competing on horizontal capabilities.


(00:00)44 - Cold open

(00:34) - Intro

(01:30) - Why recent AI progress feels like a step function

(04:13) - Model reliability & the rollercoaster of shipping 5.5

(07:33) - How OpenAI structures vertical and horizontal teams

(09:49) - Improving model efficiency and test-time compute

(12:32) - Yann Dubois' journey from Switzerland to OpenAI

(15:37) - Reasoning in 2026: Real-world utility vs verifiable rewards

(18:34) - GPT-5.5 Thinking vs Pro: Scaling test-time compute

(20:09) - How reasoning models become more efficient

(23:23) - Pre-training scaling and overcoming the data wall

(27:03) - Multimodal data, synthetic data, and embodied AI

(31:05) - Demystifying mid-training and post-training

(37:21) - Does RL create new capabilities in AI?

(38:53) - The challenges and frontier of scaling RL

(43:09) - Is building AI models a craft or a strict science?

(48:21) - How AI models generalize across different domains

(54:18) - How reinforcement learning cures AI hallucinations

(56:04) - Negative generalization and conflicting instructions

(58:05) - Can RL scale to law, medicine, and the broader economy?

(1:00:19) - The evaluation bottleneck and Model as a Judge

(1:04:21) - Continuous AI progress & continual learning

(1:08:49) - Will foundation models eat the agent harness?

(1:11:23) - Why startups should focus on the last mile of AI

More from The MAD Podcast with Matt Turck

All 44 episodes
OpenAI's Yann Dubois: Why AI Progress Suddenly Feels RealThe MAD Podcast with Matt Turck · 1 h 14 min
Listen in VO