What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)

26 Nov 2025 · 1 h 5 min · 23 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

AI progress and the “slowdown” narrative; why reasoning models are a new paradigm; how reinforcement learning (RL) and synthetic data change capabilities; what’s next for pre-training, multimodal reasoning, and generalization.

Guest backgrounds

Łukasz Kaiser is a key AI architect and co-author of “Attention Is All You Need,” inventor of the transformer architecture. He is now a leading research scientist at OpenAI, working on reasoning models behind GPT-5.1.

Key claims

Pre-training scaling still works (loss decreases with compute), but reasoning models yield larger gains per dollar. The slowdown narrative is misleading because capabilities improve via new paradigms and rapid post-training/product iterations. Reasoning models generate “chain-of-thought” tokens and are trained with RL where correctness is verifiable (math/coding/science), making them better at verification and self-correction. Multimodal reasoning lags text reasoning, and generalization remains an open question.

Notable examples

Old chat hallucinating SF Zoo opening hours; reasoning models solving math/coding better but failing simple multimodal dot-count puzzles unless they think longer or use tools (e.g., GPT-5 Pro running Python to count shared dots).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Debunking AI Slowdown Narratives

1:24 to 2:20

Lukasz discusses the misconceptions about the slowing down of AI progress.

“There was a narrative, at least in some circles, maybe outside of San Francisco, throughout the year that AI progress was slowing down, that we had maxed out pre-training, that scaling laws were hitting a wall.”

Understanding AI Progress

2:20 to 4:25

Exploration of the exponential growth in AI capabilities and its implications.

“There is this thing that's happening in AI.”

The Rise of Reasoning Models

4:25 to 6:48

Lukasz explains the shift from pre-training to reasoning models in AI.

“I feel pre-training in some sense, it's on the upper part of the S, But it's not like scaling clause for pre-training don't work.”

Capabilities of Current AI Models

6:48 to 8:02

Discussion on the impressive tasks current AI models can perform.

“I think one of the biggest things that I would say people kind of know on the inside that others don't is that already right now, it's not about the progress.”

Identifying Areas for Improvement

8:02 to 11:16

Lukasz outlines the obvious improvements needed in AI models and their training.

“And second, can you give us some examples of the obvious thing that you need to fix next and that the industry will fix?”

Deep Dive into Reasoning Models

11:16 to 14:00

An in-depth explanation of how reasoning models differ from traditional LLMs.

“I think the big question that people have in their mind is how much better will it make them?”

Reasoning Models in Science and Beyond

14:00 to 16:40

Explore the strengths and limitations of reasoning models in various domains.

“You can do this in science to some extent, right?”

Reinforcement Learning and Model Training

16:40 to 19:20

Understand the impact of reinforcement learning on model performance and reasoning.

“And maybe then it will expand to domains that go beyond where it shines today.”

The Evolution of AI Research

19:20 to 21:40

Learn about the journey of Łukasz Kaiser from mathematics to leading AI innovations.

“Going back to chain of thought, how does that actually work?”

The Development of the Transformer Model

21:40 to 28:00

Discover the collaborative efforts behind the creation of the Transformer model.

“I'd love to talk a little bit about your story.”
Show all 23 chapters

Transition from Google to OpenAI

28:00 to 30:22

Learn about the cultural differences and personal motivations behind the shift from Google to OpenAI.

“Google was also an amazing place at that time that they would very happily let you work on whatever you wanted.”

Research Team Organization at OpenAI

30:22 to 31:32

Discover how research teams are structured and how projects are prioritized at OpenAI.

“I think in general, the tech labs are more similar to each other than people think.”

Future of Pre-Training in AI

31:32 to 36:43

Understand the advancements and challenges in pre-training models for AI, including resource allocation and efficiency.

“There's always some smaller teams doing like more adventurous stuff like diffusion models at times.”

Interpretability in AI Models

36:43 to 39:41

Explore the progress and limitations in understanding complex AI models, particularly in interpretability.

“We both understand that you can distill this amazing big model and there is now enough GPUs to actually train it.”

Improvements from GPT-4 to GPT-5.1

39:41 to 42:00

Learn about the significant changes and enhancements in the transition from GPT-4 to GPT-5.1.

“I'd love now to talk about 5.1 and do a little bit of a deep dive on all the latest stuff that you guys have released in the last couple of weeks, which has been very impressive, in particular as a user.”

Post-Training Improvements and Model Naming

42:00 to 45:00

Explore how post-training adjustments enhance AI model responses and the evolution of model naming conventions.

“Some part of that is because reinforcement learning can now use tools and gather data.”

Reasoning Models and Their Limitations

45:00 to 48:20

Discuss the capabilities and limitations of reasoning models in AI, particularly in solving tasks.

“And the thinking models are the ones that do more research.”

Generalization and the Future of AI

48:20 to 51:40

Delve into the importance of reasoning for generalization in AI and the challenges of achieving it.

“I think it's quite interesting to keep in mind.”

Architectural Innovations in AI

51:40 to 55:40

Examine emerging architectural changes in AI, including multimodal approaches and the role of GPUs.

“I always thought this was the key topic in machine learning in general and in understanding intelligence.”

Long-Running Workflows and Model Engineering

55:40 to 56:00

Unpack the implications of long-running workflows in AI models and the associated engineering challenges.

“and it's fairly clear what I'm saying, please just implement it so it runs fast on this eight machine setup or 100 machine setup.”

Unpacking Codex Max and Long-Running Workflows

56:00 to 58:00

Learn about Codex Max's capabilities and challenges in running complex tasks.

“They say we'd like an AI intern by the end of next year.”

The Evolving Role of AI in Various Professions

58:00 to 1:02:30

Explore how AI is impacting different industries and the need for human oversight.

“And the other thing is transformers have this thing called context.”

Future of Robotics and Multimodal AI

1:02:30 to 1:04:50

Discuss the future of robotics and advancements in multimodal AI learning.

“Some of the topics that one may see are things like continual learning, world models, robotics, embedded intelligence.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00There is this thing that's happening in AI, and in AI, every week now, a lot is happening. Fundamentally, if you look at AI progress, it's been a very smooth, exponential increase in capabilities. This is the overarching trend. It's not like pre-training fizzled out. It's just we found out a new paradigm that at the same price gives us much more amazing development. And this paradigm is still very new. I think one of the biggest things that I would say people kind of know on the inside and others don't is that already right now, It's not about the progress. There are so many things Chad or Gemini, any LLM can do for you that people just don't realize.

0:36You can take a photo of something broken, ask how to repair it. It may tell you. You can give it a college level homework and it will do it for you. Hi, I'm Matt Turk. Welcome to the Matt Podcast. My guest today is Lukasz Kaiser, one of the key architects of modern AI who has quite literally shaped the history of the field. Lukasz was one of the co-authors of the attention is all you need paper, meaning he's one of the inventors of the transformer architecture that powers almost all the AI that we use today. He's now a leading research scientist at OpenAI, helping drive the second major paradigm shift towards reasoning models like the ones behind GPT 5.1.

1:09This episode is a deep exploration of the AI frontier, why the AI slowdown narrative is wrong, the logic puzzles that still stump the world's smartest models, how scaling is being redefined, and what all of that tells us about where AI is heading next. Please enjoy this fantastic conversation with Lukash. Lukash, welcome. Thank you very much. There was a narrative, at least in some circles, maybe outside of San Francisco, throughout the year that AI progress was slowing down, that we had maxed out pre-training, that scaling laws were hitting a wall. Yet we are recording this at the end of a huge week or couple of weeks with release of GBD 5.1, GBD 5.1 Codex Max, GBD 5.1 Pro, as well as Gemini, Banana Pro, Grok 4.1, almost three.

2:01So this feels like a major valuation of that narrative. What is it that people in Frontier AI labs know about AI progress that at least parts of the rest of the world seem to not understand? I think there is a lot to unpack there. So I want to go a little slower. There is this thing that's happening in AI. And in AI, every week now, a lot is happening. New model, coding, doing slides, self-driving cars, images, videos. It's a nice field that doesn't make you be bored for a long time. but through all of this it's sometimes hard to see the fundamental things that are happening and fundamentally if you look at AI progress it's been a very smooth exponential increase in capabilities this is the overarching trend and there has never been much to make me at least and I think my colleagues in the labs believe that this trend is not happening it's a little bit like Moore's Law, right?

3:11Moore's Law happened through decades and decades. And arguably, you would say it's still very much going on, if not speeding up with the GPUs. But of course, it did not happen as like one technology was bringing you there for 40 years. There was one technology and then another and another and another and another. And this went on for decades, right? So from the outside, you see a smooth trend. But from the inside, of course, progress is made through new developments in addition to the increase of computer power and better engineering. So all of these things come together. And in terms of language models, I think there was a big pivotal point.

3:53I mean, one point was, of course, the Transformers when it started, but the other point was reasoning models. And that happened, I think, oh, and preview was a bit a year and a month ago or something like that. So we started working on maybe three years ago, but, but you, you, you know, it's very recent. If you think of it as a, as a paradigm, that's a very recent thing. So, so it's always like these S curves, right? It starts, then it gives you amazing growth and then it flatlines a little bit though. Yeah, we'll get to the pre-training, right? I feel pre-training in some sense, it's on the upper part of the S, But it's not like scaling clause for pre-training don't work.

4:33They totally work. Like if you, what scaling clause says is that your loss will log linearly decrease with your compute. We totally see that. And clearly Google sees that and all other labs. The problem is, you know, how much money do you need to put into that versus the gains you get? And it's just a lot of money and people are putting it. But with the new paradigm of reasoning, you can get much more gains for the same amount of money because it's on this lower and there are just discoveries to be made. And these discoveries unlock insane capabilities. So it's not like pre-training fizzled out.

5:11It's just we found out a new paradigm that at the same price gives us much more amazing development. And this paradigm is still very new. It happens so fast. I think if you blink, you may miss it because it was basically you had chat 3.5, right? GPT 3.5 in chat. And it would give you answers and it used no tools, no reasoning. It would answer you something. And now you have chat. And, you know, if you were not into it, you may have blinked. And it also gives you answers. And you may say, OK, it's more or less the same, except the chat now will, you know, go look on some websites, reason about it and give you the right answer instead of something it memorized in its weights.

5:53I very much used to like this example of, you know, when, what time does the SF Zoo open tomorrow? Like the old chat would tell you, right? Totally hallucinate from its memory an hour that it read Zoo Opens on probably the zoo's website from five years ago, and it didn't know what's today or tomorrow. So it would just assume it's a weekday. Chat now knows what's today because it's in the system prompt. It goes to the zoo website, reads it, extracts the information. If it's ambiguous, probably checks three other websites just to confirm and then gives you the answer. But if you blink, you may think it's the same, but no, it's dramatically better.

6:31And, you know, as a consequence, since it can read all the websites in the world, it can give you answers and stuff that it wouldn't be able to even touch before. So there is tremendous progress, right? And it happens so fast, it may even be missed. I think I think one of the biggest things that I would say people kind of know on the inside that others don't is that already right now, it's not about the progress. Like there are so many things, Chad or Gemini, any LLM can do for you that people just don't realize. Like you can take a photo of something broken, ask how to repair it. It may tell you.

7:11You can give it a college level homework and it will do it for you. So that's absolutely amazing. So there is an education gap to some extent? Well, it just happened. I mean, if you think you said codex, right? You know, programmers are conservative a little bit. I still use Emacs from time to time. All the coding tools like, okay, it will complete one line for me. But people are very like, this is my editor. I write code here. Now people are like, no, this is codex. I ask it to do stuff. I will fix it later. Right. But I think it's the recent few months when the transition happened from, you know, people using it sometimes, but rarely, to now basically this being how a lot of people work in coding.

7:52That's quite big. I'm not sure everyone's aware of it, but it's also like if you don't do programming, why would you be aware of it? I do believe, though, that this will come to more and more domains. To the point of all of this being very new and somewhat sudden, something that you or I hear from time to time when talking to people is that part of the reason why people are so optimistic is that there is a lot of low hanging fruit, very obvious things to improve for those models in the next few months. First of all, do you agree? And second, can you give us some examples of the obvious thing that you need to fix next and that the industry will fix?

8:32Yes, there is a ton of extremely obvious things to fix. Larger part of this ton is just hard to talk about on a podcast because it is in the engineering part. Every lab has their own infra and their own bugs in the code. Machine learning is beautifully forgiving in some sense, in contrast to old software engineering, which would just yell at you when you made a mistake. You know, our Python coded will generally probably run except much slower and give you worse results if you run it wrong. So you realize, oh, no, it was wrong and you improve it and the results get better. These are huge distributed computing systems.

9:11They're very complex to run. So there is a huge amount to improve and fix and understand in the process about just how to train your model and how to do RL because RL is more finicky than pre-training. It's harder to do really right. So every day, this is our day-to-day work. On top of that, there is data. You know, we used to train on just like common crawl, basically. It's a big repository of the internet that people just scraped without regard of what. And some things came in, some didn't. It was a mess. So now, of course, every larger company has a team that tries to filter this and improve the quality.

9:51But it's a lot of work to really extract better data. Now synthetic data is becoming a thing. But when you generate synthetic data, it really matters how you do it, with what model, the whole problem. In the engineering aspects of everything, it's such a new domain that, you know, it was done somehow. It works. It's beautiful. But there is just so much to do better that people, I don't think, have any doubts that there is a lot there. And on top of that, there are the big things like multimodal. I mean, language models, they are now, as I'm sure you know, and most people realize, actually vision language models because they can also do audio.

10:36So they're multimodal models to some extent, but the multimodal part still lags behind the text part to a large extent. So that's one big area where obviously you need to do better. And it's not a huge secret how you can do better. You know, there are some methods that maybe will make it even amazingly better, but there are some very simple methods how you can do just better. But, you know, this maybe requires retraining your whole base model from scratch. And that takes a few months and it's a huge investment. So we need to organize it. So there is a lot of just work that will undoubtedly make things better.

11:16I think the big question that people have in their mind is how much better will it make them? So I'd love to do a little bit of a deep dive slash educational part on the whole reasoning model aspect, because as you just mentioned, since it's so new, some people truly understand how those work. Many people don't. At a very simplistic level, what is a reasoning model and how is that different from your sort of base LLM? So a reasoning model is like your base alarm, but before giving you the answer, it thinks what people call in the chain of thought, meaning it generates some tokens, some text that's meant not for you to read, but for the model to give you the better answer.

12:04And while it does this these days, it is also allowed to use tools. So it can, for example, in its thinking, so-called thinking process, go and browse the web to give you a better answer. So that's the superficial part of the thinking models. Now, the deep part is that you start treating this thinking process as part of the model, basically. So it's not something the model generates and it's an output for you. It's something you want to train, right? You want to tell the model you should think well, you should think so that the answer after this is good in whatever way. And this leads you to a very different way of training the model, because models were usually trained with just gradient descent, the way deep neural networks are trained, meaning you say predict the next word, and you do a gradient, you differentiate your function from the model.

13:00They're not fully differentiable, but you approximate it, and you train your weights to do that. And it was quite amazing that doing just that, you could make a chat. But with the reasoning model, you can do that because there is this reasoning part that you can't differentiate through that. So we train this with reinforcement learning. And reinforcement learning basically tells you, OK, there is just this reward and you need to do a bit of tries and reinforce, meaning push the model towards doing more of the things that lead to better answers. And this kind of training is a bit more like it has more restrictions than the training we used before.

13:36The training was before you took all of the internet, put it in, even if you didn't filter it very well, it would mostly work. Reinforcement learning, you need to be careful. You need to tune a lot of things, but you also need to prepare your data very carefully. So currently, and for at least the most basic ways we use it currently, it needs to be fairly verifiable. So there is an, is your answer correct or not? You prepare data for that. You can do that in mathematics, coding very well. You can do this in science to some extent, right? You can have test questions, correct or not. But, you know, if it comes to like writing poems, is this poem good or not?

14:11For now, the reasoning models are really shine in domains like science. And they've brought some improvements to non-science domains, but it's not quite as huge maybe yet as it could be. I mean, at least compared to mathematics and coding. Then there is the multimodal question. How do you do reasoning in multimodal? I think this is starting like, I saw some Gemini creating images in the reasoning part. It's quite exciting, but it's very, very, yeah. Pre-training and reinforcement learning part is particularly interesting from an educational point of view again, because it seems that people have come to the conclusion that there is the pre-training world and then the post-training world.

14:57And the post-training world is mostly reinforcement learning. but this idea that there is reinforcement learning in the pre-training I don't think is as understood by everyone. At the beginning of chat, let's say there was pre-training. People did not do RL, right? But then you couldn't really chat with it. So chat was RLHF applied to a pre-trained model. But the RLHF was a different kind of RL of search, right? It was very small and it was human preference that was telling you what is better. That's what the HF is, human feedback, right? You showed people like pairs of stuff. You learned a model that says, well, people seem to prefer this as an answer.

15:38You trained with that. It would very quickly, what we today say, hack this model. If you trained the RLHF too long, it would start giving things that satisfy this model that seems supposed to model human preferences. So it was a bit of a brittle technique, but it was a bit of an RL that was extremely crucial to making the models chat. These days, I think most people move towards this big RL. It's still not as big as pre-training in the scope, but it says you have a model that says whether this is correct or not. Or if it's a preference, it's a very strong model that analyzes things and says you should prefer that.

16:17And you have data that's restricted to a domain where you can do this well. And then you can also put some human preferences on top, but you make sure that you can run a little bit longer without making this whole grading fall apart. But again, this is our hour today. I do believe the hour of tomorrow will be broader. It will work on general data. And maybe then it will expand to domains that go beyond where it shines today. Now, will it shine there is a different question. Do you really need to think very much doing some of the things? Maybe not, but maybe yes. Maybe we do more thinking and reasoning than we kind of consciously call thinking.

17:02What would it take for RL to generalize? Is that better evaluations? Like you guys released GDP Val a few weeks or months ago to sort of measure performance against sort of broad economic sectors. Is that part of what the system needs? I think this is a small part of it. I think that's one part. But if you think of economic tasks, you know, making slides is important. They're following instructions, doing calculations. It's not math, but it's still very verifiable, right? What I'm thinking about is when you do pre-training, you take the internet and you just say, ask what's the next word. You know, you could think before you ask what's the next word, obviously.

17:51Now, you don't want to think before every word, probably, but I don't know if you ever looked at the training data for like a real pre-training run. Because I think people mostly don't realize how bad this is. hotels.com is a great website compared with the average chunk of 2000 words from the internet. It's a mess, right? And also a miracle that from this, the pre-training process gets you something reasonable. So you probably don't want, you know, imagine you have a hotel website telling, you know, it's a beautiful vacation. You don't necessarily want to have a very long chain of thought before that, right?

18:27If it was written by a person, there was probably some kind of thinking that went into it. Maybe not as elaborate as the math and coding thinking, but maybe there was something going on. So maybe you want a little bit of thought before at least some of the text. And that our models can't do very well yet. I think they're starting. There is a lot of generalization in this reasoning. If you learn to think for math, you will sometimes do some. Some strategies are, they transfer very much, like look up on the web and see what they say and use that information. So some of these things are very generic and they start to transfer.

19:06I feel like some are maybe not yet. Especially thinking in the visual domains is very under-trained, I believe. But we work, so we will try to push for more of that. Going back to chain of thought, how does that actually work? work, how does the model decide to create that chain of thought? And is what we see, so the little intermediary steps that we see on the screen as users, the chain of thought that's exposed to us, is that what's actually being processed by the model? Or is there a deeper, longer, broader chain of thought that happens behind the scenes? So in the current chat, GPT, you will see a summary of the chain of thought on the set.

19:54So there is another model that takes the full chain of thought and shows you a summary because the full ones are usually not very nice to read. They're more or less to say the same thing, just in more messy words. So it's better to have a more readable summary. When you start with a chain of thought, the first paper about chains of thought, you basically just ask the model, please think step by step. And it would think. So if you just pre-train a model on the Internet and ask it to think step by step, it will give you some chain of thought. Interesting and most important point is you don't stop there.

20:25You say, okay, so you start with some way of thinking, and then you say sometimes this leads to a correct answer and sometimes it leads to a wrong answer. So now I'm telling you I have some training examples. You will think 100 times and say 30 lead you to the correct answer, then I'll train you on these 30 examples and say this is the way you should be thinking. That's the reinforcement learning part of training. It changes dramatically how the models think. We see this for math and coding, but the big hope is it could also change how the models think for many other domains. Even for math and coding, you start seeing that the models start correcting their own mistakes.

21:04Earlier, if the model made a mistake, it generally just tells you what it did and insists that the mistake was right or something like that. With the thinking, it's like, oh, I often make mistakes, but I need to verify and correct myself to give the correct answer. So this just emerges from this reinforcement learning, which is beautiful, right? It's clearly a good thinking strategy to verify what you want to say. And if you think it may be an error, then think again. That's what the model learns on the most abstract level. Great. Thank you for this. All right. As a quick detour, and we'll go back to more Frontier AI topics, I'd love to talk a little bit about your story.

21:47story. I mean, you have the incredible distinction of having been at the forefront of this industry, both the Transformers paper, which was the birth of one paradigm. And now you're very much leading the charge on the reasoning model part, which is another paradigm. So this just incredible story. How did you become an AI researcher? I was a mathematician and a computer scientist, but in theoretical computer science. And that started in high school as a kid? Yes, I was definitely very into math in high school and into computers also later in high school. Yes, I did my studies in Poland. I went for a PhD in Germany.

22:27It was a theoretical computer science and mathematics PhD. So I very much am a mathematician. Yeah, I was always fascinated by, you know, how is this thinking going? What is intelligence? As a child, I always wanted to like emulate the brain. They thought, well, OK, maybe higher level explanations are more interesting. I did research in logic, but a little programming. But then there was this opportunity to join Google just as the deep learning was starting off. I already had my tenured position in France. And the French system has this beautiful thing that you can take a leave of 10 years. Yes.

23:03And you can still return anytime you want. So it's no risk. So at some point when you've solved AGI, you may return back to France and be a professor. Well, if you solve AGI, they may take you anyway. The nice part about the leave is that they will take you back even if you don't. But it's actually very important. A number of, I think there's a number of Nobel Prize winners who took this leave to just try something more risky. And, you know, sometimes it works. Sometimes it doesn't. There's a lot of luck in science and research. But it's very good to have this opportunity to take it. So I came to Google.

23:45And that was Google Brain at the time, you said, right? I came to Ray Kurzweil's group. He was my first manager. He interviewed me and was very inspiring. My first interview was to join the YouTube UI team. And I was like, OK, I'm not going. And then I had an interview with Ray. And I knew him, of course, from his books. And he's a very inspiring person. So I was like, OK, let's go. The team was separate from Google Brain at that time. Then I moved to Google Brain, worked with Ilya Skavar, another very inspiring person. There's an amazing number of great people in AI and in the Bay in general.

24:20I have to ask you at this point about the Transformer paper story, how it all came about. The eight of you, right? Seven or eight of you? How did you all get together? Well, we never got together. You never got together. Okay. I recently got a photo on Twitter of a photo session of all eight of us. And it was saying it was fake, but I knew it was fake because I don't think all eight of us were ever in the same physical room. These ideas developed from many sites before and after, like Jakob and Delia Poloschukin worked on attention, like self-attention. Of course, attention was there from the encoder-decoder side.

25:00And maybe one minute for the broad public on what attention actually means, since it's such a fundamental concept. So attention is the mechanism that tells the model, as you're doing the next thing, look into your past and find the most similar things that you see in the past to what you are seeing right now. It came from the machine translation times where people wanted to align words in one language with words in another. They were like, OK, so this word, where in this previous sentence would it be? It's an analog of alignment for deep learning. It's now called attention in AI. Just says, you know, think of like what comes to your mind as you are here now in this environment, what things from the past are similar to.

25:50And this mechanism was already used in deep learning translation before, but it was used. There was like one encoder model and the decoder would be looking at the encoder, but never at its own states. The main novelty of Transformer was self-attention, but Transformer is more than just this idea. I think that's important. I think that's the beauty of these eight weird people that somehow came together, even though not physically to do it, is that we all approached it from different sides. So there was people working on the attention idea. There is the, you know, you need to put this in the network that needs to have a lot of knowledge.

26:30So there is the feed forward layer that expands and then contracts. So Noam was working on this and nowadays used mixtures of experts, which actually came before the transformer. So how do you store knowledge in neural networks is another important question. And it's part of this model, too. And then, you know, in deep learning, people laugh that ideas are cheap. Making them work is the hard part. So how do you write the systems and the code and the baselines to actually make this train? And this is funny to say now, because nowadays you can take any deep learning framework and say, you know, X equals transformer, X train, and it will basically work.

27:11But back then, it totally did not. So you need things like learning rate warmup or tweaks to the optimizer that just work. And I did a lot of coding at that time, was working on TensorFlow and parts of the framework. And I remember distinctly that people were like, so you want to use the same model for a few different tasks? Like, why do you even do that? Like, if you have a different task, like if you do translation, you train one model. If you do parsing, you train another. If you do image recognition, you train a third. Like, you'd never train the same model for three different tasks. Why do you even write, like, APIs to do multiple tasks on one model?

Read the full transcript

27:50And I was like, no, no, we're going to do all tasks in one model. And people were like, no, no. So there was a lot of pushback against the idea? Not against the idea. Google was also an amazing place at that time that they would very happily let you work on whatever you wanted. But I don't think there was widespread belief in doing multiple tasks with the same model, not to mention, you know, this idea that you, I still find this idea that you take basically the same model as Transformer. Like now there is a bunch of changes to it, but you could in principle take the same architecture as the decoder from the paper, train it on all of the Internet, and it will basically start chatting with you.

28:30That would have back then definitely sounded as a worthy dream we maybe had as a dream, but not a reality that you expect five years later. It's very lucky that it actually works so well, right? Talk about the transition from Google to OpenAI and perhaps how those two cultures are different. So Ilya Suzkever was my manager at Brain. Then he went on to found OpenAI. He asked me a number of times whether I would like to join in the years. I found it a little bit too edgy at the time. Then Transformers came, so we had a lot of work with that. And then COVID came. And COVID was a tough time for the whole world, right?

29:13But Google totally closed. Google was reopening extremely slowly. So one part of me was I find it very hard to do remote work. I much prefer to work with people directly. That was one reason. But the other was also Google Brain, when I joined it, was a few dozen people, maybe 40, something like that. When I left, it was 4 ,000, 3 ,000 people spread across multiple offices. It's very different to work in a small group and to work in a huge company. So with all this, Ilya was like, you know, OpenAI, though, is in a much stabler state. We're doing language models. You know something about this that may look like a good match.

29:52And I was like, OK, let me try. I've never worked in any company other than Google before, other than the university. So it was quite a change to the small startup group. But I like working in smaller groups. It has its pleasure. It has a little bit of a different intensity sometimes. In general, I found it very nice. On the other hand, Google has merged, made the Gemini, and I hear it's also a very nice place. I think in general, the tech labs are more similar to each other than people think. There are some differences, but I think if I look at it from the world, you know, from the university in France, the difference between this university and any of the tech labs is much larger than between one lab or the other.

30:43How are the research teams organized within OpenAI? They're organized. They're not very organized. I mean, we do organize them. Some people have managers and sometimes talk to them. No, but mostly people find like projects, there are things to do, right? Like improve your multimodal models, improve your reasoning, improve your pre-training, improve whatever this part of the infrastructure. People work on it. You know, as we go through these parts, right, there is infrastructure, pre-training, reasoning. I think the parts are the same for most of the labs. So there will be teams doing these things.

31:27And then sometimes people change teams. Sometimes new things emerge. There's always some smaller teams doing like more adventurous stuff like diffusion models at times. Then, you know, some of the more adventurous stuff like video models gets big and then maybe they need to grow. Do people compete for GPU access? I don't think it's so much people that compete. I think it's more projects that compete for GPU access. There's definitely some of that. On the other hand, like on the big picture of GPU access, a lot of this is just determined by how the technology works, right? Currently, pre-training just uses the most GPUs of all the parts.

32:13So it needs the most GPUs, right? RL is growing in the use. Now, video models, of course, use a lot of GPUs too. So you need to split them like this. Then, of course, people will be, oh, but my thing would be so much better if I had more GPUs. And I've certainly said that a number of times, too. So then you kind of push it, you know, I really need more. And then some people may say, well, but, you know, there's only so much. There is never enough GPUs for everyone. So there is some part of the competition, but the big part is just decided by the fact how the technology works currently. Great. What is next for pre-training?

32:53We talked about data. We talked about engineering, a big GPU compute aspect to this. What happens to pre-training in the next year or two? Pre-training, as I said, I think it has reached this upper level of the S-curve in terms of science, but it can scale smoothly. Meaning if you put more compute, you will get better losses if you do things right, which is extremely hard. And that's valuable. You don't get the same payoff as pushing HRL, but it generally just makes the model more capable. And that's certainly something you want to do. I think what people underestimate a little bit in the big narrative is, you know, OpenAI three, four years ago, I joined even before that, was a small research lab with a product called API.

33:50But, you know, it was not such a big, there was no GPU constraint on the product side, for example. All GPUs were just used for training. So it was very easy as a decision for the people to say, you know, we're going to train GPT-4. This will be the smartest and largest model ever. And what do we care about small models? I mean, we care of them as to make like to debug the training of the big model, but that's it. So GPT-4 was the smartest model and it was great, right? But then it turned out, oh, there is this chat and now we have a billion users. And, you know, people want to chat with it a lot every day and you need GPUs.

34:30So you train the next like huge model and it turns out you cannot satisfy this. Like people will not want to pay you enough to chat with the bigger model. So you just economically need the smaller model. And this happened, of course, to all the labs, because like the moment the economy arrived and it became a product, you had to start thinking about price much more carefully. than before. So I think this caused the fact that instead of just training the largest and largest thing you can for the money you have, we said, well, no, we're going to train the same thing, but same quality, but smaller, cheaper.

35:10The pressure to give the same quality for less money is very large. In some sense, as a researcher, it almost makes me a little sad. I have a big laugh for these huge models that people say human brain has 100 trillion synapses or orders of magnitude, of course, are not that exactly calculated, but our models don't have 100 trillion parameters yet. So maybe we should reach it. I would certainly love it, but then you need to pay for it. So I think this may be why people kind of think that pre-training has paused because a lot of effort went into training smaller models. Now on the side, people kind of rediscovered how amazing distillation is.

35:51Distillation means you can train a big model and then put the same knowledge from the big model as a teacher to the little model. People knew about distillation. It's a paper for a long time ago. But somehow, at least for OpenAI, I think maybe it was more in Google's DNA when Aureole's there. But people kind of rediscovered how important that is for the economics. But now it also means that, oh, training this huge model is actually good because you distill all the little ones from it. So now maybe there is a bit more of a return to like, it's also a matter, you know, once you realize you have the billion users and you need the GPUs, you need to invest into them.

36:31And of course, everyone sees this. There's a huge investment, but the GPUs are not online yet. So when they come back online, and I think this may play into this, what people call resurgence of pre-training. We both understand that you can distill this amazing big model and there is now enough GPUs to actually train it. So it's resurging. But all of this fundamentally happens on the same scaling curve, right? It's not like we didn't know that you could do this. It's more like the different requirements of different months have changed, sometimes changed the priorities. But I think it's good to step back from it and think of the big picture, which is that pre-training has always worked.

37:15And the beautiful thing is it even stacks with RL. So if you run this thinking RL process on top of a better model, it works even better than if you run it on top of a smaller model. One question that I find fascinating as I hear you speak and the evolution of the modern AI systems has been this combination of LLM plus RL plus a lot of things going on. And it used to be at some point, and maybe that was back in the deep learning days, that people would routinely say that they understood how AI worked at a micro level, like the matrix multiplication aspect, but didn't fully understand once you had everything together, what really, really happened at the end of the day in the model.

38:03And I know there's been tons of work done on interpretability over the last couple of years in particular, but particularly for those very complex systems, is it increasingly clear what the models do or is there some element of black box that persists? I would say both. There is a huge progress in understanding models. Fundamentally, I mean, think of the model that is chat. It talks to a billion people about all kinds of topics. It gets this knowledge from reading all of the internet. Obviously, you cannot identify, like, I cannot understand what's going on in there. I don't know the whole internet.

38:39What we can identify is there was a beautiful paper just, I think, last week from OpenAI about if you tell the model that lots of its weights should be zeros, it should be very sparse, then you can really trace when it's thinking about one particular thing, then you can trace what it's actually doing. So if you say limit ourselves to this and to really study this inside a model, then you can get a lot of understanding. And there's circuits in the models. Antropic had great papers on that. So the understanding of what the models are doing on a higher level has progressed a lot, but then it's still an understanding of what smaller models do, not the biggest ones.

39:19But it's not so much that these patterns don't apply to bigger models. They do. It's just the bigger models just do so many things at the same time that there is some limit to what you can understand. But I think this limit is a bit more fundamental than people think. It's like every very complex system. You can only understand so many things and then you don't, right? Thank you for all this. I'd love now to talk about 5.1 and do a little bit of a deep dive on all the latest stuff that you guys have released in the last couple of weeks, which has been very impressive, in particular as a user. I think that the 5.1 moniker doesn't do justice to the evolution between 5.1 and 5.

40:03It feels like a much larger improvement than the number would indicate from, again, my perspective as a user. Walk us maybe through the evolution of from GPT-4 to 5 to 5.1. What has actually changed? That's a very tough question. I think less than you think. I think it's, no, I mean, from GPT-4 to 5, I think the biggest thing that changed is reasoning, meaning RL and synthetic data. As I told you, the pre-training part in that timeframe was mostly about making things cheaper, not making things better. So, of course, the price has changed dramatically too, right? A thousand times, I think, are some of these orders of magnitude.

40:48The main improvements from four to five is adding reasoning with reinforcement learning and this allowed to generate synthetic data, which also improves the model. So that's the big picture. In addition to that, ChatGPT is now a product used by a lot of people. So the post-training team has learned a tremendous number of lessons and it's added, you know, things clearly experimented, wanted the model to be very nice to you. and turned out to be too nice. Then now when a lot of people use it, you need to be really careful about safety, right? There may be people that are in distress using the model.

41:26The model needs to do something reasonable in these cases. It was not trained for it before. Now it is, and it makes the model much better. But, you know, in the same time, you don't want to refuse to answer any question that has any sign of anything. So as you work on these things, you make the model much better in use, not just for the people in distress, but for everyone who wants questions answered, but the answers to be reasonable. And, you know, there was these things called hallucinations. It's still with us to some extent, but dramatically less than two years ago. Some part of that is because reinforcement learning can now use tools and gather data.

42:06And it also encourages the model to, like, know, you know, verify what it's doing. so that's an emergent thing from this reinforcement learning of reasoning but also you just add data because you realize sometimes the model should say I don't know so you add this to the post-training data you say like well we really need to give it a thought how the model should answer people in various situations. The price to find point one is mostly this kind like it's mostly a post-training improvement. Yes, to double click on this, because it is super interesting. So indeed, as part of 5.1, there's the ability to choose different kind of styles from nerdy to professional.

42:52And that's, I guess, in reaction to the fact that some people were missing the sycophantic aspect of earlier models when chat GPT, when GPT-5 came out. And so adding more tones, that's all post-training stuff. So you tell the model, those are examples of like how you should be responding, which is more like a sort of like super fast tuning kind of a paradigm. Or is that RL like right or wrong with rewards? How does that work? I don't work on post-training and it certainly has a lot of quirks. But I think the main part is indeed RL, where you say, OK, is this response cynical? Is this response like that?

43:35and you say, okay, if you were told to be cynical, this is how you should respond. If you were told to be funny, try this one. So I do think the RL is a big part of it. In between models or different versions of the models, are the releases aligned with pre-training efforts or sometimes you have like one big pre-training effort and like several models that come out based on that? There used to be a time, not that long ago, we have a year distant past where where where the models were did have an alignment with technical stuff right so they would align whether with either with rl runs or pre-training runs that's why you had a beautiful model called 4.0 which was aligned with a pre-training run which was obviously worse than the o3 aligned with an rl run that was the follow-up to a one naturally because you couldn't use the name O2.

44:31But it was slightly better than the O4 Mini because that one was Mini. And, you know, we had this beautiful model picker and people kind of thought this was not the best naming for some whatever reason. So, no, I mean, it was fairly obvious that this was very confusing, right? So now the naming is by capability, right? GPT-5 is a capable model. 5.1 is a more capable model. Mini is the smaller model that's slightly less capable, but faster and cheaper. And the thinking models are the ones that do more research. In that sense, the naming is detached from any technical, in particular, 5.1, maybe just a pre-training, sorry, post-training thing, but maybe 5.2 is the newly pre-trained model or maybe not, but the naming is detached from the technology.

45:24which also gives some, you know, as OpenAI has grown, there is a number of projects, right? There is RL and pre-training and maybe, you know, something just to make slides better or whatnot. And with distillation, you have the ability to put a number of projects into one model. It's kind of nice that you don't need to wait on all of them to complete at the same time. You can try to periodically put this together, actually make sure that as a product, It's nice to the users and good and do this separately from, you know, waiting on the new full pre-training run that takes months and so on. So I feel like even though a little tear in my eye goes for the times where it was that pre-trained model number that was the number.

46:07As it's a product serving a billion users, it's maybe inevitable that you should name it by what the user should expect from it rather than. In 5.1, you have additional granularity in terms of telling the model how long it should think. By default, how does the model decide how long it should think? So the model sees the task. It will decide on its own a little bit how long it should think. But you can give it an additional, it's trained with an additional information that can think even harder and then it will think longer. So you have now the ability to steer that. But I still think it is important to realize.

46:48So this is the fundamental change that came with reasoning models, that using more tokens to think increases your capability. And it increases it, given the computation, way faster than pre-training. So if you give GPT-5 the ability to think for long, it can solve tasks that are, You know, we had these gold medal at Mathematical Olympiad and Computer Science Olympiad. So they're amazing abilities. At the same time, the fundamental training method of reasoning is very limited to science data. So it's not as broad as the pre-training, which I think like pre-training models felt kind of almost uniformly good or bad at things.

47:38I mean, this was still not uniform because it's not like teaching humans, right? But the reasoning models are even more, people call it jagged, right? They have amazing abilities somewhere and then close by, not so much. And that can be very confusing. It's something I always love that. It's weird because you can say the model is amazing at Mathematical Olympiad. But at the same time, I have a math book for, I have a first grader daughter in the first grade. She's five years old. I took one exercise from this math book and none of the frontier models is able to solve it. And you would be able to solve it in 10 seconds.

48:17So that's something to keep in mind. Models are both amazing and they're tasks that they cannot do very well. I can show you this as an example. I think it's quite interesting to keep in mind. Let me start with Gemini 3, just to blame the competitors. Yes, please. So it has, you see, two groups of dots on both sides. And the question is, is the number of dots even or odd? And if you look at it, you see, oh, they're like two identical things. So that would be even. That's what the five-year-old is supposed to learn. But there is one dot that's shared. So now that must be odd. For this simple one, which has like, you know, I don't know, 20 dots or so, Gemini 3 actually does it, right?

48:58It finds out that it's an even number of dots and it says that and that's great. And then you have another puzzle, which is very similar, except now there are two mountains of dots. And there's also one dot shared at the bottom now. And right in context, right after that, you ask, OK, how about this one? And then it does some thinking and it just totally misses that there is a shared dot and it says the number is even. And it's like, in context where you've seen this first example, how would you ever miss that? You know, and here is the same, the exact same prompt for GPT 5.1 thinking. And it also solves the first.

49:37It sees the dot. It says it's odd. And then it sees the mountains. And somehow it doesn't see the dot. And it says it's even. The nice thing is, if you let it think longer, or if you just let it think again, It will see it. So if you use GPT-5 Pro, it takes 15 minutes. So, you know, this is the human five-year-old takes 15 seconds. The GPT-5 Pro will run Python code to extract these dots from an image and then it will count them in a loop. So that's not quite... And why is that? What trips up the model? I think this is mostly multimodal part. The models are just, they're starting. Like you see the first example they managed.

50:26So they've clearly made some progress, but they have not yet learned to do good reasoning in multimodal domains. And they have not yet learned to use one reasoning in context to do the next reasoning. What is written in context is, you know, learning in context happens, but learning from reasoning in context is still not very strong. All of these, though, are things that are very well known. And the models are just not trained enough to do this. It's just something we know we need to add into training. So I think these are things that will generally improve. I do think there is a deeper question whether, so, you know, like multimodal will improve.

51:09This will improve. Like we keep finding these examples. So as the frontier will move, it will certainly move forward. Some things will smoothen, but the question is, will it still be just other things that you don't need to, you know, teach the human like every, you know, okay, now you know how to use a spoon and a fork. But now if the fork has four instead of three ends, then you need to learn a new. That would be a failure of machine learning. You know, I am fascinated by generalization. I think that's the most important topic. I always thought this was the key topic in machine learning in general and in understanding intelligence.

51:47Pre-training is a little different, right? Because it increases the data together with your increase in model size. So it doesn't necessarily increase generalization. It just uses more knowledge. I do believe that reasoning actually increases generalization. But now we train it on such narrow domains that it may still be to see. But I think the big question in all of AI is, is reasoning enough to increase generalization or do you need like more general methods? I think the first step is to make reasoning more general, as we talked before. That's my passion. That's also what I work on. There is still something there, right?

52:20We push the models. They learn. They learn things that are around what we teach them. They still have limitations because they don't live in the physical world, because they're not very good at multimodal, because reasoning is very young and there's a lot of bugs in how we do it yet. But once we fix that, there will be this big question. Is that enough or is there like something other big to make models generalize better? So we don't need to, you know, teach it every particular thing in the training data that it just learns and generalizes. I think that's the most fascinating question. But I also think a good way to approach a question like that is to first solve everything that leads up to it.

52:59You know, you cannot know whether there is a wall or not until you come close to it, because otherwise, you know, we AI is moving very fast. Someone said it's like driving fast in a fog. You never know how far or close you are. So we're moving. We are learning a lot. And does that mean, so that central question of basically learning with very little data, the way a child would, and the fact that a child is able to do things that even the most powerful model cannot do. So this, as you said, to unpack this, making progress on reasoning and showing how far we can get into generalization with reasoning.

53:38And then the separate question is, as you said, whether we need an entire different architecture. And that's where we get into, for example, Yann LeCamp's work. Do you see promising, fundamental architectural changes outside of transformers that have caught your attention and feel like they could be a serious path to explore in the future? I think there is a lot of beautiful work that people are trying out. You know, the ARC challenges inspired one set of people. There are models now that are very small and solve them very well. But with methods that I'm not sure are actually general, you need to see.

54:18Jan Lecun has been pushing for other methods. So I feel like his approach is more towards the multimodal part. But maybe if you solve multimodal right, maybe if you do JEPA, it also helps your other understanding. There is a lot of people pushing fundamental science. It's maybe not so much in the news as the things that push. But whatever you do, it will probably run on some GPU. If you get a trillion dollar of new GPUs, the old GPUs will be much easier to get also. So I think this growth in LLMAI on the more traditional side is also helping people to have an easier time to run more experimental research projects on various things.

55:02So I think there is a lot of exploration, a lot of ideas. It's still a little hard to implement them at a higher scale. The engineering part is the biggest bottleneck. I mean, GPUs are a bottleneck, too, when you scale really up. But implementing something that's larger than one machine, it's an experimental research project, so you don't have a team to do that. I think that's still harder than it should be. But, you know, codecs may get their code. This is the thing where AI researchers have great hope to help themselves and also other researchers. is that if you could just say, hey, Codex, this is the idea, and it's fairly clear what I'm saying, please just implement it so it runs fast on this eight machine setup or 100 machine setup.

55:50That would be amazing. It's not quite capable of doing that yet, but it's capable of doing this more and more. I think that's what OpenAI says. They say we'd like an AI intern by the end of next year. That's how I understand this. Can someone help us? Is part of Codex's, the path for Codex to be able to do some of this, does that revolve around how long it can run? Context behind the question being that, again, like two days ago as we recorded this, you guys released GPT 5.1. Codex Max, described as a frontier-agentic coding model, trained on real-world software engineering tasks, designed for long-running workflows and using compaction to operate across multiple context windows in millions of tokens.

56:39So I'd be interested in unpacking some of this. What does that mean to run for a very long time? Is that an engineering problem or a model problem? And then maybe a word on compaction. So it is both an engineering and model problem. You know, you want to do some engineering tasks, like write a, you have some machine learning idea. You want codecs to implement it for you, test it on some simple thing, find the bugs. So it needs to run this thing. This is not something you would do in an hour, right? It's something you'd spend a week on. So the model needs to spend a considerable amount of time because it needs to run things, wait for the results, then fix them.

57:18The model is not like it's going to come up with the correct code out from thin air, right? It's just like us. It needs to go through the process. And oftentimes in the process, since it was not trained on anything very long in its training, or maybe very few, but certainly nothing that went on for a week, it can get lost. It can start doing loops or doing something weird. That's, of course, not something you want. So we try to train in a way that makes it not happen. But it does. So, you know, how can you make the model actually run a process that requires this larger feedback loop without tripping up?

58:00And the other thing is transformers have this thing called context. So they remember all the things that they have seen in the current run. And that can just exceed the memory available for your run. And the attention matrices are n by n, where n is this length. so they can get huge. So instead of keeping everything, you say, well, I'm going to just ask the model on the side to summarize the most important thing from the past, put it in context and you and forget some part of it, right? So it's a very basic form of forgetting the compaction, right? And that allows you to run for much longer if you do this repeatedly.

58:37But again, you need to train the model to do that. And when you do it, it works to some extent. I don't think it works well enough to replace an AI researcher yet. It made a fair bit of progress. I think another part of progress that's a little understated on the research side, but is very important, is allowing the model to connect to all of these things. So models now use tools like Web Search and Python routinely, but to run on a GPU, to have access to a cluster. It's hard to train models with that because then you need to dedicate for the model to use. And that has security problems. And this thing, like how do models connect with the external world?

59:15It's a fundamentally very hard problem because, you know, when you connect in an unlimited way, you can break things in the real world. We don't like models to break things for us. So that's the part where people work a lot. It overlaps with security, right? You need to have very good security to allow models to go on and train on the things they need to train. One theme that people like me, VCs and founders and startups think about a lot as we see all the progress at OpenAI is as the models keep getting more general with more authentic capabilities, the ability to run for a very long time, you know, going to areas like science and math.

59:56And, you know, recently it was reported that there were some investment bankers hired to help improve the model's capability to do grunt investment banking work. All of that taken into account, is there a world where basically models or maybe just one model does everything? And I don't know if that's AGI, let's not necessarily go into that debate. But what's left for people that build products that sit on top of models? I just showed you a five-year-old exercise that the model doesn't do. I think we need to keep that in mind. So you're saying there's hope. There is hope the next model will do it.

1:00:36That hope, okay. Well, for me, yes. I still think we have some way on the models. Progress has been rapid. So there is good hope there will be less and less of this. But on the other hand, for now, you don't need to do a deep search to find things where you'd really want a human to do that task because the model's not super good. On the other hand, Transformer paper started with translation. I recently went to a translation industry conference. The translation industry has grown considerably since then. It has not shrunk. There's more translations to be done. Translators are paid more. questions, why would you even want a translator if the model's so good in most of the cases?

1:01:19The answer is sometimes, imagine you do a listing for a newspaper, but in a language you don't know. And GPT-5 will almost certainly translate it correctly for you. If it's into Spanish or French or any high resource language, would you still publish it without having a human who speaks that language look at it? Would you publish it if it's a UI of chat GPT that a billion people are going to see. It's a question of trust. Probably right. But if you have a million users, a billion users, maybe you will pay the$50 for someone to just take a look over it before you translate it. So this is in an industry that fundamentally is totally automated, right?

1:01:57There is still the question of trust. And I think that's a question we will grapple with for a long time, there are also just things you want a person to do. Like, I don't think we will have no things to do. But that doesn't mean that some things we do may not dramatically change. And that, you know, that can be very painful for people who do these things. And so this is a serious topic that happy people are engaging with. But I don't think like there will be this global lack of anything for people to do. And maybe as a last question to help us get a sense for what people at the frontier of AI are currently thinking about or working on.

1:02:38Some of the topics that one may see are things like continual learning, world models, robotics, embedded intelligence. So what do you personally find very interesting in addition to what you mentioned upfront multimodal? But what do you personally find really interesting as a research area? Well, I find this general data reinforcement learning is my pet peeve and what I work on, luckily. For example, robotics is probably just an illustration that we are not doing that well in multimodal and that we're not doing that well in general. RL in general reasoning yet. The moment we do really well in multimodal and we manage to generalize reasoning to the physical domains that the robot needs, I think it will see amazing progress.

1:03:31When it does, I have a feeling given that a lot of companies are launching hardware that's kind of teleoperated or glove operated or something. So my suspicion is by the moment we make this progress, which maybe it will be next year, maybe it will be in a few more years. But the hardware will maybe be there by then. And having a robot in the home may be like a big visible change. More visible than, you know, chat. I mean, given how quickly we got used to the self-driving cars in San Francisco, maybe it will be only visible for like the first two days and then be like, yeah, sure, the robot's there.

1:04:12It's always been cleaning since I can remember like the last three months. And it's stunning to me how quickly we get used to these things, right? The self-driving cars in San Francisco are something people got used to so quickly. So maybe this will happen for robots too. Nevertheless, I do think it will be quite dramatic in our perception of the world when it happens. Hardware is hard though, right? Robots may have accidents in the house. You need to be very careful. So maybe it will take longer to deploy them and actually make it a scalable business. We'll see. It is amazing that, you know, we are at this point where we can start thinking like, yes, maybe that will come soon.

1:04:54Lukasz, it's been absolutely wonderful. Thank you so much for spending time with us today. Thank you very much, Matt. Thank you for the invitation. Great to talk to you. Hi, it's Matt Turk again. Thanks for listening to this episode of the MAD podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests. Thanks and see you at the next episode.

From the publisher

We’re told that AI progress is slowing down, that pre-training has hit a wall, that scaling laws are running out of road. Yet we’re releasing this episode in the middle of a wild couple of weeks that saw GPT-5.1, GPT-5.1 Codex Max, fresh reasoning modes and long-running agents ship from OpenAI — on top of a flood of new frontier models elsewhere. To make sense of what’s actually happening at the edge of the field, I sat down with someone who has literally helped define both of the major AI paradigms of our time.


Łukasz Kaiser is one of the co-authors of “Attention Is All You Need,” the paper that introduced the Transformer architecture behind modern LLMs, and is now a leading research scientist at OpenAI working on reasoning models like those behind GPT-5.1. In this conversation, he explains why AI progress still looks like a smooth exponential curve from inside the labs, why pre-training is very much alive even as reinforcement-learning-based reasoning models take over the spotlight, how chain-of-thought actually works under the hood, and what it really means to “train the thinking process” with RL on verifiable domains like math, code and science. We talk about the messy reality of low-hanging fruit in engineering and data, the economics of GPUs and distillation, interpretability work on circuits and sparsity, and why the best frontier models can still be stumped by a logic puzzle from his five-year-old’s math book.


We also go deep into Łukasz’s personal journey — from logic and games in Poland and France, to Ray Kurzweil’s team, Google Brain and the inside story of the Transformer, to joining OpenAI and helping drive the shift from chatbots to genuine reasoning engines. Along the way we cover GPT-4 → GPT-5 → GPT-5.1, post-training and tone, GPT-5.1 Codex Max and long-running coding agents with compaction, alternative architectures beyond Transformers, whether foundation models will “eat” most agents and applications, what the translation industry can teach us about trust and human-in-the-loop, and why he thinks generalization, multimodal reasoning and robots in the home are where some of the most interesting challenges still lie.


OpenAI

Website - https://openai.com

X/Twitter - https://x.com/OpenAI


Łukasz Kaiser

LinkedIn - https://www.linkedin.com/in/lukaszkaiser/

X/Twitter - https://x.com/lukaszkaiser


FIRSTMARK

Website - https://firstmark.com

X/Twitter - https://twitter.com/FirstMarkCap


Matt Turck (Managing Director)

Blog - https://mattturck.com

LinkedIn - https://www.linkedin.com/in/turck/

X/Twitter - https://twitter.com/mattturck


(00:00) – Cold open and intro

(01:29) – “AI slowdown” vs a wild week of new frontier models

(08:03) – Low-hanging fruit: infra, RL training and better data

(11:39) – What is a reasoning model, in plain language?

(17:02) – Chain-of-thought and training the thinking process with RL

(21:39) – Łukasz’s path: from logic and France to Google and Kurzweil

(24:20) – Inside the Transformer story and what “attention” really means

(28:42) – From Google Brain to OpenAI: culture, scale and GPUs

(32:49) – What’s next for pre-training, GPUs and distillation

(37:29) – Can we still understand these models? Circuits, sparsity and black boxes

(39:42) – GPT-4 → GPT-5 → GPT-5.1: what actually changed

(42:40) – Post-training, safety and teaching GPT-5.1 different tones

(46:16) – How long should GPT-5.1 think? Reasoning tokens and jagged abilities

(47:43) – The five-year-old’s dot puzzle that still breaks frontier models

(52:22) – Generalization, child-like learning and whether reasoning is enough

(53:48) – Beyond Transformers: ARC, LeCun’s ideas and multimodal bottlenecks

(56:10) – GPT-5.1 Codex Max, long-running agents and compaction

(1:00:06) – Will foundation models eat most apps? The translation analogy and trust

(1:02:34) – What still needs to be solved, and where AI might go next

More from The MAD Podcast with Matt Turck

All 44 episodes
What’s Next for AI? OpenAI’s Łukasz Kaiser (Transformer Co-Author)The MAD Podcast with Matt Turck · 1 h 5 min
Listen in VO