How GPT-5 Thinks — OpenAI VP of Research Jerry Tworek

16 Oct 2025 · 1 h 16 min · 30 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

How OpenAI’s reasoning models work, what “thinking”/chain-of-thought means, how long models should think, and how OpenAI’s research and RL program evolved from O1 to O3 to GPT-5.

Guest

Jerry Tworek, VP of Research at OpenAI; member of the Metis list of top AI researchers. Background: Grew up in Poland; studied mathematics at University of Warsaw; briefly pursued trading (JP Morgan internship; later hedge fund attempts in London and Amsterdam). Joined OpenAI in 2019, working on reinforcement-learning robotics (dexterous manipulation; neural-network control solving a Rubik’s Cube).

Key claims

Reasoning is getting to an answer you don’t yet know, akin to search but not naive search; chain-of-thought is verbalized computation learned from training on human thinking. Longer thinking usually improves results, but there’s a quality-vs-latency trade-off (“cheap, fast, or good”). O1 was mainly a puzzle demo; O3 was a “tectonic shift” via tool use and perseverance; GPT-5 is framed as O3.1, with further jumps aimed at thinking longer and using more systems/sources.

Notable examples

“Let’s solve it step by step” prompting; Rubik’s Cube dexterous manipulation; RLHF via human thumbs-up/thumbs-down preferences; RL scaling compared to steel vs semiconductors; GRPO (DeepSeek open-sourced) accelerating US labs’ reasoning-model research.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Understanding Reasoning in AI

0:54 to 2:09

Jerry explains what reasoning means in the context of AI and ChatGPT.

“Please enjoy this great conversation with Jerry.”

Chain of Thought Process in Language Models

2:09 to 4:34

Discussion on how language models use chain of thought for reasoning.

“Is that a logical tree and it eliminates option after option?”

User Experience and Model Thinking Time

4:34 to 7:22

Exploration of how models balance quality and speed based on user expectations.

“They will try to predict the next token, but they fail.”

Evolution of Reasoning Models at OpenAI

7:22 to 10:28

Jerry discusses the evolution of reasoning across different models like O1, O3, and GPT-5.

“So there was 01, then there was 03, then most recently HGP5.”

Jerry Tworek's Journey into AI

10:28 to 14:00

Jerry shares his personal background and journey into the field of AI research.

“But before we do that, let's talk about your journey.”

Journey from Academia to Trading

14:00 to 16:58

Discover how Jerry transitioned from academia to a career in trading using mathematics.

“It felt a little bit too rigid, a little bit too structured in a way that I didn't know if I will feel good.”

Discovery of AI and Reinforcement Learning

16:58 to 19:34

Learn about Jerry's introduction to AI and his fascination with reinforcement learning.

“And the depth of what you can go of trying to understand it model is very deep.”

Joining OpenAI and Early Projects

19:34 to 21:07

Explore Jerry's journey to OpenAI and his initial projects focusing on reinforcement learning.

“Like, you know, Google search, where are places where you can do reinforcement learning in this world?”

Robotics Focus at OpenAI

21:07 to 22:47

Understand Jerry's work on robotics at OpenAI, including dexterous manipulation projects.

“the world, what scaling up reinforcement learning can do.”

A Day in the Life of a Researcher

22:47 to 24:03

Get insights into Jerry's daily responsibilities and collaborative work at OpenAI.

“So fast forward to today, still in the same vein of like the behind the scenes of, you know, you all had OpenAI and sort of life there.”
Show all 30 chapters

Structuring Research Projects at OpenAI

24:03 to 26:52

Learn about how research priorities are determined and organized at OpenAI.

“Do people suggest ideas and others vet them?”

Collaboration vs. IP Protection

26:52 to 28:00

Explore the balance between collaboration and protecting intellectual property at OpenAI.

“Because I would imagine putting myself in your shoes, OpenAI shoes, there is probably a tension between wanting to be collaborative in general.”

Collaboration and Culture at OpenAI

28:00 to 30:02

Explore the collaborative culture within OpenAI and how it fuels research productivity.

“Those are humans and those are human things.”

Rapid Model Development at OpenAI

30:02 to 32:38

Learn about the factors enabling OpenAI's rapid iterations and releases of AI models.

“Why are you guys able to ship so quickly?”

Reinforcement Learning Explained

32:38 to 35:32

An accessible explanation of reinforcement learning using dog training as an analogy.

“So is the right way to think about modern AI systems at OpenAI.”

Understanding Reinforcement Learning Dynamics

35:32 to 40:00

Delve deeper into the mechanics of reinforcement learning and its components.

“Yeah, I usually the metaphor and the analogy I have to reinforcement learning is like training a dog.”

Evolution of Reinforcement Learning

40:00 to 42:00

Discuss the historical context and evolution of reinforcement learning techniques in AI.

“Mostly how does modern RL differ from historical RL?”

Early Challenges in Reinforcement Learning

42:00 to 43:35

Discusses the initial struggles and insights in applying RL to language models.

“But it was in some way at that end of doing RL without pre-training.”

The Breakthrough of RLHF in GPT-4

43:35 to 45:24

Explains how the RLHF technique improved GPT-4's performance significantly.

“early trials and errors weren't super successful at the moment when we when we trained GPT-4 And there was an interesting moment where we trained GPT-4 and everyone today thinks, oh, GPT-4 is such a great model.”

The Role of Human Feedback in AI Training

45:24 to 47:25

Describes how human feedback was integrated into the training process of AI models.

“together as a package like you know deliver that gpt moment to the world that everyone sees And as much as it is a big success of pre-training, it actually also was pretty big success of RL in the form of RLHF.”

Understanding Unsupervised Learning

47:25 to 48:23

Clarifies the differences between unsupervised and supervised learning in AI.

“The interesting bit about data labeling industry, I'm not sure how much we want to go on that tangent.”

The Complexity of Scaling Reinforcement Learning

48:23 to 53:10

Explores the challenges and intricacies involved in scaling RL.

“But why pre-training is called unsupervised because in some definition of it, you don't need any extra labels to the data that you feed into the model.”

Future of Agentic AI and Problem Solving

53:10 to 56:00

Discusses the potential of AI to enhance automation and problem-solving capabilities.

“where the emphasis has been on the sort of second part of the plan, which is scaling RL.”

The Evolving Capabilities of AI

56:00 to 58:01

Explore how AI has developed the ability to think for longer durations and tackle complex problems.

“and through AI doing good things for us, the things that we want.”

Reinforcement Learning and Alignment

58:01 to 1:00:44

Delve into the relationship between reinforcement learning and the alignment of AI models with human values.

“Is there a concept of, I guess, online RL that happens where, as the agent does something and learns from the real world, the RL happens in real time?”

AI's Success in Programming Competitions

1:00:44 to 1:01:10

Learn about OpenAI's success at the ICPC World Finals and the implications for AI's capabilities.

“explaining to it the things that we that we want from it uh but it's a very definitely very important and central part of any, should be of any AI research program.”

Insights from AI's Competitive Programming Performance

1:01:10 to 1:05:19

Understand the technical background of AI's performance in programming contests and its broader significance.

“But taking a quick sort of going down the rabbit hole a little bit about math.”

Challenges in Reinforcement Learning

1:05:19 to 1:09:14

Discuss the complexities and challenges associated with reinforcement learning in AI training.

“And I think this is like where we want to be, like solving competitions is cool, but people solve competitions to prove that they can go to the actual frontier level job and solve new technical problems.”

Pathway to AGI

1:09:14 to 1:10:05

Examine the potential pathways to achieving Artificial General Intelligence (AGI) through innovative training methods.

“So maybe to zoom out, to close this conversation, you said the other day, you tweeted, we all collectively believe AGI should have been built yesterday.”

Exploring the Path to AGI

1:10:05 to 1:15:38

Learn about the ongoing research and philosophical questions surrounding AI and AGI.

“And there will surely be a few things more.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00O1, to be perfectly honest, it was really mostly good at solving puzzles. It was almost more like a technology demonstration. O3 has really been something like a tectonic shift, the trajectory of AI. GPT-5 in some way can be considered as like O3.1, which I am after as something that would be the next pretty significant jump. We understand that there's only one time in history where AI is being built and deployed and developed. We are together in this goal that's larger than every one of us. Hi, I'm Matt Turk from FirstMark. Welcome to the Matt Podcast. Today, my guest is Jared Forek, VP of Research at OpenAI and a member of the Metis list of the world's top AI researchers.

0:43In this episode, we go deep on how models actually reason. We also go behind the scenes at OpenAI, how a few big bets get staffed, why everyone knows everything, and how that culture ships fast. Please enjoy this great conversation with Jerry. Hey, Jerry, welcome. Hello, very happy to be here. We are going to talk about reasoning a lot in this conversation. At a high level, what does reasoning actually mean? When we talk to ChatGPT and ChatGPT says it's thinking, What actually is happening behind the scenes? I think that the thinking process is at least a good analogy. As we were in the early days of AI, always had this goal, dream of trying to teach models to reason.

1:29We were thinking about it, spending more time to get better results. If like human is posed with a very hard problem in front of them, very rarely they have answers straight away. Sometimes they need to find that answer. sometimes they need to perform certain computation sometimes they need to like look up some information sometimes they need to teach themselves something and the process is reasoning is like getting to an answer that you don't yet know like in some way it can be called search but it's not really like a very naive search search is a loaded word but reasoning is the process of getting to an answer and a work that you need to do that is like longer than what usually is considered answering a question i think that that difference is here like answering a question usually means you already know the answer and you just just just elicit the answer you know and the process of reasoning is getting to the answer that you don't know and usually the longer you spend on getting to this answer for whatever you need to do to get there the better the better it gets and uh we've all become familiar since you guys uh released uh one i guess a little over a year ago and in september 2024 with the concept of chain of thought, which is in layman's term, the little messages that you see when you query ChatGPT and it tells you, it shows its work, it tells you what it does.

2:52What does that actually do? Is that a logical tree and it eliminates option after option? What actually happens? Language models do on their own fundamental level is they are often called as next token prediction machines. And that's not completely accurate in the age of reinforcement learning, but they still operate on mostly on tokens that are mostly text. The language models, again, are these days also multimodal and they operate on mostly text. But to simplify a little bit for a second, language models generate text. And what chain of thought is, is their thinking process verbalized using human words and human concepts.

3:36so the the magic that we are seeing why this is all possible is that while you are training on all of internet on a lot of human knowledge and human thinking process the model starts learning in some ways to think how humans do and in some ways get to the answers how humans do from seeing humans do it a lot in the text that was that was pre-generated and that that was based in a training data on humans. And then the chain of thought is basically eliciting that capability in language models of thinking and getting to an answer like humans do. A lot of what early chain of thought work was doing was kind of solving math puzzles.

4:24And the first most famous prompt to elicit chain of thought in language model was so-called, let's solve it step by step. There is this very classical result in language models that if you ask them, what is some mathematical expression or some puzzle, they will try to give you an answer. They will try to predict the next token, but they fail. It's a hard thing. They can't compute it in one token job. But if you ask them, please do it step by step, they will start thinking, okay, I don't know the answer, but the first step of getting to the answer is this. And then they write like chain of thought, which is a series of text, series of tokens doing the first part of the computation, the second part of the computation, the last part of the computation.

5:06Then they connect those things and then they can get to the answer. So the chain of thought is basically a process of thinking encoded in words, how humans would solve a problem on a piece of paper going step by step from start to the end. And since time and by that, I mean the time spent thinking is so important to that concept of reasoning. How does a model decide how long to think when we're in chat GPT-5 and we're in auto mode and it says that it's going to decide automatically how long to think? What happens there? It's basically part of our optimization process, partially for the happiness of the users and what they want to expect.

5:56Because when you have a thinking process, you need to balance two things, which is the quality of the result. As we said, there have been those pretty great scaling laws that we demonstrated with the release of O1. The longer the model thinks, the better result you get. But also, people don't like waiting. Waiting is time lost that you could do something. Everyone wants to get results as quickly as possible. And there is this saying you can get cheap, fast, or good, and you can take two. And that applies to language models as well. There is a trade-off, and it's delicate. That's why we also expose some of that trade-off to the users, where you can have a high-rezoning model and a low-rezoning model.

6:41And this is, in the end, the same model. We just tweak the parameter, which says we want you to think longer or shorter. We try to encode some heuristics of what we think the users will want when thinking on an answer a little bit longer and getting to a better answer is worth it, waiting, versus not. But it's a bit of a trying to guess the anticipation of the users. What's the right amount of thinking for them in this particular situation? Fascinating. So it's more user-driven. So it's more like a user experience kind of thing. In the end, it is, because the question is like, how long do you want to wait for answer?

7:18You can always wait longer and get an even better answer. It's been a little over a year since the release of the world's first reasoning model, O1, which is an effort that you led. What has been the journey since? So there was 01, then there was 03, then most recently HGP5. How would you characterize the evolution of reasoning specifically across the three models in the last year? In some way, how I characterize our reasoning or scaling up reinforcement learning research program is we do a series of scale up runs that are progressively more and more ambitious. Everyone, we try to do something more, something larger scale, something that should result in a better trained model than the last one.

8:11And obviously, we don't release all the models that we train. some we release, some we think need to wait a little bit longer for the moment where they will have their time to shine in the hands of the users, but 01 was the first model we decided to release to demonstrate to the world those models, and 01 to be perfectly honest, it was really mostly good at solving puzzles and maybe a few kind of thinking problems here and there, but it wasn't yet very useful model. It was almost more like a technology demonstration than actually really polished product. But we were thinking we have something cool and we wanted to share it with the world as OpenAI.

8:5703, I think, changed that pretty significantly. In some way, it is a model that is meaningfully useful and a little bit self-serving, but it was a moment when I started using ChatGPT quite a bit and And I am basically a user completely hooked on reasoning model. And in chat GPT right now, I use basically exclusively reasoning models because those are the only models that I trust. The output and the result and I think O3, its ability of using tools and getting to answer, leveraging a lot of contextual information from various sources and persevering towards getting to that. has really been something I think there was a little bit of a tectonic shift in the trajectory of AI and I think we did something really really great there GPT-5 in some way can be considered as like 0.3.1 it's a little bit of iteration of the same thing and the same concept and what I am after and my team right now is something I think next that would be the next pretty significant jump of how we interact with models that are even more capable, thinking even longer and interact with even more systems and sources of information on their own journey.

10:27but but like separately in the meantime we continue to build a lot of things on top of all free technology like codex which i think like coding agents are at the moment the first like really successful agentic products built on top of ai there are things like computer using agent it's called gpt agent right now i think and like the pre-search and a few a few other things that that we will we'll keep on building on like all free generation technology great all right so we're going to go into all of this in much greater detail in a minute. But before we do that, let's talk about your journey. I think it's a super fascinating topic for all of us.

11:09You guys are changing the world. So I think I'm curious and we're all curious, I think, about the people, the human aspect of who those people are that are just having such an impact. So you, starting from the beginning, you grew up in Poland, I believe, right? Yes, I grew up in Poland. Walk us through your formative years and how you got to get started in this field. Yeah, yeah, yeah. Happy to do that. An interesting fact, it's almost like a crystal starts from something and you put a little bit something in the beginning. there's i think one part that was like important and part like like the starting point of my journey where i i didn't know where it came from because it was with there with me from the very beginning of my life in a moment that i don't really know when it started it just was always there with me i always thought that being a scientist and doing science is the highest calling a human can have and i don't don't really know where it came from like you know my parents maybe like were singing the light bright lullabies to me like when i was when i was one or something like that but basically since i remember i wanted to be a scientist in the early years i also discovered like you know i have i have talent for for those things i like i was going to school and i saw i get things slightly faster than people around me at least like you know in a regular school and in the middle of poland and which which like you know made me made me kind of like doing those things like studying maths and science a little bit more because because like you know I felt it felt good in a way it felt like this is this is something that naturally naturally fit me and like you know I grew up as a like you know very regular kid just just like you know being slightly nerdy guy and trying to balance my my side of being interested in science programming maths and like having some social life and i definitely have had some kind of like a party arc in my in my life but i think i think that the most important part and moment was when i actually went to university college university of warsaw is where i went and i decided again to study mathematics at around that time of being 18 my idea of life was to be a mathematician with a pencil sitting in a room with a piece of paper and solving equations this is kind of my 18 year old dream of how how life should be lived and what i what i want to do in in my life and like you know i i do have like you know my my personality is built again in a way of really appreciating like you know solid science pursuit of truth great great engineering and all the all those aspects but I definitely have also a little bit of like a misfit kind of rebellious tinge to it and that resulted after a few years of studying mathematics I what I realized about myself and about the world is that I really like maths and I am I'm quite good at it but I didn't like academia that much and I realized I don't want to like stay in academia I don't want to stay in university and that this would not be environment where I thought I would be long term very happy and fitting.

14:31It felt a little bit too rigid, a little bit too structured in a way that I didn't know if I will feel good. And in some way for a young me, I was around 21 years old at that moment, that was pretty big crisis of faith for me. I had a moment like, you know, lost purpose of life. So I just did like, you know, a very simple, like first principles thinking, you know, I am graduating with a degree in mathematics. I need to get a job to get food. And like, you know, what job can I do to use mathematics and in that job? And looking at the job market, that moment was 2011, I think, or 2010, somewhere around that.

15:22I decided to become a trader and trade for a living as the one way where I can do what I like, which is mathematics, and get a career. year i got i got a quick internship at jp morgan uh investment bank trading floor and equity derivatives group spent six months there learning a little bit how how does trading work and what does what does what does it look like a little bit i finished my degree uh i got i got a message from like a boss of my boss at jp morgan something saying hey jerry you were like you know one of our best interns ever that we had we really really liked you working with us and we are we are living bank and starting a new hedge fund would you want to come with us and like for for 20 21 or 22 year old jerry that sounded like kind of like a cool adventure type of story that's that that i was interested in and going there and do it it had enough like enough interesting problems to be solved and at the same time it had like yeah this this kind of trying trying something new trying something ambitious kind of kind of that i generally like so it was in london that company didn't really work out unfortunately but it was it was hard and ambitious and not everything works out i did try that again starting another hedge fund from scratch with a few other people in amsterdam i worked there for a few more years and eventually eventually i got bored like i I generally working in trading is an interesting and exciting problem.

17:03Market is very hard. And the depth of what you can go of trying to understand it model is very deep. And I worked with pretty smart people overall, but I stopped feeling I am growing after a few years of doing that. And at the same time, together with a friend I was working with, we just started chatting about AI and about this. artificial intelligence. And what really drew me to artificial intelligence was reinforcement learning. And specifically the DQN agents trained by people in DeepMind in 2013, but I think it was a few years later that I actually learned about those results. From my perspective, and again, this is just how my brain works, the 2012 ImageNet results weren't that significant like during during my university years i learned a bunch about like how classical ai the neural networks weren't very fashionable back then but i still learned about what they are i learned about svms and all kind of methods how you train classifiers and for me it was kind of like obvious and natural if you have enough parameters and tweet it hard enough you will fit a classifier to whatever you want it was was kind of obvious well it's like what was to me not obvious is I never considered classifiers a smart thing classifiers that you learn a function on some set of inputs to have some some set of outputs and you can can keep training it to approximate better and better what but kind of what was something that I missed back then is that when you can fit any function better and better you can start shaping behaviors and strategies and when I really saw that was in the in the like DQN results where they applied the same things that worked in ImageNet like neural networks and they weren't particularly big or impressive neural networks with a classical field of reinforcement learning to solve simple computer games and turns out those simple neural networks with a simple learning algorithm started learning pretty complex computer their games and exhibiting very interesting behaviors.

19:20I saw those behaviors, I saw those results, and I was like, this is what I want to do for the rest of my life, which is like, you know, not a very long horizon with a 20-something things about, but I was like, this is what I want to do. Where do I do that? Like, you know, Google search, where are places where you can do reinforcement learning in this world? Like, you know, Google DeepMind and OpenAI came up with this kind of, at that moment, pretty small and like somewhat known but they are yeah you joined you joined open ai in 2019 right so like very much uh yeah very much in the early days still very much in the kind of like non-profit era yes of open air so how did you connect with them i just applied for the website i had the most it was like boring and uninteresting thing in the world which is like you know open ai.com jobs like you know apply send resume and hope they respond and then you know like luckily enough, I was, they did.

20:15I don't know how many resumes OpenAI I was getting in that time. I think it's definitely much less than today. But I was like, I came there and I was like, that doesn't matter what do I do as long as it's reinforcement learning. So you joined in 2019 and with a passion for reinforcement learning. So was that around the Dota 2 moment? Because OpenAI, interestingly, in those early days of 2019 did a lot of reinforcement learning focused work, right? And then there was a whole unsupervised learning, GPT moment that happened afterwards, but it started from roots in reinforcement learning. So did you work on that project specifically or was it too advanced by the time you showed up?

20:59So the project that I worked was robotics project at OpenAI, which shared the same code and same methods as the Dota project. Like in one hand, the Dota project was OpenAI's way to demonstrate the world, what scaling up reinforcement learning can do. And in some way, it was like taking the 2013 DQN agents and just doing all the hard work of making it bigger and bigger and solving harder and harder problems. And OpenAI generally from the very beginning was aware and really, you know, it was simple but genius insight that you need to have a large scale system to learn really, really interesting, complex behaviors.

21:43And that was one way of what Dota was trying to show that by scaling up reinforcement learning, we can solve pretty complex environments. And then there was another project, there were, I think, three reinforcement learning projects at OpenAI at that time. And the second one was robotics, which is applying the same methods that we now knew or were proving that can solve pretty complex computer games. can they solve all the practical problems OpenAI was always optimistic and ambitious and trying to see if we can scale our own to solve data can it load my dishwasher, can it fold my clothes, can it build a house and this is what we are doing, the project I was working on was focused on dexterous manipulation which was back then and still continues to be an elusive challenge for trained policies And we got to a showcase of demonstrating that a hand controlled by a neural network was able to solve Rubik's Cube, which is a pretty delicate and complex task to do.

22:47So fast forward to today, still in the same vein of like the behind the scenes of, you know, you all had OpenAI and sort of life there. What's a day in the life of Jerry? Like what does somebody like you do? Like you read papers, you train models, you manage teams. What's your day like? Yeah, my days are surprisingly uniform, which is I come to the office early in the day after driving my kids to school. Then what do I do all day is basically talk to other researchers. I talk to other researchers all day, every day. And this is basically exclusively what I do. I take ideas from people, bounce with them, brainstorm with one partner, then move to another one and do the same thing over and over and iterate.

23:41and in that way, keep refining our research program. Sometimes those are group meetings and group meetings are there as well and have their own team dynamics. But that is basically exclusively what I do. The only thing that changes is the topics of research from meeting to meeting and from person to person. How are priorities in research determined this range of possible projects? Is that top-down? Is that bottoms-up? Do people suggest ideas and others vet them? How does that work? Yeah, yeah, yeah. It's the art of structuring, organizing, and leading a research project is something that I generally learn to appreciate very quickly in OpenAI's journey and in my career.

24:34There is something we are good, is structuring research projects. And I think it's a unique mix. We can't say it's top-down. We cannot say it's bottom-up. It's a mix of those two, which is balancing all the important aspects. One thing OpenAI embodies and determines is that we all work on a very few projects total. There are not that many projects. OpenAI is not trying to do everything. We are not trying to have portfolio. We are trying to have multiple different bets. always the idea is we do a few core things really really well and put a lot of effort there which means there are there needs to be a lot of people working together on the same large scale large ambition project and we have we have a few of those all number probably three or four depending depending on how do you how do you call it and that's and that's it and from that perspective like people don't have ultimate freedom it's not that that people come to open ai and say hey i want to do this and they just they just do this because you need to do something towards the goal of one of those four projects and then within those projects like we try to be like relatively bottoms up in in a way as long as it feeds again into into those goal and the most important part of the research lead is keep to making sure all the researchers are working towards this one shared goals and that don't fracture in their own ways of thinking and doing things.

26:04And so it's an incredibly hard thing. It's a very, very hard job and not always easy to visible how delicate it is. But that's a lot of what it is. I don't think top-down structuring of research doesn't work in research organizations. I really don't believe in it because you are not kind of hiring some of the smartest people in the world. And OpenAI has incredibly, incredibly smart people to kind of tell them what to do. They need to figure out what to do, but they cannot figure out in the whole space of things what are cool things to do. They need to figure out from within the space of what the project needs and what could advance the research goals of OpenAI the most.

26:51And to which you're saying, is there a collaboration between the teams working on the three or four projects at the same time? Because I would imagine putting myself in your shoes, OpenAI shoes, there is probably a tension between wanting to be collaborative in general. But equally, I mean, this is probably the most important IP in the world. So you probably want to make sure that not everybody knows everything about everything. Well, perhaps not. I'm speculating here. How do you think about that collaboration versus some protection of IP? You'd be surprised, but the truth is in research at OpenAI, which is around slightly less than 600 people at the moment, everyone knows everything, really it does.

27:37And we always have been fully transparent. And in some way, you are a little bit shooting yourself in the foot if there is a researcher that doesn't at least have the chance to learn about everything because they don't have the best information to their job in a best way and like it is like yeah it is some like a risk of losing ip but i think i think the risk of not doing the right thing and of people not being informed about about research and not being able to do the best research is much higher in my in my personal opinion and how how i i approach those things so we are we are extremely internally transparent within research and and that is that is one of our operating principle as the goal is to do the best research we can and train the best models we can consequently and the culture generally is very collaborative like you know it is always the case when you have 600 people when you have groups of people they're always like one person doesn't like the other person because they looked at them weirdly or one person thinks the other smells bad or just doesn't like their ideas.

28:42Like that does happen. Those are humans and those are human things. But generally in large scale, I think we really are, have this belief like, you know, we are together in this goal that's larger than every one of us. It is a very positive sum game because the AI seems to be getting like, you know, only more and more significant and the success of open AI is far from guaranteed. It depends on us doing great work every day. So there's a lot of like feeling of shared fate and the fact that we all need to rely on each other to do our job to achieve this shared mission. So I generally think with all the caveats of human nature getting in the way sometimes, I think on a large scale OpenAI is very, very collaborative.

29:27How do you all manage to keep that pace of releases? It seems to me from the outside that there's a tension again between research, which in some ways kind of feels like it could be a long-term kind of thing. And on the other hand, you guys seem to be just shipping and shipping and shipping and shipping across the organization, but including in terms of core models, again, to the point that you went from 01.203 to GPT-5 in like a year. How do you balance all of that? Why are you guys able to ship so quickly? I think the fundamental reasons for it is, in general, OpenAI, in my at least worldview, is a generational company in a way that we have incredible momentum behind us.

30:20We know that we were doing pretty great in the past and we need to continue that. We have incredibly smart people. Literally, the most talented people in the world are all coming and want to work at OpenAI right now. which means like people every output per single person is incredibly high and everyone really every every single person does a whole lot so so we have we have like a momentum that carries us forward we have really great people that work together we have like good operating like way of structuring research and and and and can borrow a lot from silicon valley how how to get things done quickly and people are generally very excited about work everyone feels the weight and potential of what we are doing what we are trying to do and because of that people at open ai have a tendency to work pretty hard and and like you know having great people excited about what they are doing or working together reasonably well results in doing a lot of things we understand that there's only one time in history where AI is being built and deployed and developed and people want to do it in the best way that is possible.

31:35Do you all use a lot of your own tools? I think Fiji, Simo was tweeting the other day that I think in the latest, what you all announced at Dev Day today, a lot of it was written by Codex. Is that part of the daily experience? Do you use a model to come up with new ideas for a model? Do you use a codex to write the code? How does that work? Yeah, we definitely use codex a lot for coding and this is only getting better. As I said, I use ChatGPT a lot, although not surprisingly, not that much for actually coming with ideas. But for a lot of questions that I have, I think I am a pretty heavy user of ChatGPT right now, happily paying like$200 a month for it.

32:22And I think I'm getting... They make you pay? For a while, it's worth it. They are making you pay. And I'm kind of like pretty, pretty okay with this because then you get pretty generous like usage limits and not really bottlenecked on it. Thank you for all of that. Let's switch tags and go back to how all of this works. So is the right way to think about modern AI systems at OpenAI. By modern, I mean as of October 2025 versus the old days of, you know, nine months ago. So exactly. So is the right way to think about it as a combination of pre-training and RL? First of all, is that the right way to think about it?

33:10And second, if so, just at a high level, how does the articulation between both of those work? And then after that, I'd love to do a little bit of a deep dive on RL to make this very educational for folks. Today's language models basically can be thought as first they are pre-trained, then you do reinforcement learning on it. The reinforcement learning would not work without pre-training. And I think in a similar way, pre-trained models have a lot of limitations that are very hard to resolve without doing something that looks like reinforcement learning. So I think both of those bits are here to be and to stay like i think the way how they are like combined and do that may and probably will evolve in the future nothing should be treated as dogmatic and fixed and we need to we need to keep generally figuring out the way how to train better models and this is what we are trying to do the interesting thing and you know i i can credit that to to ilia how much foresight that he had but whenever I was like started at OpenAI early 2019 right and and and I remember there was like research all hands or something like that where Idyia came on stage and talked about like what is what is OpenAI's research program what are we trying to pursue and what he said and at the beginning of 2019 was to train large generative model on all data we can and then do reinforcement learning on it that was that was the open ai research plan at the beginning of 2019 and this is exactly what we are doing today now the algorithms change architectures change like i don't think he was even thinking about transformer at that moment like you know the gpt was like there was some gpt but it was like a toy example that someone was playing with but the goal training large generative model all the data in the world and then doing reinforcement learning with it was already there at the core DNA of OpenAI and that's what is happening right now.

35:12So let's do, if you will, a little bit of reinforcement learning 101 to make this really interesting to like a broad group of people listening to this. So in very simple terms, like explain it to me like I'm 10, what is reinforcement learning? Yeah, I usually the metaphor and the analogy I have to reinforcement learning is like training a dog. It's very, very close. And I used to have a dog when I was a teenager. And even I remember my parents did. I didn't know anything about raising a dog. But they kind of invited through some friend of a friend, a fireman, who I think was working with like service dogs.

35:57And he came to me and he basically told me a little bit about how do you train your dog. And what most dog owners that are ambitious about training your dogs know, it is always extremely important to have a bag of treats in your pocket. That's what you always do. And whenever you see your dog behave well, what you should be doing, you should smile and you should give your dog a treat. Whenever you see your dog do something bad, you basically like give your attention away turn away and become sad and before the years of breeding the dogs discover it's a it's a it's a like you know bad reward and bad behavior and this is exactly doing that but with models we elicit a lot of different behaviors in the models put them in challenging situations and then we give them cookie if they do something we want if they do a good thing and give them some kind of punishment and negative reward if they do something like that we that we don't want and that we don't like in a good way a good way to do rl is if you balance those things so if you kind of give cookies half of the time and punish the other half the time but this is almost like a mathematical kind of uh kind of aspect of it but that's that's the most important part, which is like elicit behaviors, reward the good ones, and then going forward, the model will be most more likely to do what you want and less likely to do what you not want, and for that it improves.

37:27It is the way how to train models to elicit actual behaviors. That is not versus next token prediction. If you pre-train them all, you literally train them all to predict the next token. RL is a completely different gradient and a completely different set of what we want to get out of them all. And getting the model to do what you want just for some vocabulary and semantics, you hear sometimes the term policy. So in RL, you hear terms like agent, environment, action, reward, and policy. So I think a lot of those are sort of self-explaining, but policy is what? That's a strategy, that's a behavior of the model?

38:11Yeah, policy is the behavior of the model as the model weights represent what it does when put in a different thing. Like model in the end is a mathematical object and you can define it. Policy is a mathematical function that maps observations to actions. What you see and then what you do with what you see. Yeah, so agent is a model. Action is what the model does. Reward is how you say whether that's good or bad. environment you you hear a lot of things uh these days about designing the right environment for rl what does that mean environment is like in some way it is everything that the model sees but but the interesting thing about difference about rl environments and most other types of like what you can call supervised learning or unsupervised learning is that reinforcement learning environments you want them to be interactive you want like them to evolve as the model does things in general like like similarly how like if you want to like learn how to play guitar you kind of take a guitar and you strum it and what happens is you hear sound of that and then you hear it and then you can do like learn to play like like with with actual like feedback of what is happening with the guitar and it's in that way you know the environment is like how does the world react to your actions and a lot of like what drives your actions is the is what is happening in your in your environment and it's in in your in your world and that's kind of like the only way how to like really teach agents to like learn to react to changes in the environment is through it's reinforcement learning can you give us a little bit of a bird's eye view of the evolution of rl over the years?

40:03Mostly how does modern RL differ from historical RL? Yeah, yeah. Again, there was super historical RL. Not even that old, but the main tectonic shift was when combining neural networks with reinforcement learning. Reinforcement learning predates neural network as a general mathematical method of optimizing behaviors in like mathematically defined environment and as a method of study. That's what is known as deep reinforcement learning. Is that right? Yes, and then the deep reinforcement learning that basically like deep mind, like invention of combining neural networks with reinforcement learning, the DQN moment I talked to you about.

40:55And then from there, like there was a moment where like there was a pretty active like area of research of reinforcement learning on games like when even when i started like 20 2019 the reinforcement learning was was kind of fashionable at that moment although not like very successful but we were the reinforcement learning was able to solve a lot of games but the bottleneck was there that the models were not pre-trained in any way we were we're training a lot of behaviors playing games like we even got alpha go moment out of that which a lot of people got very excited about it but was still like learning behaviors without the models that were that were meaningfully smart about those behavior there were still a lot of kind of like you know you don't want to call it caveman intelligence but but something something in that regard about models not being really smart even though being being pretty heavily reinforced and there was like a long research in that and A lot of cool results and theoretical understanding of RL comes from those days because people were researching RL actively.

42:00But it was in some way at that end of doing RL without pre-training. And then in my moment when I finished working on robotics, I started working on teaching language models to code. But having pre-trained models was a really big deal. And the GPT era of scaling and of large-scale ingesting lots of data to really train great models enabled us already at that moment to start RL. And that was one of the first things I did almost immediately. Whenever GPT-3 was trained, I tried to do RL on it. And there were always bottlenecks. The systems were kind of clunky. It was hard to figure out. What are the right algorithms?

Read the full transcript

42:48what are the right problems to work on it and what is the right algorithm to train it on and what opening I did at that moment and what's kind of like how research goes we kind of cargo-culted a lot of things that was used for games and almost the same things are for robotics and the first RL I was doing on large-angle model was kind of the same PPL we used for everything and like it gave some results but those like early results weren't completely mind-blowing in RL and there was a long time where we're keeping on investing it and you know personally I always believed there will be a really really big moment for RL and language models but the early trials and errors weren't super successful at the moment when we when we trained GPT-4 And there was an interesting moment where we trained GPT-4 and everyone today thinks, oh, GPT-4 is such a great model.

43:52But when we trained GPT-4, we were pretty underwhelmed internally. And there was a lot of moments, oh, we trained this model, we spent a lot of money on it. And it's kind of like, you know, pretty dumb. At least, at least, like, you know, we have GPT-3, GPT-3 already does all that stuff. And GPT-4 doesn't really seem to be that much better. And we had this kind of question, it kind of seemed smart on evals that were one token long. It seemed to be able to give a pretty detailed answer to complex questions where it was one token. But if you actually let it speak for longer, it wasn't very coherent or really gave a very long answer.

44:28We needed to answer this question, how do we actually make the language model that seems to have some smartness in its way? actually sound smart and actually be good and talking to it and that was that was the moment where a technique that was developed already a few years earlier like really shown which was called rlhf just basically doing ppo on large language models with the reward given from human preferences of seeing two parts of that text thumbs up and thumb down yeah thumbs up thumbs down like whatever whatever human preferences is and that's a very good reward because there are a lot of things how the model can generate bad text and how early gpt4 was generating bad text in a lot of ways and rlhf was able to catch those things and correct it and you know reinforce good behaviors reinforce generating good text and then and then punishing bad text and in the end gpt4 plus rlhf together as a package like you know deliver that gpt moment to the world that everyone sees And as much as it is a big success of pre-training, it actually also was pretty big success of RL in the form of RLHF.

45:37Amazing. And just to double click on that, the RLHF, so we're all familiar as users with, I mentioned, Thumbs Up and Thumbs Down, that's on the interface. But the actual RLHF happened post-training. Is that right? Yes. And what did that look like as an effort? Or did you have a bunch of just humans sitting down in front of the model, industry specialists maybe, and give it feedback? How did that actually work? RLHF was a research program that was already happening in the background for a while. I think we did RLHF, at least I remember GPT-2 being RLHF for quite a bit. That was already there and already happening.

46:23it's like gathering data for even for rlhf it's own like research domain basically and always thinking what is the right data to train them all what is the right data to train like your rewards and how do you how to shape your rewards it's it's a research that we've been doing and it's it's a very open-ended and very deep in many many different ways and like what like i think there are papers written on what what how what rlhf is but there's a lot of depth to it but like you know long story short is you have what we call ai trainers these days and they look at outputs of their models and they give them scores and then you'll learn basically a model of those scores and use that for training and that's part of for people who may be curious like the entire data labeling industry so scale ai and and a bunch of others that's uh what they do right yes yes I think in a way, I think it's getting more and more to be a thing of the past as the models are getting smarter and smarter.

47:22This is becoming less of a thing. But I think a few years back, and especially in GPT-4 days, this was the thing. The interesting bit about data labeling industry, I'm not sure how much we want to go on that tangent. It has to constantly reinvent itself because the AIs are getting smarter at some moment. Certain things you don't want to label with humans if AI already can do it. So you move the frontier and you change the type of data you are labeling as you already RLH-tef the previous part. We've been talking about RL, but the first phase of all of this is the creation, the pre-training of the models.

47:57That is unsupervised learning, right? Do you want maybe for, again, to make this broadly interesting to people, define unsupervised versus supervised and in what way was the pre-training unsupervised versus self-supervised or whatever nuance? Yeah, I think those are nuances and I don't think they are as stark and as sharp as some people like to determine them. But why pre-training is called unsupervised because in some definition of it, you don't need any extra labels to the data that you feed into the model. You just feed the text as this. In some way, you may argue that the data is already labeled because it is self-labeled.

48:45If you give the model from the text, predict the next part of text. In some way, it is a label, but it is self-supervised because we don't clearly tell the model what is right or what is wrong or what do we want from it or what do we not want? We want it to just predict the other part of data. They can do the same thing with images. You can mask the part of an image and tell them all like predict the next bit of image. But whenever there is this like classic machine learning notion of like targets and labels, like I guess we were talking about classifiers. So supervised learning was like, you have some notion of targets, what your targets are and some notion of labels.

49:24Supervised learning was like predict those labels from targets. and this is like a like some type of mapping but actually what's what's interesting is that there are many more bits usually in the targets than in the labels and studying the structure of targets itself and yields much more learning and much more intelligence than learning the mapping itself so like spending a whole compute on just just learning the data itself without the labels is the right thing to do and like what is often called like representation learning and studying studying the data and its properties. Okay, great. All right, so going back to RL, you tweeted the other day, the GRPO release has been, in a large way, has accelerated the research learning program of most US research labs.

50:12So what is a GRPO? It was a little bit of a tongue-in-cheek moment. I am extrapolating here a little bit what exactly happens because I haven't been in most US research labs, but I have some mental model of what happened and how. And long story short, GRPO was the open source released from DeepSeek. And there was like everyone who like is terminally online follows AID scores, knows that DeepSeek moment of whatever it was when the Chinese company that seems to be doing really, really great work released new model. And it was also a pre-trained model, a reasoning model. They open source the algorithm, they open source a lot of things they did.

50:53overall like really really great and technically excellent release and like there was a lot of discourse about about like you know that they they pre-trained their model particularly cheaply and that was that was part of the discussion about that deep seek moment but the other part of the discussion was that they that they kind of like released their reasoning process it was like not very far after our 01 release as far as i know like our 01 release mostly caught a lot of u.s labs by surprise they didn't have like similarly advanced rl research program to my knowledge basically no one and like and i think like the only company in the world that again i am as i am aware there's probably a lot of things i i didn't know but you talk to people sometimes you hear rumors so this is my my version of the world is that if you look at the other papers of deep seek like that company was doing pretty similar in some ways rl research to what we are doing and i you know i think i have to clarify like what we what's openai is doing is not exactly grpo it is slightly different in many different ways but like some parts are are definitely are definitely similar and what's what's most important those are both like large-scale policy gradient algorithms and like deep seek was doing company was doing research in a slightly adjacent area they were they were they were not very far and when we whenever we released o1 and we told the world that you can get pretty like those great results with scaling up reinforcement learning on language models i think it was like not very big hop for for the deep seed company to kind of realize okay we are not very far from getting similarly good results and they did it they trained their reasoning model and they released it and they told the world how pretty like not very not much later than we released all one and i think like you know for a lot of u.s research lab that didn't yet know they didn't have a research program how to train like reasoning models.

52:48They looked, oh, there's this Chinese company they released how to do it. It helped us kickstart and train reasoning models much faster than we would have to otherwise if we would have to find all those bits ourselves. What does it take to scale RL? So if there was a phase where OpenAI was very focused on pre-training and then, if I understand correctly, the last 12, 18 months or whatever time period where the emphasis has been on the sort of second part of the plan, which is scaling RL. Is that a question of just giving RL more compute, more data, more labeling, as we were saying? What does it take?

53:30The first thing that is important to know and understand, RL is hard. Like, conceptually, if you think about it, and there's still a lot of depth to it, But very conceptually, mathematically speaking, pre-training is that simple. It is kind of the simplest thing you can do. And there has been a lot of thought and a lot of optimization already put through that, through a few years of optimizing and doing very well at very large scale, very simple mathematical operation. In terms of RL, it is much, much more complex. There's many more things going on in a reinforcement learning run. there's many more things that can go wrong in in doing it especially as you as you scale up and many more types of bottlenecks failures uh like it's it's it's a much more much more delicate thing and there's much more uh like much much more room for error in some way you know i don't want to go too deeply in the parallel because it's a little bit overblown but in some way like you know just just to give some coincidence you can have a steel factory which makes steel and like the process is relatively standardized and you make blocks of steel and they are they are uniform and nice and well defined what it is versus like building semiconductors which there are there are very very few companies in the world that can do it because there are so many things that can go wrong and you have to put a lot of attention to details to make great semiconductor and it's very complex internally.

55:03And in many ways, this is kind of like, I don't want to diminish because there is a lot of very hard technical difficulty to do pre-training at a large scale. But there's just many more moving pieces and many more elements of the reinforcement learning stack that need to get the right to get a large scale run successful. You mentioned working on such LGBT agent, like the agentic AI, where does that all fit the tool use, like the whole agentic autonomy versus reasoning RL, like help us to reconcile what does what and what impacts what. I think what is the important thing is I believe and I think that there can be a lot of positive impact of AI on our world and on our lives through automation, through problem solving, and through AI doing good things for us, the things that we want.

56:05And for a long time, and for a long time again, it's not that long, but the last two years or so, or maybe approaching three, we've been like living in this world where we kind of ask questions to AI and gave us an answer at the beginning instantly. Now it can think for like a minute or two, which feels long, but in many ways, what can you do for two minutes if you think of how many problems you want to solve. And AI is probably a little bit faster in the things it can solve, but it's still a limit of what it can do. There are still a lot of tasks that you know that would take AI to do much longer.

56:44When I prompt Codex, it works for a while, again, a few minutes. There are a lot of things we have internally and we are doing that allow the models to work for much longer. We still didn't figure out the right product to deploy them, but the models can think for 30 minutes, hour, two hours these days on certain types of tasks and problems, even longer than that. They generally are capable of doing so, and we need to figure out how to make that process more useful and more being able to actually come to various problems in real life, whatever it is, coding or booking travel or making plans or even designing houses or new electronic devices or whatever else we would like models to do, we would like them eventually to be able to do for us.

57:42And a lot of this comes through the models thinking independently for longer periods of time and considering more of our alternatives, bits, and just sometimes going through a slog of very long lists of tasks. So the agentic part is powered by fundamental reasoning. Is there a concept of, I guess, online RL that happens where, as the agent does something and learns from the real world, the RL happens in real time? So generally, all of RL is happening. Most of RL that you hear talk to language models is online, but it's done online in a way that is still a training run. It's still being trained kind of separately from the users.

58:32There have been a few models in the world, and I've learned recently that I think Cursor is trying to train some models online with their users in the loop, and it's theoretically possible to train models like EnchatGPT or every other product just responding to the users and reinforce through whatever rewards you get in there. But this is not what I am aware, at least not what OpenAI is doing at the moment. And it can be great, but it can be also dangerous because you are not really very much controlling what you are reinforcing in that loop and what could happen. And so, yeah, at least until we have a really good safeguard, I don't think we should try to do that in anything like as complex and large scale as ChatGPT.

59:19Yeah, interesting. And very much on that note, talking about alignment for a minute, is alignment an RL thing? I mean, do you create alignment in the model by teaching it what is right and wrong? Kind of, yeah. It's a little bit of yes and a little bit of no. In a way, alignment is about steering the model to certain behaviors. And that is definitely an RL thing and an RL problem. But also, you want the models to know what is right and what is wrong and understand the world. And not all of them are reasoning and RL problems. Those are very often just as AI problems. and we like in the way to be aligned the model needs to know right or wrong to choose right i don't think like you can just tell them all like you show with a few good things to do and it will do them all needs to deeply understand its action and consequences to really to really be able to choose the right thing and it's a i think it's a never-ending pursuit because like even even for humans it's not super easy to uh to define what's what what do we consider aligned i think as our as our civilization will evolve it will the notion of of alignment and and and the goals of humanity will keep evolving and we'll need to keep nudging the model towards those things and keep keep explaining to it the things that we that we want from it uh but it's a very definitely very important and central part of any, should be of any AI research program.

1:00:55Yeah, which brings a whole, you know, next series of questions of where RL is efficient versus less. So it seems that it's been particularly good for like math and coding. And then the next obvious question is, what about the rest of the world? But taking a quick sort of going down the rabbit hole a little bit about math. So just in September, like just a few weeks ago, you guys did something unbelievable with the ISPC world finals. Do you want to talk about what that was and what went on from a model technical perspective behind the scenes? There happens surprisingly little from our perspective, from the model perspective.

1:01:39We just have a pretty smart model and then when we ask them to solve programming problems they are correct. and like what's what's a little bit of a backstory in it is that i think we use like specifically programming puzzles for a while as a very nice research test bed of our of our ideas those are those are nice problems to experiment on them and they they weren't ever like considered part of the product but it's it's those are pretty like complex problems and they require a whole bunch of thinking are very nice to give rewards to it so all of researchers just just liked working on those problems as a way of trying out their oral ideas like you always need a data set i'll i'll take a data set of programming puzzles and try it and i think i think like because of that a little bit our moles just were always very very good at competitive programming as a kind of byproduct we never we never tried to be good at it but like researchers were trying their ideas on it and because of it's like every every training run whatever whatever we are doing just end up being very, very good at those type of puzzles.

1:02:48And then it was a little bit of a formality for us to go and submit that to a competition. It's largely about demonstrating to the world what is the level of capability in those models. But I think it is important and true to acknowledge that in not all domains, at least like comparing to human baseline, at this moment we can be as nice and as good as in programming competition problems in many ways because those were like, you know, tried for a long time by like, you know, many, many researchers. And researchers don't always spend as much time as they could. and I would like them to uncover practical problems that people go to touch GPT or our models with.

1:03:41Great. So it sort of came out of the box. There was no specific training for it. And just to remind people, so what I'm referring to, what we were discussing, is the ICPC World Finals that was just in September 2025, which is International Collegiate Programming Contest, ICPC, that happened in Baku, Azerbaijan, where OpenAI solved 12 complex algorithmic problems within a five-hour time limit, basically making it take the equivalent of first place in front of human teams. So just for context. We did a little bit of a round tour of various competitions. We did ICPC, we did also like IOI, International Olympic Informatics earlier this year, and Adcoder Heuristics competition as well where we went second behind a single human that is also a Polish person that used to be employed by OpenAI some time ago.

1:04:36Funny coincidence but I think we were looking for a moment in time where we are, where our models are kind of smart enough to be able to compete in those competitions with some of incredibly smart and talented humans but it was never our particular goal and focus it's kind of like we think if we are doing good research for training smart models they should be smart enough to do those things and we kind of like take that milestone and now we keep on moving forward and i hope we will see and i think we are already seeing more and more practical and tangible things coming out like every week or every other week on twitter I do see what I think are credible reports of actual scientists using some of our reasoning models to help them perform calculations, solve hard technical problems with our models.

1:05:34And I think this is like where we want to be, like solving competitions is cool, but people solve competitions to prove that they can go to the actual frontier level job and solve new technical problems. and this is kind of what we want from our models as well. We alluded to this a second ago. So mentally, at least for somebody like me, I understand how RL could be used very effectively to train against math problems or coding problems. I think one of the big questions right now is how do you do that for the rest of the world in contexts and disciplines where the answer is not right or wrong, maybe a little more murky and the rest of the economy.

1:06:22Like you guys as an organization came up with a GDP Val the other day, which is a way of evaluating performance against different industries. What is your thinking in terms of generalization of RL as a path to success for the rest of the world? I think the short and quick answer is somehow humans can learn all those things. And as long as there is any way to evaluate performance and figure out if something is going right or wrong and you can compute that feedback, like you need to be able to somehow calculate how well something did, then you can optimize it. And then you can do reinforcement learning with it.

1:07:08I think there can be an argument that if there is no notion of what is right or what is wrong, then also humans are not able to improve and learn because there needs to be a learning signal that's coming somewhere. There is mostly a question of how convenient and how easy is it to get that feedback. And everyone doing reinforcement learning should strive and try to be able to train on more and more complex and interesting training signals to do. and like very often there comes a notion of what is often called reward hacking and like what's what happens when doing reinforcement learning it happens a lot and it's a it's an important problem you shape your reward in some way to reward certain behaviors but sometimes it is the case that's what your reward is not what actually you want there there is like one thing you need to train them all to do the behavior's reward but there there's also a natural like mismatch about the reward that you give them all and what you actually want.

1:08:09And there are sometimes moments where the mall does what you reward, but it's not in the spirit of what you had wanted and you need to fix it. It's almost like a parenting challenge. And in some way, you can say it's a limitation of reinforcement learning. But when I was thinking about it, I realized a lot of that happens in human systems as well. there are a lot of like incentive system and reward systems and even even happens in workplaces in all kind of human groups that humans have rewards that are not always optimized for the for the ultimate goals of the system and they and they hack rewards constantly in many different ways and there is a constant whack-a-mole game between between setting the right rewards and seeing if the system does it and it's a huge like issue in any policy making almost and any incentives programs.

1:09:02And this is the same kind of like whack-a-mole game in reinforcement learning research, trying to make sure your rewards are better and better representing what you actually care about the model to be doing. All right. So maybe to zoom out, to close this conversation, you said the other day, you tweeted, we all collectively believe AGI should have been built yesterday. And the fact that it hasn't yet is mostly because of a simple mistake that needs to be fixed, which is super awesome as a tweet. Do you think that the combination of pre-training and scaled RL texts us to AGI? There's always an interesting question of like, what do we consider something that is not pre-training on RL?

1:09:55And like, where is the limit? I generally think something that we are doing like pre-training today is necessary. I think something that we are doing RL today is necessary. And there will surely be a few things more. And we have a lot of very ambitious research programs on some of those things. And I don't think the question of distance in research space is hard to say. like with some for some people like what we what we are what we want to do and what we are planning to build is not very far from those things for someone will say oh it's completely different and it's like a very much not that so i don't want to go into into like debates whether it's the same or not but like we are and want to be constantly changing the way how we train the models to more represent what we think the the right form of intelligence is and the most useful useful form of learning is and constantly are researching various things.

1:10:55And then the distance from like, what is the distance from AGI is also like a very complex question. I really like someone said it to me, but I think it is right that, you know, if you talk to someone from 10 years ago and show them chat GPT from today, they would probably call it AGI, but we are not today because it still has a lot of limitations and we are all very aware of those limitations and we are pretty sure we can resolve those limitations. there will be probably some further limitations of the future models that that will need to be fixed there is an ultimate question which is very hard to answer when when is the moment that the model can like improve itself without that much like external output and without humans working on it and fixing it and i think i think it's a it is a very hard question it is a serious question that like you know we need to try to answer humanity needs to need to try tries to answer because because like you know almost that one will still like largely depend on um on like our infrastructure and our systems but we'll be able to start start fixing itself without without us having to fix it and that's that's like the predictions of what really ai will be able to do and we'll be able to solve it that moment start becoming like a little bit murkier than what we can do right now, which I think we still can do that pretty well.

1:12:19But philosophically, you may have heard Richard Sutton, you know, the other day on the Dworkish podcast, which is a wonderful episode that people should really listen to, effectively saying that the only path to AGI was going to be pure RL and that fundamentally LLMs, and I hope I'm characterizing what he said appropriately, but that LLMs were a flawed premise because effectively that was imitation of reality whereas RL was enforcement of reality. Do you have any thoughts sort of like philosophically on that question? Yeah, I haven't had a chance to fully listen to that episode yet so I also don't get all the details of that thought but what I can say is that we are doing quite serious RL on language models these days and like i don't like in terms of a pure rl i don't think like really pure rl makes sense rl needs pre-training to be successful and i think pre-training as i said before needs rl to be to be successful as well i don't think without rl it would make sense that the research program we are we are doing but but we are like open ai is and i'm pretty sure all other all other al labs as well very serious about about doing doing a lot of reinforcement learning on our models and i think like what what what kind of like a lot of people are saying that that whether llms are an on-ramp or off-ramp of agi very often they they do mean like pre-training but it's also clear that that like you know the current way how we are doing things like also it's not yet enough and it's not yet everything and there will need to be a further further changes to to the setup but sometimes like people say oh if you are if you are doing rl it's not llm it's something else sometimes say oh if you can write program in your in your rollout and it's a chain of thought it's not an this is a neural network only it's a neural symbolic system so it's like you know it's easy to get like some people consider something in llm and the other thing not uh but but personally my view is what we have is a pretty good foundation for the next step of like we did have transformers first train for transfer translation then we were then we were pre-training them on large-scale data then we were doing our lhf on them now we are doing large-scale reinforcement learning we'll do a few more uh more and more complex things there is a chance somewhere along the line the architecture will start changing more or less significantly and and you know i think i personally think we are on the right path and it will feel less like completely turning around and more like keep on adding more things and maybe dissolving some old elements that carried us to that particular level of intelligence and were not needed anymore.

1:15:20Well, that feels like a wonderful place to leave it. You've been very generous with your time and thoughts and giving us a glimpse into to open AI, what you work on, what it looks like behind the scenes and the key aspects of pre-training and scaling reinforcement learning. So it's been a wonderful conversation. Jerry, thank you so much. Really appreciate it. Thank you very much. I enjoyed being here a lot too. Hi, it's Matt Turk again. Thanks for listening to this episode of the Matt Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from.

1:15:58This really helps us build a podcast and get great guests. Thanks and see you on the next episode.

From the publisher

What does it really mean when GPT-5 “thinks”? In this conversation, OpenAI’s VP of Research Jerry Tworek explains how modern reasoning models work in practice—why pretraining and reinforcement learning (RL/RLHF) are both essential, what that on-screen “thinking” actually does, and when extra test-time compute helps (or doesn’t). We trace the evolution from O1 (a tech demo good at puzzles) to O3 (the tool-use shift) to GPT-5 (Jerry calls it “03.1-ish”), and talk through verifiers, reward design, and the real trade-offs behind “auto” reasoning modes.


We also go inside OpenAI: how research is organized, why collaboration is unusually transparent, and how the company ships fast without losing rigor. Jerry shares the backstory on competitive-programming results like ICPC, what they signal (and what they don’t), and where agents and tool use are genuinely useful today. Finally, we zoom out: could pretraining + RL be the path to AGI?


This is the MAD Podcast —AI for the 99%. If you’re curious about how these systems actually work (without needing a PhD), this episode is your map to the current AI frontier.



OpenAI

Website - https://openai.com

X/Twitter - https://x.com/OpenAI


Jerry Tworek

LinkedIn - https://www.linkedin.com/in/jerry-tworek-b5b9aa56

X/Twitter - https://x.com/millionint


FIRSTMARK

Website - https://firstmark.com

X/Twitter - https://twitter.com/FirstMarkCap


Matt Turck (Managing Director)

LinkedIn - https://www.linkedin.com/in/turck/

X/Twitter - https://twitter.com/mattturck



(00:00) Intro

(01:01) What Reasoning Actually Means in AI

(02:32) Chain of Thought: Models Thinking in Words

(05:25) How Models Decide Thinking Time

(07:24) Evolution from O1 to O3 to GPT-5

(11:00) Before OpenAI: Growing up in Poland, Dropping out of School, Trading

(20:32) Working on Robotics and Rubik's Cube Solving

(23:02) A Day in the Life: Talking to Researchers

(24:06) How Research Priorities Are Determined

(26:53) Collaboration vs IP Protection at OpenAI

(29:32) Shipping Fast While Doing Deep Research

(31:52) Using OpenAI's Own Tools Daily

(32:43) Pre-Training Plus RL: The Modern AI Stack

(35:10) Reinforcement Learning 101: Training Dogs

(40:17) The Evolution of Deep Reinforcement Learning

(42:09) When GPT-4 Seemed Underwhelming at First

(45:39) How RLHF Made GPT-4 Actually Useful

(48:02) Unsupervised vs Supervised Learning

(49:59) GRPO and How DeepSeek Accelerated US Research

(53:05) What It Takes to Scale Reinforcement Learning

(55:36) Agentic AI and Long-Horizon Thinking

(59:19) Alignment as an RL Problem

(1:01:11) Winning ICPC World Finals Without Specific Training

(1:05:53) Applying RL Beyond Math and Coding

(1:09:15) The Path from Here to AGI

(1:12:23) Pure RL vs Language Models

More from The MAD Podcast with Matt Turck

All 44 episodes
How GPT-5 Thinks — OpenAI VP of Research Jerry TworekThe MAD Podcast with Matt Turck · 1 h 16 min
Listen in VO