OpenAI's Noam Brown, Ilge Akkaya and Hunter Lightman on o1 and Teaching LLMs to Reason Better

2 Oct 2024 · 45 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Summary: OpenAI's Noam Brown, Ilge Akkaya, and Hunter Lightman on o1 and Teaching LLMs to Reason Better

Overview In this episode of *Training Data*, hosts Sonya Huang and Pat Grady engage with OpenAI researchers Noam Brown, Ilge Akkaya, and Hunter Lightman to discuss *o1* (also known as Strawberry), a model that combines large language models (LLMs) with AlphaGo-style deep reinforcement learning (RL). The episode explores the development, capabilities, and implications of *o1* as it pertains to reasoning, problem-solving, and various benchmarks.

Episode Highlights

Introduction

  • Time: 00:00 - Introduction to the episode and guests.

Conviction in o1

  • Time: 01:33 - The team discusses their initial skepticism and eventual confidence in the model's capabilities.

Reasoning and How o1 Works

  • Time: 04:24 - The discussion delves into the definition of reasoning and how *o1* utilizes chains of thought and backtracking to solve problems.

Lessons from Gameplay

  • Time: 07:02 - Comparisons are made between the gameplay of AlphaGo and the reasoning strategies of *o1*.

Generation vs. Verification

  • Time: 09:14 - The conversation touches on the generator-verifier gap and its implications for various tasks.

o1's Performance

  • Time: 10:31 - The researchers reveal what has surprised them about *o1* thus far, including its strengths in STEM subjects.

Implications of Inference-Time Scaling Laws

  • Time: 28:39 - Discussion around the scaling laws for computational power during inference and the model's future capabilities.

The Importance of Reasoning

  • Time: 25:29 - Emphasizes how crucial reasoning is for achieving AGI (Artificial General Intelligence) and accomplishing economically valuable tasks.

o1-mini

  • Time: 41:13 - Introduction of *o1-mini*, a smaller variant designed for applications requiring reasoning without extensive world knowledge.

Founders' Perspectives

  • Time: 42:15 - Guidance for founders on when to utilize *o1* versus other models like GPT-4, and insights into the future trajectory of *o1* and its iterations.

Key Concepts and Discussions

Reasoning

  • Defined as the ability to think through problems and consider multiple options, often benefiting from extended thinking time.

AlphaGo and Deep Reinforcement Learning

  • The success of AlphaGo has set a precedent for deep reinforcement learning's application in broader problem-solving contexts, particularly in *o1*.

Challenges and Expectations

  • The researchers acknowledge the challenges in achieving AGI, particularly the gap between the model's capabilities and human reasoning.

Human-AI Collaboration

  • Insights into how researchers have begun utilizing *o1* as a brainstorming partner, demonstrating its potential as a collaborator in fields like cancer research.

Future of AI and AGI

  • Discussion on what defines AGI and how *o1* represents a step towards more generalizable reasoning capabilities, with a focus on scaling inference-time compute to unlock further potential.

Conclusion The episode concludes with a reflection on the early days of *o1*, the excitement around its potential applications, and the anticipation of future developments in AI reasoning capabilities. The researchers express a commitment to iteratively improve their models based on user feedback and real-world applications.

Mentioned Resources

  • Learning to Reason with LLMs: [Technical Report](#)
  • Agent57: Outperforming the human Atari benchmark (DeepMind, 2020)
  • The Last Question by Isaac Asimov
  • Various benchmarks and competitions including the IOI (International Olympiad in Informatics)

---

This summary encapsulates the primary discussions, concepts, and insights from the podcast episode, providing a clear understanding of the themes surrounding *o1* and its implications for the future of AI.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00One way to think about reasoning is there are some problems that benefit from being able to think about it for longer. You know, there's this classic notion of system one versus system two thinking in human system one is the more automatic, instinctive response and system two is the slower, you know, more process driven response. And for some tasks, you don't really benefit from more thinking time. So if I ask you like, what's the capital of bootcamp, you know, you could think about it for two years. It's not going to help you get it get it right with higher accuracy. What is the capital of bootcamp?

0:32I actually don't know. But you know, there's some problems where there's clearly a benefit from being able to think for longer. So one classic example that I point to is the Sudoku puzzle. It's you you could in theory just go through a lot of different possibilities for like what the Sudoku puzzle might be. what the solution might be, and it's really easy to recognize when you have the correct solution. So in theory, if you just had like tons and tons of time to solve a puzzle, you would eventually figure it out.

1:13We're excited to have Noam Hunter and Ilgo with us today, who are three of the researchers on Project Strawberry or O1 at OpenAI. O1 is OpenAI's first major foray into general inference time compute, and we're excited I'd like to talk to the team about reasoning, chain of thought, inference time scaling loss, and more. I'll go hunter and gnome. Thank you so much for joining us and congratulations on releasing a one into the wild. I want to start by asking, did you always have conviction this is going to work?

1:46I think that we had conviction that something in this direction was promising, but the actual path to get here was never clear. And I, you know, you look at all one, it's not like this was an overnight day. It actually there's a lot of years of research that goes into this. And a lot of that research didn't actually pan out. But I think that there was conviction from OpenAI and like a lot of the leadership that something in this direction had to work. and they were willing to keep investing in it despite the initial setbacks. And I think that eventually paid off. I'll say that I did not have as much conviction as known from the very beginning.

2:31I've been staring at language models trying to teach them to do math and other kinds of reasoning for a while. And I think there's like a lot into research that's ebb and flow. Sometimes things work, sometimes things don't work. when we saw that the methods we were pursuing here started to work. I think it was a kind of aha moment for a lot of people, myself included, where I started to read some outputs from the models that were approaching problem solving in a different way. And that was this moment, I think, for me, where my conviction really set in. I think that OpenAI in general takes a very empirical data driven approach to a lot of these things.

3:12and when the data starts to speak to you, when the data starts to make sense, when the trends start to line up and we see something that we wanna pursue, we pursue it and that for me was when, I think the conviction really set in. Well, that you, Elga, you've been at OpenAI for very long time, five and a half years. Five and a half years. What do you think, did you have conviction from the beginning that this approach was gonna work? No, I've been wrong several times since joining about the path to AGI. We originally thought that robotics was the way forward. That's why I joined robotics team first.

3:47Embodied AI, AGI, that's where we thought things were going to go. But, yeah, I mean things hit roadblocks. I would say like during my time here, Chachi PT, well, I guess that's kind of obvious now that was a paradigm shift. We were able to share very broadly with the world, something that is a universal interface. And I'm glad that now we have a new path potentially forward to push this reasoning paradigm. Yeah, it was definitely not obvious to me for the longest time. Yeah. I realize there's only so much that you're able to say publicly for very good reasons about how it works. But what can you share about how it works even in sort of general terms?

4:36So the all one model series are trained with Aral to be able to think and you could call it's reasoning maybe also and it is fundamentally different from what we're used to with LLM's and we've seen it really generalize to a lot of different reasoning domains as we've also shared recently. So we're very excited about this paradigm shift with this new model family. And for people who may not be as familiar with what stated the art in the world of language models today, what is reasoning? How would you define reasoning? And maybe a couple words on what makes it important? Good question. I mean, I think one way to think about reasoning is there are some problems that benefit from being able to think about it for longer.

5:26You know, there's this classic notion of system one versus system two thinking in human, system one is the more automatic, instinctive response, and system two is the slower, more process -driven response. And for some tasks, you don't really benefit from more thinking time, so if I ask you, what's the capital of Bhutan? You could think about it for two years. It's not going to help you get it right with higher accuracy. What is the capital of Bhutan? I actually don't know. But there's some problems where there's clearly a benefit from being able to think for a longer. So one classic example that I point to is the Sudoku puzzle.

6:00It's, you could, in theory, just go through a lot of different possibilities for like, what the Sudoku puzzle might be, what the solution might be. And it's really easy to recognize when you have the correct solution. So in theory, if you just had like tons and tons of time to solve a puzzle, you would eventually figure it out. And so that's what I consider to be, I think a lot of people in the AI community have like different definitions of reasoning. And I'm not claiming that this is like the conical one. I think everybody has their own opinions, but I view it as the kinds of problems where there is a benefit from being able to consider more options and think for longer.

6:37You might call it like a generator verifier gap where there's like, it's really hard to generate a correct solution, but it's much easier to recognize when you have one. And I think all problems exist on this spectrum from really easy to verify relative to generation, like a Sudoku puzzle versus just as hard to verify as it is to generate a solution, like name of the capital of Gutan. I wanna ask about AlfaGo and know your background having done a lot of great work and poker and other games. To what extent are the lessons from gameplay analogous to what you guys have done with O1 and how are they different?

7:18So I think one thing that's really cool about O1 is that it does clearly benefit by being able to think for longer. And when you look back at like many of the AI breakthroughs that have happened, I think AlphaGo is the classic example. One of the things that was really noticeable about the bot though I think underappreciated the time was that it thought for a very long time before acting. It would take you know 30 seconds to make it to make a move. And if you tried to have it act instantly, it actually wasn't better than top humans. It was noticeably worse than them. And so it clearly benefited a lot by that extra thinking time.

7:56Now, the problem is that the extra thinking time that it had, it was running multiple outreach research, which is like a particular form of reasoning that worked well for go. But for example, doesn't work in a game like poker, which my early research was on. And so a lot of the like methods that existed for being able to reason, for being able to like think for longer, was still specific to the domains, even though the neural nets behind it, the system one part of the AI was very general. And I think one thing that's really cool about O1 is that it is so general. The way that it's thinking for longer is actually quite general and can be used for a lot of different domains and we're seeing that by giving it to users and seeing what they are able to do with it.

8:48Yeah, one of the things that's always been really compelling to me about language models and this is nothing new is just that because their interface is the text interface, they can be adapted to work on all different kinds of problems. And so what's exciting, I think, about this moment for us is that we think we have a way to do something, to do reinforcement learning on this general interface. And then we're excited to see what that can lead to. One question on that. You mentioned, I thought that was well put, sort of the, the, I forget exactly how you phrased it, but the gap between generation and verification, there's sort of a spectrum in terms of how easy things are to verify.

9:25does the method for reasoning remain consistent at various points in that spectrum or are there different methods that apply to various points in that spectrum? One thing I'm excited about for this release has been to get O1 in the hands of so many new people to play with it, to see how it works, what kinds of problems it's good at and what kinds of problems it's bad at. I think this is like something really core to open AI's strategy of iterative deployment. we put the technology that we build, the research that we develop out into the world so that we can see We do it safely and we do it so that we can see how the world interacts with it and what kinds of things we might not always understand fully ourselves And so in thinking about like what are the limits of our approaches here?

10:11I think it's been really enlightening to see like Twitter Show what it can and what it can't do I hope that that is like enlightening for the world that's useful for for everyone to figure out what these new tools are useful for and then I also hope where it will take back that information and use it effectively to understand our processes, our research, our products better. Speaking of which, is there anything in particular that you all have seen in the Twitter verse that surprised you, you know, ways that people have figured out how to use O1 that you hadn't anticipated? There's one thing I'm super excited about.

10:45I've seen a lot of MDs and researchers use the model as a brainstorming partner and what they are talking about is that like they've been in cancer research for so many years and they've been just running these ideas by the model about what they can do about these gene discovery gene therapy type of applications and they are able to get like these really novel ways of research to pursue from the model. Clearly the model cannot do the research itself but it can just be a very nice collaborator with humans for in this respect. So I'm super excited about about seeing the model just advanced a scientific path forward.

11:24That's not what we're doing in our team, but that is the thing, I guess, like we want to see in the world, the domains that are outside hours that gets really benefit by this model. No, I think you tweeted that deep RL is out of the trough of disillusionment. Can you say more about what you meant by that? I mean, I think there is definitely a period starting with I think Atari, you know, the Deep Mine Atari results, where Deep RL was the hot thing. I mean, I was in a PhD program. I remember what it was like in like, you know, 2015 to 2018, 2019, and Deep RL was the hot thing. And in some ways, I think that was, I mean, a lot of research was done, but certainly some things were overlooked.

12:10And I think one of the things that was kind of looked was the power of just training on tons and tons of data using, you know, something like the GPT approach. And in many ways, it's kind of surprising because if you look at AlphaGo, which was in many ways like the the crowning achievement of DeepRL, yes, there was this RL step, but there was also, I mean, first of all, there's also this reasoning step. But even before that, there was this large process of learning from human data. And that's really what got AlphaGo off the ground. And so then there was this like increasing shift. There is, I guess like a view that this was an impurity in some sense that so a lot of DeepRL is really focused on learning without human data, with just learning from scratch.

13:00Yeah, Alpha Zero, which was a great, which was an amazing result and actually ended up doing a lot better than AlphaGo. But I think probably because of this focus on Looney for Scratch, this GPT paradigm kind of flew under the radar for a while. And except for OpenAI, which saw some initial results for it and, you know, again, had the conviction to double down on that investment. Yeah, so there was definitely this period where DeepRL was the hot thing. And then I think, you know, when GPT -3 came out and some of these other like large language models. And there was so much success without deep RL.

13:44It there was like, yeah, a period of disillusionment where a lot of people switched away from it or kind of lost faith in it. And what we're seeing now with O1 is that actually there is a place for it and it can be quite powerful when it's combined with these other elements as well. And I think a lot of the DPRL results were in kind of well -defined settings like gameplay. Is O1 one of the first times that you've seen DPRL used in much more general kind of unbounded setting? Is that the right way to think about it? Yeah, I think it's a good point that a lot of the highlight DPRL results were really cool, but also very narrow in their applicability.

14:28I mean, I think there were a lot of quite useful deep RL results and also quite general RL results, but there wasn't anything comparable to something like GPT -4 in its impact. So I think we will see that kind of level of impact from deep RL in this new paradigm going forward. One more question in this general train of thought. I remember the AlphaGo results, you know, at some point with the in the lease of the old Tournament, there was move 37 and that move surprised everybody. Have you seen something of that sort where O1 tells you something and it's surprising and you think that it's actually right and it's better than any top human good think of?

15:09Have you had that moment yet with the model? Or you think it's O2, O3? One of the ones that comes to mind is we spent a lot of the time preparing for the IOI competition that we put the model into. looking at its responses to programming competition problems. And there was one problem where it was, someone was really insistent on solving the problem in this kind of weird way, with some weird method, I don't know exactly what the details were. And our colleagues who are much more into competitive programming were trying to figure out why it was doing it like this. I don't think it was quite a, like, this is a stroke of genius moment.

15:44I think it was just like the model didn't know the actual way to solve it. And so it just like being ahead until it found something else. Did it get there? Yeah, yeah, it solved the problem. It just it just it used some it was like it was some method that would have been really easy if you saw something else I wish I had the specific one, but I remember that being kind of interesting. There's a lot of the things in the in the in the programming competition results I think somewhere we have the IOI competition programs published Where you can start to see that the model doesn't approach thinking quite like a human does, or doesn't approach these problems quite like a human does it has slightly different ways of solving it for the actual IOI competition.

16:23There was one problem that humans did really pour on that the model was able to get half credit on. And then another problem that humans did really well on that the model was like barely able to get off the ground on. Just showing that it kind of has a different way of approaching these things than then maybe a human would. I've seen the model solve some geometry problems and the way of thinking was quite surprising to me such that you're asking the model just like give me this like sphere and then there are some points on the sphere and asking for probability of some event or something and the model would go let's visualize this, let's put the points and then if I think about it that way or something so I'm like oh you're just using words and visualizing something that really helps you contextualize.

17:09I would do that as a human and seeing all one do it too, just really surprises me. Interesting. That's fascinating. So it's stuff that's actually understandable to a human and would actually kind of expand the boundaries about humans would think about problems versus some undecisurable machine language. That's really fascinating. Yeah, I definitely think one of the cool things about our own result is that these chains of thoughts the model produces are human interpretable. And so we can look at them and we can kind of poke around how the model is thinking. Were there, were there aha moments along the way?

17:42Or were there moments where, you know, Hunter, you mentioned that you were not as convinced at the outset that this is the direction that was going to work? Was there a moment when that changed where you said, oh my gosh, this is actually going to work?

17:56Yeah, so I've been in an open air about two and a half years and most of the time I've been working on trying to get the models better at solving math problems. And we've done a bunch of work in that direction. We've built various different the bespoke systems for that. And there was a moment on the O1 trajectory where we had just trained this model with this method with a bunch of fixes and changes and whatnot. and it was scoring higher on the Matthew Vals than any of our other attempts. Any of our bespoke systems and then we were reading the Change of Thought and you could see that they felt like they had a different Character in particular.

18:34You could see that when it got stuck it would say wait this is wrong Let me take a step back. Let me figure out the right path forward. And we called this backtracking and I think for a long time I'd been waiting to see in instance of the models backtracking and I kind of felt like I wasn't going to get to see an other regressive language model backtrack because they're just kind of predict next token, predict next token, predict next token. And so when we saw this score on the math test and we saw the trajectory that had the backtracking, that was the moment for me where I was like, wow, this is like something is coming together.

19:07I didn't think was going to come together and I need to update. And I think that was when I grew a lot of my conviction. I think the story is the same for me. I think it was I've probably run the same time actually. Like, I, you know, I definitely, I joined with this idea of like, you know, chat to BT doesn't really think before responding. Like, it's very, very fast. And there was this like powerful paradigm of like, in these, in these games of AI is being able to think for longer and getting much better results. But, and there's a question about how do you bring that into language models that I was really interested in.

19:41And you know, that's like, it's easy to say that, but then there's like, there's this difference between just like saying that, oh, there should be a way for it to think for longer than actually like delivering on that. And so we, you know, we, I tried to, I tried a few things and like other people were trying a few different things. And in particular, yeah, one of the things we wanted to see was this ability to backtrack or to like recognize when it made a mistake or to like try different approaches. And we had a lot of discussions around how do you enable that kind of behavior? And at some point, we just felt like, okay, well, one of the things we should try, this is a baseline, is just to have the AI thing for longer.

20:21And we saw that, yeah, once it's able to, to think for longer, it develops these abilities, almost like, emergently, that we're very powerful, and contain things like backtracking and self correction, and all these things that we were wondering how to enable in the models. And to see it come from such a clean, scalable approach, that was for me the big moment when I was like, okay, it's very clear that we can push this further. And it's so clear to see where things are going. No, my thing is understanding how strong and effective his conviction and test time compute was. I feel like all of our early one on ones when he joined we're talking about test time computer and it's power and I think multiple points throughout the project known would just say, why don't we let the model think for longer?

21:12And then we would and it would get better and he would just be, he would just look at us kind of funny like we hadn't done it until that. One thing we noticed in your e -vails is that, you know, O1 is noticeably good at STEM. It's better at than the previous models. Is there a rough intuition for that? Why that is? I mentioned before that there's some tasks that are reasoning tasks that are easier to verify than they are to generate a solution for, and there's some tasks that don't really fall into that category. And I think STEM problems tend to fall into the what we would consider hard reasoning problems.

21:48And so I think that's a big factor for why we're seeing a lift on STEM kind of subjects. Makes sense. I think relatedly we saw that in the research paper that you guys released, that 01 passes your research engineer interview with pretty high pass rates. What do you make of that? And does that mean at some point in the future, OpenAI will be hiring 01 instead of human engineers? I don't think we're quite at that level yet. I think that there's more. It's hard to be the 100 % though. Maybe the interviews need to be better. Okay. I think that the 0 -1 does feel at least to me, I think other people on our team like a better coding partner than the other models.

22:35I think it's already authored a couple of PRs in our repo. And so in some ways it is acting like a software engineer. Because I think software engineering is one of the STEM domains that benefits from longer reasoning. I don't know, I think that the kinds of rollouts that we're seeing from the model are thinking for a few minutes at the time. I think the kinds of software engineering job that I do when I go on right code, I think for more than a few minutes at a time. And so maybe as we start to scale these things further, as we start to follow this trendline and let O1 think for longer and longer, it'll be able to do more and more of those tasks.

23:12And we'll see. you'll be able to tell that we've achieved AGI internally when we take down all the jobless things. And you know, the company's doing really well, really poorly. What do you think it's going to take for a one to get great at the humanities? Do you think being good at raising and logic and STEM kind of naturally will extend to being good at the humanities? As you scale up in front of time, or how do you think that plays out? You know, we're like we said, we released the models and we were kind of curious to see what they were good at and what they weren't as good at and what people end up using it for.

23:47And I think there's clearly a gap between the raw intelligence of the model and how useful it is for various tasks. In some ways it's very useful, but I think that it could be a lot more useful and a lot more ways. And I think there's still some iterating to do to be able to unlock that. like more general usefulness. Well, I'm kind of scone that. Do you view, I'm curious if there's a philosophy at OpenAI or maybe just a point of view that you guys have on, how much of the gap between the capabilities of the model and whatever real world job needs to be done? How much of that gap do you want to make part of the model?

24:28And how much of that gap is sort of the job of the ecosystem that exists on top of your APIs, like their job to figure out? Do you have a thought process internally for figuring out what are the jobs to be done that we want to be part of the model versus where do we want our boundaries to be so that there's an ecosystem that exists around us? I always heard that opening I was very focused on EGI. I was honestly kind of skeptical of that before I joined the company. Basically the first day that I started, there was an all hands of the company and Sam got up in front of the whole company and basically laid out the priorities going forward for the short -term and the long -term, it became very clear that AGI was the actual priority.

25:14And so I think the clearest answer to that is AGI is the goal. There's no single application that is the priority other than getting us to AGI. Do you have a death mission for AGI? Everybody has their own definition for you. Exactly. That's what I'm curious. I don't know if I have a concrete definition. I just think that it's something about the proportion of economically valuable jobs that are models and our AS systems are able to do. I think it's going to ramp up a bunch over the course of the next however many years. I don't know. It's one of those, you'll feel it when you feel it and will move the goalpost back I can be like, this isn't that for however long until one day we're just working alongside these AI coworkers and they're doing large parts of the jobs that we do now and we're doing different jobs and the Holy Co system of what it means to do work has changed.

26:14One of your colleagues had a good articulation of the importance of reasoning on the path to AGI, which I think paraphrases as something like any job to be done is going to have obstacles along the way. And the thing that gets you around those obstacles is you're going to reason through them. And I thought that was like a pretty nice connection between the importance of reasoning and the objective of AGI and sort of being able to accomplish economically useful tasks. Is that is that the best way to think about what reasoning is and why it matters or there are other frameworks that you guys tend to use?

Read the full transcript

26:54I think this is a TBD thing just because I think it a lot of the stages of the development of these AI systems of these models we've seen different shortcomings, different failings of them. I think we're learning a lot of these things as we develop the system, as we evaluate them, as we try to understand their capabilities and what they're capable of. Other things that come to mind that I don't know how they relate to reasoning or not are like strategic planning, ideating, or things like this, where to be a, to make an AI model that's as good as an excellent product manager, you need to do a lot of brainstorming ideation on, what users need, what all these things are, is that reasoning, or is that a different kind of creativity that's not quite reasoning, it needs to be addressed differently, then afterwards, when you think about operationalizing those plans into action, you have to strategize about how to move an organization towards getting things done, is that reasoning, there's parts of it that are probably reasoning, and then there's maybe parts that are something else, and maybe eventually it'll all look like reasoning to us, or maybe we'll come up with a new word, and there'll be new steps we need to take to get there.

28:01I don't know how long we can, we'll be able to push this forward, but whenever I think about this general reasoning problem, it helps to think about the domain of math. We've spent a lot of time reading what the model is thinking, when you ask it a math problem. And then it's clearly doing this thing where like it hits an obstacle and then it back tracks just has a problem all the way. It's maybe I should try this other thing. So when you see that's thinking process, you can imagine that it might generalize the things that are beyond math. That's what gives me hope. I don't know the answer, but hopefully.

28:39The thing that gives me pause is that the O1 is already better than me at math. But it's not as good at me as being a software engineer. And so there's some there's some mismatch here. There's still a job to be done. Yeah. There's still some work to do. Yeah. If my whole job were doing a me problems and doing high school competition, I'd be out of work. There's still some stuff for me for right now. Since you mentioned sort of the the chain of thought being able to watch the reasoning behind the scenes, I have a question that might be one of those questions you guys can't answer. But just for fun.

29:08Was it for first off, I give you props for in the blog that you you guys published with the release of a one, explaining why chain of thought is actually hidden and literally saying like partly it's for competitive reasons. I'm curious if that was a contentious decision or like how controversial that decision was because I could see it going either way and it's a logical decision to hide it but I could also imagine a world in which you decide to expose it. So I'm just curious if that was a contentious decision. I don't think it was contentious. I mean, I think for the same reason that you don't want to share the model weights necessarily for a frontier model.

29:44I think there's a lot of risks to sharing the thinking process behind the model. I think it's a similar decision actually. Can you explain from a layman's first maybe to a layman like what is the chain of that and what's the example of one? So for instance, if you're asked to solve an integral, most of us would need a piece of paper and a pencil and we would kind of lay out the steps from getting from a complex equation and then there will be steps of simplifications and then going to a final answer. The answer could be one, but how do I get there? That is the train of thought in the domain of math.

30:24Let's talk about that path forward. Inference time, scaling laws. To me, that was the most important chart from the research that you guys published and It seems to me like a monumental result similar to the scaling loss from from pre -training and Sorry to be happy Do agree that like the implications here. I think they're pretty profound and you know, what does it mean for for the field as all? I think it's pretty profound and I think One of the things that I wonder when we were preparing to release a one is whether people would recognize its significance We included it, but it's kind of a subtle point.

31:06And I was actually really surprised and impressed that so many people recognized what this meant. There have been a lot of concerns that like, A, I might be hitting a wall or plateauing because pre -training is so expensive and becoming so expensive. And there's all those questions around, like is there an update to train on? And I think one of the major takeaways about O1, especially on preview is not what the model is capable of today, but what it means for the future. The fact that we're able to have this different dimension for scaling that is so far pretty untapped. I think is a big deal and I think means that the ceiling is a lot higher than a lot of people have appreciated.

31:53What happens when you let the model things for hours or months or years What do you think happens? We haven't had a one for years. So we haven't been able to let it think that long yet. Is there a job just running in the background right now? It's just just just still thinking about. Solve world peace. Okay, I'm thinking. Thinking. Yeah, there's a there's a awesome of story like that called the last question. Where we are you they asked this big computer sized AI. Something about like how do we reverse entropy? And it says I need to think longer for that and like the story goes and then a 10 years later They see and it's still thinking and then a hundred years later than a thousand years later and then 10 thousand years later Yeah, there is as yet meaningful and not enough information for me NASA or something like yeah, like it's still yeah Do you have a guess empirically on you know what'll happen?

32:49What you know or I guess right now I think the model has I've seen some reports like 120 IQ so like very very smart is there a ceiling on that as you scale up in friends time confused? Do you think you get to infinite IQ? One of the important things is that like, it's 120 IQ on some tests someone gave. This doesn't mean that it's got like 120 IQ level reasoning at all the different domains that we care about. I think we even talk about how it is below 40 on some things like creative writing and whatnot.

33:24So there's definitely, it's like, it's confusing to think about how we extrapolate this model. I think it's an important point that, you know, we talk about these benchmarks and we, one of the benchmarks that we highlighted in our results was GPQA, which is this, you know, questions that are given to PhD students and like to be going to PhD students can answer. and the AI is outperforming a lot of PhDs on this benchmark right now. That doesn't mean that it's smarter than a PhD in every single way imaginable. There's a lot of things that a PhD can do that there's a lot of things that a human can do, a period that the AI can't do.

33:59And so you always have to look at these e -vals with some understanding that it's measuring a certain thing that is typically a proxy for human intelligence when he measured, you know, when humans take that test, but means like different when the AI takes that test. Maybe a way of framing that as an answer to the question is that I hope that we can see that letting the model think longer on the kinds of things that it's already showing it's good at will continue to get it better. So one of my big Twitter moments was I saw a professor that I had had in school, a math professor was tweeting about how he was really impressed with the one because he had given it a proof that had been solved before by humans, but never by an AI model.

34:44And it just took it and ran with it and figured it out. And that to me feels like we're at the cost of something really interesting where it's close to being a useful tool for doing novel math research, where if it can do some small lemma's in some proofs for like real math research, that would be really, that would be really really a breakthrough. And so I hope by letting it think longer we can get better at that particular task of being a really good math research assistant It's harder for me to extrapolate what it's gonna look like Well, we'll get better the things that it's not good at now What would that path forward look like and then what would the infinite IQ or whatever look like then?

35:21When it thinks forever on problems that it's not good at but instead I think you can kind of ground yourself in a Here are the problems. It's good at if we let it think longer at these Oh, it's going to be useful for math research Oh, it's going to be really useful for software engineering. Oh, it's going to be really and you can start to play that game and start to see how I hope the future will evolve. What are the bottlenecks to scaling test time compute? I mean, for pre training, it's pretty clear you need enormous amounts of compute. You need enormous amounts of data. The stuff requires enormous amounts of money.

35:51I guess it's pretty easy to imagine the bottlenecks on scaling, pre training. What constraints sort of the scaling of inference time compute? When GPD 2 came out and GPD 3 came out, it was pretty clear that if you just throw more data and were GPUs at it, it's going to get a lot better. It still took years to get from GPD 2 and GPD 3 and GPD 4. There's just a lot that goes into taking an idea that sounds very simple and scaling it up to a very large scale. I think that there's a similar challenge here. where, okay, it's like a simple idea, but there's a lot that work that has to go into actually scaling it up.

36:34So I think that's the challenge. Yeah, I think that one thing that I think maybe doesn't any more surprise, but one thing I think might have used to surprise more academic oriented researchers who join OpenAI is how much of the problems we solve are engineering problems versus research problems. Building large -scale systems, training large -scale systems, running algorithms that have never been invented before on systems that are brand new. As skill no one's ever thought of is really hard and so there's always a lot of just like hard engineering work to make these systems scale up. Also one needs to know what to test the model on.

37:12So we do have these standard e -vals as benchmarks but perhaps there are ones that we are not yet testing the model on. So we're definitely looking for those where we can just spend more compute on test time and get better results. One of the things I'm having a hard time wrapping my head around is, you know, what happens when you give the model, you know, near infinite computes, because as a human, I am, you know, even if I'm Terrence Tao, like I am limited at some points by my brain, whereas you can just put more and more compute at inference time. And so does that mean that, for example, all math theorems will eventually be solvable through this approach or like where is the limit do you think?

37:55Infinite compute is a lot of compute. Near infinite. It goes back to the Asimov story if you're waiting 10 ,000 years but maybe. But I say that just to ground it in a like we don't know yet quite what the scaling of this is for how it relates to solving really hard math theorems. It might be that you really do need to let that think for a thousand years to solve some of the unsolved core math problems. Yeah, I mean, I think it is true that if you let it think for long enough, then in theory, you could just go through, you formalize everything in lean or something and you go through every single possible lean proof.

38:33And eventually you stumble upon the theorem. We have algorithms already that can solve any math problem as maybe what you were about to get at. Given infinite time, you can do a lot of things. So clearly it gets into diminishing returns as you think for longer. Yeah, very fair. What do you think is the biggest misunderstanding about 01? I think a big one was like when the name strawberry leaked, people assume that like it's because of this popular question online of like the models can't answer how many ours are in strawberry. And that's actually not the case is when when we saw that question actually we were really concerned that there was some internal leak about the model.

39:08And as far as we know, there was, and it was just like a complete coincidence that our project was named strawberry and there is also this like popular reasoning about strawberries. As far as I can tell the only reason it's called strawberries is because at some point at some time someone needed to come with a code name and someone in that room was eating a box of strawberries. And I think that's really the end of it. It's more relatable than Houston. I think I was pretty impressed with how well understood it was actually. Yeah. We were actually not sure how it was going to be received when we launched.

39:44There was a big debate internally about like, are people that are going to be disappointed that it's like, you know, not better at everything, are people going to be like impressed by, you know, the crazy math performance. And what we were really trying to communicate was that it's not really about about the model that we're releasing, it's more about where it's headed. And I think I was, yeah, I wasn't sure if that would be well understood, but it seems like it was. And so I think I was actually very, very happy to see that. Is there any criticism of O1 that you think is fair? It's absolutely not better at everything.

40:21It's a funky model to play with. I think people on the internet are finding new ways to prompt it to do better. So there's still a lot of weird edges to work with. I don't know, I'm really excited to see someone had alluded earlier to like the letting the ecosystem work with our platform to make more intelligent products, to make more intelligent things. I'm really interested to see how that goes with the one. I think we're in the very early days, it's kind of like, I don't know, at some point a year ago, people started to really figure out these LMPs, these language model programs with GPD4 or whatever, and it was enabling smarter software engineer tools and things like that.

41:07Maybe we'll see some similar kinds of developments with people building at Hub of 01. You can watch one of the things that we have not talked about is 01 Mini. And I've heard a lot of excitement about 01 Mini because people are generally excited about small models and if you can preserve the reasoning and extract some of the world knowledge, you know, for which deep neural nets are not exactly the most efficient mechanism. Like, that's a pretty decent thing topless. I'm curious, what's your love look, segment about O1 Mini and kind of the general direction that that represents? It's a super exciting model also for us as researchers.

41:45If a model is fast, it's universally useful. So yeah, we also like it. Yeah, they kind of serve different purposes and also, yeah, we are excited to have like a cheaper faster version and then kind of like a heavier slower one as well. Yeah, they are useful for different things. So yeah, definitely excited that we ended up with a good trade off there. I really like that framing because I think it highlights how much progress is, like how much you can move forward times how much you can iterate. And at least for our research, like Elga gets at Owen Mini lets us iterate faster, hopefully for the broader ecosystem of people playing with these models.

42:26Owen Mini will also allow them to iterate faster. And so it should be like a really useful and exciting artifact, at least for that reason. For founders who are building in the AI space, how should they think about, you know, when they should be using GPT -4 versus O1. Like, do they have to be doing something STEM -related, coding -related, math -related for TU -ZO -1? Or how should they think about it? I'd love if they could figure that out for us. Ha -ha -ha. One of the motivations that we had for releasing O1 preview is to see what people end up using it for and how they end up using it. Either way, there was actually some question about whether it's even worth releasing a one preview.

43:12But yeah, I think one of the reasons why we wanted to release it was so that we can get into people's hands early and see what use cases it's really useful for, what it's not useful for, what people like to use it for it, and how to improve it for the things that people find useful for. And the thing you think people most under -appreciated about 01 right now? It's like somewhat proof we're getting a little bit better naming things. We didn't call it like GPT 4 .5 thinking mode. Uh -huh. Yeah. A lot of that. It was strawberry. It was Q star. So I don't know thinking mode that kind of has a kind of has a ring to it.

43:58What are you guys most excited about for or O2, O3, whatever may come back. O3 .5, whatever. We're not at a point where we are out of ideas. So I'm excited to see how it plays out, just keep doing our research. But yeah, most excited about getting the feedback because as researchers, we are clearly biased towards the domains that we can understand, but we'll receive a lot of different use cases from the usage of the product. And we're gonna say maybe like, Oh, yeah, this is an interesting thing to push for. And yeah, like beyond our imagination, it might get better at different fields. I think it's really cool that we have a trend line, which we post in that blog post.

44:43And I think it'll be really interesting to see how that trend line extent. Wonderful. That's a good note to end on. Thank you guys so much for joining us today.

From the publisher

Combining LLMs with AlphaGo-style deep reinforcement learning has been a holy grail for many leading AI labs, and with o1 (aka Strawberry) we are seeing the most general merging of the two modes to date. o1 is admittedly better at math than essay writing, but it has already achieved SOTA on a number of math, coding and reasoning benchmarks.
Deep RL legend and now OpenAI researcher Noam Brown and teammates Ilge Akkaya and Hunter Lightman discuss the ah-ha moments on the way to the release of o1, how it uses chains of thought and backtracking to think through problems, the discovery of strong test-time compute scaling laws and what to expect as the model gets better. 
Hosted by: Sonya Huang and Pat Grady, Sequoia Capital 
Mentioned in this episode:

Learning to Reason with LLMs: Technical report accompanying the launch of OpenAI o1.

Generator verifier gap: Concept Noam explains in terms of what kinds of problems benefit from more inference-time compute.

Agent57: Outperforming the human Atari benchmark, 2020 paper where DeepMind demonstrated “the first deep reinforcement learning agent to obtain a score that is above the human baseline on all 57 Atari 2600 games.”

Move 37: Pivotal move in AlphaGo’s second game against Lee Sedol where it made a move so surprising that Sedol thought it must be a mistake, and only later discovered he had lost the game to a superhuman move.

IOI competition: OpenAI entered o1 into the International Olympiad in Informatics and received a Silver Medal.

System 1, System 2: The thesis if Danial Khaneman’s pivotal book of behavioral economics, Thinking, Fast and Slow, that positied two distinct modes of thought, with System 1 being fast and instinctive and System 2 being slow and rational.

AlphaZero: The predecessor to AlphaGo which learned a variety of games completely from scratch through self-play. Interestingly, self-play doesn’t seem to have a role in o1.

Solving Rubik’s Cube with a robot hand: Early OpenAI robotics paper that Ilge Akkaya worked on.

The Last Question: Science fiction story by Isaac Asimov with interesting parallels to scaling inference-time compute.

Strawberry: Why?

O1-mini: A smaller, more efficient version of 1 for applications that require reasoning without broad world knowledge.

00:00 - Introduction
01:33 - Conviction in o1
04:24 - How o1 works
05:04 - What is reasoning?
07:02 - Lessons from gameplay
09:14 - Generation vs verification
10:31 - What is surprising about o1 so far
11:37 - The trough of disillusionment
14:03 - Applying deep RL
14:45 - o1’s AlphaGo moment?
17:38 - A-ha moments
21:10 - Why is o1 good at STEM?
24:10 - Capabilities vs usefulness
25:29 - Defining AGI
26:13 - The importance of reasoning
28:39 - Chain of thought
30:41 - Implication of inference-time scaling laws
35:10 - Bottlenecks to scaling test-time compute
38:46 - Biggest misunderstanding about o1?
41:13 - o1-mini
42:15 - How should founders think about o1?

More from Training Data

All 110 episodes
OpenAI's Noam Brown, Ilge Akkaya and Hunter Lightman on o1 and Teaching LLMs to Reason BetterTraining Data · 45 min
Listen in VO