OpenAI cofounder Greg Brockman on the scaling hypothesis and refactoring as a killer AI use case

18 Jun 2025 · 32 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Cheeky Pint Podcast Episode Summary

Episode Title

Greg Brockman on the Scaling Hypothesis and Refactoring as a Killer AI Use Case Host: John Collison Guest: Greg Brockman, Co-founder of OpenAI and former CTO of Stripe

Episode Overview In this episode, John Collison interviews Greg Brockman, the co-founder of OpenAI and the first engineer at Stripe. The conversation dives into various aspects of AI development, including the scaling hypothesis, lessons from deep learning, the future of AI, and the challenges of energy bottlenecks.

Key Discussions

  1. Scaling Hypothesis
  2. Greg discusses the scaling hypothesis and reflects on whether OpenAI was the first to take it seriously.
  3. OpenAI’s approach was to prove its viability through their Dota 2 project, where scaling their compute power led to exponential performance improvements.
  1. Deep Learning Insights
  2. Dota 2 Project:
  3. The early struggles and milestones during the Dota project provided major lessons in AI development and management.
  4. Greg shares that the traditional management style of setting outcome-based milestones failed, revealing the importance of input control and adaptable strategies.
  1. Turing Test and AI Advancements
  2. Greg posits that while significant achievements have been made in AI, the original Turing Test has yet to be surpassed.
  3. He emphasizes the need for AI to deliver tangible economic value, particularly in areas like personalization, which could be seen as the next frontier for AI.
  1. AI in Coding
  2. The conversation touches on the role of AI in programming, highlighting refactoring as a notable use case.
  3. Greg believes that AI can take on more mechanical aspects of coding, allowing humans to focus on more complex tasks.
  1. Energy Bottlenecks and Future Scaling
  2. Acknowledges potential energy bottlenecks as a future challenge for AI scaling.
  3. Greg advocates for increased energy production to support the computational demands of AI technologies.
  1. Research-Driven Product Development
  2. Discusses OpenAI’s unique approach of pursuing technology without initially having a clear problem to solve, contrary to typical startup methodologies.
  1. Personalization in AI
  2. Greg highlights the shift towards personalized AI interactions, noting how past models lacked user context.
  3. Personalization is seen as crucial for the advancement of AI capabilities and user experience.
  1. Future of AGI and Predictions
  2. Greg provides insights on the timeline for achieving Artificial General Intelligence (AGI), suggesting that meaningful advancements could occur within the next 2 to 5 years.
  3. He reflects on the unpredictable nature of AI progress, affirming that the focus should remain on achieving step-function improvements annually.

Key Takeaways

  • Experimental Learning: The need for continuous adaptation and learning from failure is essential in tech and AI development.
  • Personalization as a Frontier: Many believe personalization in AI is a critical next step for enhancing user experience and interaction.
  • Energy as a Bottleneck: Energy limitations could hinder the scaling of AI technologies in the future, necessitating an increase in energy production.
  • Innovative Uses of AI: Refactoring and coding assistances represent practical applications of AI that can significantly enhance productivity in software development.

Conclusion This episode provides a comprehensive look at the current landscape of AI development through the lens of Greg Brockman’s experiences, discussing both the challenges and opportunities that lie ahead in the evolving field of artificial intelligence. The conversation emphasizes the role of deep learning, scaling, and the necessity for creativity in problem-solving as integral to future advancements.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00This is totally backwards from how you're supposed to do a startup, right? You're supposed to have a problem. And we had no idea what the problem was. Is there a world where the AI becomes the manager and that, you know, give you ideas and give you some time to do? This was probably the hardest project I've ever done because it felt totally doomed, right? It's like, I know like every instinct, every builder instinct of mine, actually, he filled him. Oh, it was it. Greg in 2010 dropped out of MIT to become our first engineer. When we went on to become Stripes CTO in 2015, after he left, he co -founded OpenAI.

0:29Open AI is really cooking at the moment and Greg is one of the most productive people I know. Cheers!

0:37Wow, you really have some... What's called photos? Some deep cuss. GDV Tril, yeah. So... Yeah, we figured we'd put people at ease, you know, make people feel at home. Very nice, thank you. Okay, well I just drive straight into all the questions I have which is lost. Alright. If you were not working in AI, how could one have known that something was about to start working? People were telling you that AI was the future in the 1970s and the 1980s and the 1990s, and then very quickly in the late 20s, I was never considered to have met. Well, I was someone who was not in the field, and so I remember very much what it was like.

1:102013, 2014, it felt like every day on Hacker News, there'd be a new deep learning for X article. And I remember being like, what is deep learning? And I knew like one person in the field, and I asked them to introduce me to more people in the field. And I just kept getting introduced to a bunch of my smartest friends from college. Now, if you actually look at the work that was being done, you're 2012 basically, you know, image recognition for the first time, you could solve with the neural net much better than anything else. And I just blew all these traditional computer vision to approach it out of the water.

1:42It's like this like learned system that is able to outperform 40 years worth of let's like write down all the rules and try to like, you know, sort of handcraft the algorithm for the task. And it's very easy to then be like, okay, well, this approach sure works for computer vision, but it's never gonna work for machine translation in, you know, 2014 suddenly you're getting great results in machine translation. And I think that this pattern was applied in subfield after subfield. One thing I've been wondering about is so many different things are finally working at the same time. And so we have LLNs, which are obviously amazing.

2:14Then we also separately have image models really working And we also have text -to -speech -to -text working way better than they were before. And so, what's the common factor behind everything starting to work at the same time? Well, it's deep learning. Right? I think deep learning is the core for a long time. Like, you know, why didn't deep learning work in the 1980s? Well, so, okay. So, you look at the number of orders of magnitude of compute that we've gone through from 1940 to today. I mean, it's just astounding. Using their all -explain by compute scale -ups, I think. I think that's the right algorithms.

2:44Of course, the type of algorithm changes in some of those results aren't even deep learning, but I think that fundamentally it is about compute, and you needed an algorithm that is scalable that can actually absorb that compute. Was open AI the first company to take the scaling hypothesis really seriously? I think that claiming the first is always difficult, but I think that it is clear that we are sort of succeeded much more wildly sooner than anyone else. And so I think that we have real conviction behind what we needed to do. Some people think that OpenAI set out to prove the scale hypothesis, whereas it was almost the other way around of the scale hypothesis is what we observed as the thing that was working for us.

3:19And you know, we really saw it for the first time actually during our Dota 2 project. We started out with 16 cores to train a little agent on Yacob and Shimon who were getting the project from an ML perspective on their desktop. And then they scaled the 32 cores. And it felt like every week I'd come back to the office and they scaled up by another 2X and we had 2X performance. And just so clear, you just need to keep going. Where does this thing peter out? and it just never did. Founders get too much credit because you have an initial product that's a pretty reasonable idea. And then you listen to the customers and you follow what's working on.

3:52So you were saying that was kind of open AI with the scaling hypothesis where you started trying to make Dota AI's work and you noticed that adding more compute worked really well. And you said, where else will just throw more compute at it, yield benefits? I think that's like to first order correct. And I think one thing that distinguishes OpenAI from, you know, sort of the typical startup is we did everything in reverse. Right. It's like you're supposed to have a problem to solve. No one cares about the technology. Form DNC up front. Exactly. Yes. And for us, we really chased the technology without any idea of how it would be applied.

4:29Yes. And a lot of pursuing the technology really is, you have to let reality hit you cold hard in the face. There's just no other way to achieve results. You can't sort of will it into existence. You can't convince people that this is the thing. You have to actually make the system work. We have to figure out what is the right frontier, what are the problems, what are the things during the edge of working and to really double down on those. What else do you take away from the data work? How else does this? You could say you could just start with how it ends. We could have skipped that period in the wilderness, but it sounds like it was somewhat formative for the OpenAI organization.

5:03Well, so I think Dota had many lessons, one of which actually was in management lesson for me. I remember when he started out the project, I tried to set a list of milestones, right? It's like, okay, this date we're going to be this player, this date we're going to be this player. And it did not work at all. I remember our first milestone came and you realized that you cannot control the outcome, right? You cannot set outcome based milestones. What you can do is you can control the inputs of, we're going to try these experiments by this state. We're going to implement this feature by this state.

5:37And that is what actually worked. And I remember it was one of those things that was just like a story that I could not have sort of written in any better if we'd intended to, where we had a, we bid our in -house like best player. And then we were playing as a semi -pro, and he was just trouncing us, trouncing us. And then suddenly we were starting to get pretty good. So we showed up at the international, this tournament blind. First day, we had three players that we played against. We went 30, 30, and then two won. We were like, oh no, we lost. What happened? It turned out that this pro that we were playing against, that he had used an item we never trained against.

6:14And we were like, oh no, we're totally going to be hosed. So what do we do? Well, we just need to change the training. And so people stayed up all night to get this done. They added this extra item in there, 4 a .m., they finally got the job running. that Wednesday we're supposed to play against the number two and the number one player in the world. And our semi -pro plays against it. And he's like, this bot is totally broken. And we're like, oh no, we like, clearing head of bugs, something terrible has happened. And you're like, look, it's taking all this damage it doesn't need to, I'm gonna go kill it.

6:42It goes into kill it, he loses. And he's like, that was weird. He'd realized that what it happened was it had learned a baiting strategy. And then we realized, well, we have a super bot, but it's so bad at the beginning, because it's trying to do the baiting. So what if we just stitched the two bots we have together? And then that bot was just undefeatable. And we played against this number one player in one. And to me, this is like the story of how deep learning works. It's like you can't control where you're gonna go. You can control everything that goes in. You can put these metrics and these measurements and you can have sort of the evaluations and that being able to gauge where you're at is almost as important as being able to make the forward progress.

7:18But if you get all those elements right, then you can do true magic. And you're also describing something that works really well for an organization where it was motivating to stay up all night. Like, you know, if the prize was impossibly far away, it wouldn't have been as motivating, but the fact that there was a near -term reward function and you were able to show concrete progress. I think so. Yeah, I mean, I think some of my favorite engineering stories have the same character. I mean, when I stay up all night to get our ISO 85, the 83 integration stuff. Something that's staying up all night for like critical projects that actually have important history in my hands out.

7:49Yeah, all startups. I'm glad to hear the tradition of the live -o -out on the wild opening. So Dota's fallen, Chess's fallen, Go has fallen. We've passed the Turing test. I think by anyone's measure with people commenting how the little fanfare wouldn't be dead, but we seem to have pretty good. What's a good new Turing test? Well, I'll tell you two things. One is that if you look at the strict version of the Turing test, I would actually claim we haven't done it yet. So no one's really gone that extra mile to say, can we actually have an AI that is fully indistinguishable for me human? And it's not clear if it's even a good task, right?

8:25But I think that the right question to your point is like, well, what is the milestone that we should be chasing in terms of capability? Like I remember talking to one of our board members in 2018, and he said, look, I get that you were outside about near -term age AI, but it just doesn't feel like it's on track. And I asked, well, what do you mean? He said, in a world with the near -term age AI, you would expect massive economic value to be delivered by AI already. And where is it? And in 2018, I think it was a very fair criticism. And clearly that's starting to change now. If it was like one thing that's may really change the AI Marcus is personalization, up to quite recently, when you asked Chacha Dea question, it was like walking into, you know, a shop off the street, they've never met you before, they know nothing about you, whatever.

9:12That's obviously not ideal for this, you know, close part of your digital life. I'm curious how you're thinking about personalization from a product point of view. It feels to me like the most meaningful change since the chat interface. Two and a half years ago. Two and a half years ago. I mean, I think it's absolutely critical. And I think it is very rightly considered to be kind of a next frontier. Like I am someone who always, when I just Google something, I go into incognito mode. Because I don't even want my computer to remember that history. And I always used to go for contemporary chats on chat GPT.

9:45But now my usage has totally reversed. I want Chatchee, BT to know to remember everything. I want to remember all of my interactions, because it's useful. Okay, so you guys figured out from a product point of view how to make the memory actually work better. And this is actually a product point of view, but also really research point of view. And I presume there's a flip -flop between the product and research where when you find something that's useful for a product point of view, then the product people say, and I'm just a product person, like, you researchers go actually make this good, and then that kind of kicks off more research.

10:10It's like, well, we actually, so to some extent, that's a failure mode in our mind, right? that I think that we really don't want to have that kind of silo. We really want to blur the lines and have people cross -collaborate. So it's very different mindsets from how you would traditionally build a product versus how you do research. Part of what had happened actually was that we had GPT -3. We knew we needed to build a product in order to be able to see and raise funding. We were like, well, what product do we build? We wrote down a list of like 100 different products. We could do a medical thing.

10:40Then you're like, well, now we have to sell the hospitals. You can have to hire doctors and you just realize you give up on the G and AGI, right? You're gonna like go for a specific thing. And so someone had the idea of saying, well, why don't we just make a API and let people figure it out? And again, this is totally backwards from how you're supposed to do a startup, right? You're supposed to have a problem and we had no idea what the problem was. Yeah, yeah. And so we're gonna back into the problem. And so this actually felt like this was probably the hardest project I've ever done because it felt totally dimmed.

11:10Right, it's like I know like every instinct every builder instinct of my actually feel didn't oh felt it wasn't just like Open ended or something no fault. You're still doing it. I mean, it's like at some point if you have there's no There's definitely no other path. There was no other path was the only shot we had I remember someone also saying like I can't imagine anyone paying for samples from this model and I was like might be right And I'm still trying to imagine myself. And it was just not clear where we above threshold, or below threshold. And we showed it to people and people were interested, but the very different from people being like, I will build my company on top of this.

11:48And I suppose the first use case to get any traction. So AI Dungeon. And what was that again? There you go. So AI Dungeon was a tech -based, tech -based adventure game. Oh, sure, yay. Yeah. OK, but that was real revenue. or that was zero revenue. Yes, and in fact, I believe they were our first paying user. And I could you confuse where you're like, oh, clearly the future of OpenAI is gaming. I know, I know, I know. I'll go to our roots. Exactly. And it's interesting too, because yet we had dreaming of all of these applications and medicine and all these things. And you start with the gaming application.

12:25But we could see signs of life on so many other things. I think in many ways, JPT3 was like the best world's best demo machine, right? when we bring these to the API, people were coming with all these cool things you could do, but making them reliable was so hard. It really wasn't until the next generation of GPD 4 until we started to figure out how to do post training well, that then you were actually able to build real businesses on top of these things. Bill Gates was saying recently that GPD 4 was the best demo he'd ever seen since Xerox Park. He said it to me. The night he saw it, yes. So that's high praise.

13:00I want to say to the medicine thing you mentioned this. So like you said, I think your family has personal stories you've talked about getting very valuable, diagnostic help. We ourselves actually, it's a bit more minor in our family, but we managed to fix a cat thanks to debugging it within the LLM. And I think this is an interesting example, right? Because so many people that I know have had some kind of experience like this. And maybe it's because you actually don't get that much time from a doctor. Are there other examples like this medicine application where you're seeing a lot of success that many people have similar stories, but we just hear less about it.

13:42Yeah, I think it's a great question. And by the way, I think medicine is an example of one where I kind of thought it was going to be one of the last domains that we successfully built to add value in. but it turns out that the bar is so low and you just need to exceed WebMD. And so I think that we have seen sort of other areas that are like a real common theme. Like one that's very interesting right now is like the life coach, like life advice kind of application. Or you just talked about it. You're actually really taking off? Yeah, it really is. And so I think that there's things like education is another area that just like clearly like is really having an impact.

14:15And there are studies coming out now that actually show that people are able to learn better through the use of these tools. Let's be expected, right? Like it is the bloom to sigma effect in a product. Yes. And that, for example, is like why, you know, Sal Khan, that's why he started Khan Academy, was to think about if you can give personalized tutoring to everyone. Yes. And we showed him GPT -4. He's like, this is the thing. Yeah. Like, we're going to become a GPT -4 app. And so I think that there are these really amazing applications that are affecting everyone's like daily lives. You know, obviously programming is another one.

14:46I think that people are seeing all across the board in professional context. We're heading to a world where just like, you know, if you want to do productive work, you don't have access to computer, like you're going to be hampered, and so similarly, not having access to AI is heading in the same direction. Okay, so speaking of not having access to AI, I will posit that these days, it feels like AI product development is mostly OS limited. Is that how you feel? Like are we stuck at the moment? I do feel a little the stuckage, but it's not to worry. It is overcomable. But yeah, I think it is true.

15:25Look, like two years ago, we released plugins in ChadGPT. Do you remember them or those? Yeah. Yeah. And that was like trying to make it so anyone could write apps that then ChadGPT could access. And the models were just not that good. Right. Then we limited it to like three plugins at a time, because there's only so many functions and stuff. And it just wasn't that reliable. And now we're in a world where MCP basically, is really taking off and is a way to hook up your AI to different tools and very much like kind of try and take that same type of idea and really make it work. Now the world that we're in is very similar where there's certain interfaces we don't have or being able to access your phone and all those APIs.

16:07And there's a question of is the model above threshold to actually use them or not? And my observation has been that basically I think that there is maybe a lag of like six months of different interfaces that are hard to access. But once we have a model that's good enough, we will find a way, people will find a way. And so I think that we're in a world where I have every expectation that we will get the future that has been promised, but it's just gonna take some work. I feel like there are many moments where I'm using my phone and I want a single button where it's just like, you know, check to you.

16:38What do you think of this? Like I need your comment, I need your fact check, I need your explanation, something like that. And you take a screenshot and you're like, go into chat GBT and you click upload photo and it feels like very 1993 versus the button on my phone that just says, hey, chat GBT, what do you think about this? And obviously you guys are not empowered to go build best. That's what I mean by it feels somehow like we're a little operating system. Well, I definitely get it, but I'll say, I think that there are two dimensions. So this is kind of how I've been thinking about things since we released the API back in 2020.

17:08There is capability and convenience. What you're referring to is the convenience, right? but it's pretty inconvenient to do the screenshot and paste it. But the thing is, if the capability is good enough, you are willing to accept any sort of inconvenience. It's like if this, by taking the screenshot, showing you the chat, you could give you amazing insight. You can tell you how to build stripe in some way, and it takes you a month to do it. You have to crawl to the top of the mountain. You'll do it. The convenience will not stop you. And so the point that I'm trying to make is that if the capability is high enough, people will start doing a specific thought.

17:45They'll discover the use cases and the convenience will just catch them. And then in the convenience, there's so much pressure. There's pressure on the phone manufacturer. There's pressure on us. There's pressure on everyone in order to bring down the security. So it seems to be patient. It'll be great in three years time. Yes. Really in it. New is the AI. A criticism people like to love you of AI is, yeah, it's great and handy and all, but it hasn't come up with a single novel advance in mathematics. Have you? or science. Well, you could have if you had become a mathematician, but humanity has for keeping the score board.

18:20Why don't you make it that criticism? Just wait. Okay, so you think like it would take one of the millennium prizes, you think we possibly will see that. I think for sure. I mean, there's no question. Like two years, five years, ten years? I think that is the question. It's just timing, right? That is my question. I would put it two to five years as the right number. And I think ultimately, this comes back to the question of benchmarks, right? Is that actually being able to solve a linear problem is pretty high bar. Yeah. Pretty high bar. And once you can do that, there's so many other things that will definitely be possible.

18:56And I think that we're starting to see the leading edges of this. And to me, if we look at our definition of AGI, we recently started talking about the framework of thinking about levels of AGI, starting from chat bots to reasoners to agents to innovators to organizations, five levels. And we're basically somewhere in level three right now. Yeah. And level four, this innovators, like that's going to be different. You know, I recently posted some pictures of our visit to Appling, Texas, where we're building these big data centers together with our partner Oracle. And imagine taking that whole data center and just thinking hard about one problem.

19:34Imagine just thinking about how to solve a Millennium problem or how to cure a specific kind of cancer. Maybe it needs access to some apparatus. Maybe it needs access to robotic wet labs. Maybe it needs access to different tools in the world. But that level of computational power couples with the ability to experiment and learn from your ideas. That is going to be something the world has never seen. Okay, so yet again, we just we haven't put a a respectable amount of compute from these problems compared to what we will be doing. We're still least tiny little computers. So that actually gets to, in terms of these scaling laws, do they eventually run out because we just run out of compute or like do you eventually get to the point where we're inventing new kinds of nuclear energy and so that is what unlocks the next level.

20:18A lot of energy that comes online now is for data centers, which is not true when you guys started training GPT -2. And so, isn't that just the upcoming bottleneck? Right. As it should be, right? It really should be that it's energy manufactured into intelligence and that's your only bottleneck. But I'm saying that'll be like quite a plateau compared to the exponential growth we've seen over the past few years. Well, I think, unless the key change in terms of promising and plans for building and everything like that, this is, I think, the core, right? is that if you look at every trend in this field, there is these exponentials, these S curves that sum up to exponentials.

20:57Sure, but these exponentials were mostly existing in like tech, Silicon Valley space, where it's pretty easy to have exponential growth. It's pretty hard in permitting and real estate and damning rivers and building nuclear power plants. It's harder to have exponential growth. Let's see how fusion pans out. Yeah, okay, but even fusion, most industry observers would say is still kind of five years away. And so, what it's the next five years of power growth. So, look, I think that it is very possible that we end up bottlenecked on energy. And that's actually one reason that we've been spending a lot of time really trying to advocate for the fact that we just need far more power.

21:33And I think that my observation of the market is that ultimately the capitalist markets do provide. Yeah, I think there's this absolute tsunami of demand that is coming our way. But I feel some confidence that, again, And when there's enough pressure, when there's enough clarity of this is the bottleneck, and it's not just for any company, right? It's really for national competitiveness. And you look at other countries that are just building huge amounts of power far more than we are. I think that actually for America to remain competitive, there's just no choice but to build out for real power.

22:06We do. Speaking of bottlenecks, everyone was talking about the data wall in 2023. I think this is an interesting thing where knowing it's talking about the data wall anymore, And yes, it doesn't feel like AI progress has slowed down. Just test time confused, is this, like, people were wrong about the data wall and where it presented a bottleneck. Like, these are actually still a data wall, but it's two years away. It's basically all of these things, right? It's really, it truly is, right? It's like, you keep changing the paradigm, right? That is the real core of the Kurzweil view of the world, right?

Read the full transcript

22:37Is that, fine, this one way of doing things, taps out, and if you just look at that one way of doing things, you feel hopeless, right? you feel like this is it, but somehow you will find a new S -curve. I think that's what's happened, for example, synthetic data, for example, reinforcement learning. If you think about the RL paradigm, fundamentally that's a data production mechanism, right? Just that the AI happens to be training on its own data, and then you learn it very rapidly, and then you learn on that. Each of these has taken us much further. I think there are lots of algorithmic ideas, lots of techniques, lots of ways of even using the existing data better.

23:11And so I think that fundamentally the S curves continue. And if you zoom out, it all looks smooth and uninterrupted. So it's kind of like chip monetization for each generation. People are like, okay, well that's the smallest. You could possibly make a chip. That's it, we're done. But monetization. And somehow we've regretted that. Yes. And now one difference with chips is, I think at the end of the day, there is some concomit. Like, we're right, but we've never been that close to that. Right, right. Yes, yes. And where does AI coding go? And in particular, vibe coding is all the rage right now.

23:44It's kind of the term of 2035. It's sort of working. It's very impressive. No one is really fully letting AI software engineers run end -to -end in production. So I'm just curious, what are your one to two -year predictions on what happens with AI coding? Well, my general observation is that once something kind of works in this field, the next gen is going to be great. And so I think that's where we are right now for AI coding. And so I think what we're going to see is AI is taking more and more of the drudgery, more of this like pain, more of the kind of parts that are not very fun for humans.

24:20Now one thing that's very interesting is that I think that so far the vibe coding has actually taken a lot of code that is actually quite fun and left behind the review and the deployment and these things that are not fun at all. And so I'm hopeful that we're actually going to be able to make a lot of progress on these other areas as well. But fundamentally, we should really end up with a full AI co -worker. And I think it really will be anything you want to create. You can be the manager, right? And you can really have this team of software engineering agents. Now, the thing that I think will be very interesting to see is, is there a world where the AI becomes the manager?

24:51And it gives you ideas and gives you some tasks to do. And that's something that, again, it's just like totally backwards in terms of how we think about it. But are there ways in which you can actually have outcomes for companies and actually have people who've jobs become much more meaningful because they have an AI who really deeply understands them in the same way that your AI doctor really deeply understands all of your needs. But is in part of the common thread that we're talking about here, often places where AI tools underperform, it's because they're trying to do something generally, you know, like a voice recognition is not that good because it's trying to recognize all voices as opposed to trying to recognize, you know, my voice in particular.

25:29and similarly with AI coding, they work well in places where you need no context at all, and we're single -shorting an app based on publicly available libraries, and then places where you have to understand a million -line codebase. They haven't fully figured out how to do a good job of that. Like, is that a fair parallel to draw between all these challenges? Well, I think there's two things in there. One is that I think this is already changing, So if you look at something like Codex, it's actually great at operating in a big code base. I ask it for where functionality is implemented, and it's better than I am at finding it.

26:07Which is kind of a wild fact. It's super cool to see it, like wrapping around and just going and exploring. Actually, this is one thing that we really shot for with Codex, was to build a tool for software engineers who are not necessarily vibe coding. It's not about building a new app for scratch, which is a cool demo, but that's not actually how much software gets written. And actually, I think maybe the killer enterprise feature is refactors, right? It's like rewriting your copal app or changing your Facebook to hip hop, you know, to do a static PHP. Exactly. Right? If you think about it, like the amount of like deep sophisticated thought that is required to accomplish a refactor is actually not that high.

26:45There's a lot of mechanical work that's just the sheer volume of it that's hard. And like that's an AI -shaped problem for sure. So I think we're going to see a lot more productivity on all sorts of AI tasks as a result. Now, we are in a world where you said a second thing, which is maybe you need to narrow down more. I think that the way these models work is you actually do want one model that knows more and more things, and you want it to have some personalization to you. But the fact that it kind of has this one base model that kind of knows everything is actually a very useful starting point.

27:18So I do think that you're going to see a world where we'll have more and more capable base models and figuring out how do you really connect it to all of your organization's code and context and history. How does OpenAI decides plus products to do? I'm just curious how you think about when to develop specific products or when you think, oh, you can truly do that with Chatchy PT and that's good enough. Yeah, this is a really tough question, right? It's something we really struggle with. I think that over time, I guess actually when we first launched at GBT, we're glad for this, well, we're an enterprise business and we're a consumer business.

27:53And that seems terrifying, terrifying as it started up. And I remember talking to one of my board members who said, it just feels like what you have is an unfocused strategy at first, because you're just doing all these different things. But if you think about it, maybe an analogy is to a company like Disney, where you make one core asset, like Little Mermaid, right? and then you productize it in all these different ways, right? Think about Little Mermaid, the ride, the lunch box, the t -shirt. I think that we have some element of that, we have the core model, and then we have a question of, well, what are the applications that this can add a lot of value to?

28:28Quickly, right? With like a small amount of additional work. So I think the question of what areas to go into are, how far does it take us off the general path on the return on, and how important is this domain area, especially for achieving the bigger goal? How much synergy is there across it with respect to other things we work on? So coding is one where there's very clear synergy, very clear ROI, right? Because if we can speed ourselves up, that's something that accelerates everything. How is being from North Dakota? How's it shaped you? Well, North Dakota was an amazing place to grow up. Well, then actually, come on.

29:03It actually helped be there. I've been there. And it was great. Okay, it was incredibly safe. Our doors didn't even have working blocks. Like it was that kind of place. I had a lot of freedom academically. So sixth grade, my dad taught me some algebra. Seventh grade was the first time they split you into advanced math and so I was going to be taking pre -algebra. So my mom took me to go see the teacher. The teacher and we asked, can you skip? And the teacher looked at us very condescendingly and said, every parent believes that their child is special. I can guarantee your son will be plenty challenged in my class.

29:36And so after a month, the me sitting in the back just playing games on my calculator and you know she calling me randomly to try to trip me up and just looking at the board and be like, two X. She said, okay fair enough, your son is nothing to learn in this class. And so they moved me into eighth grade algebra. But then eighth grade rolled around and I had no more math left in my middle school. And so you went to college right? Well so I did in high school, in high school start going to University of North Dakota and took a bunch of classes there. But also, I was connected to a lot of people who were the top math kids in the country through things like math camp and the math competitions.

30:11So you were saying the social scene was not too distracting in North Dakota? Not too distracting, but it was definitely fun. Last question. Do you remember we were going to camp by sea in 2017 and asked you how far AGI was away, and you said two or three years. and stuff like that. You did, yeah. And I don't see the recording. Well, I just try to never think like, was I right? Were you right? How should we grade that? Because we didn't get a GI, but we didn't not get a GI either. And so I'm curious if you have any reflections from your own AGI prediction journey? I think that we are, I will say, I think that AI is surprising.

30:58I think that that is like the single most consistent theme is that the thing we were picturing, we got something different, but we got something better, more magical, something that is more helpful. I'm actually quite happy with that. Now, predicting where you go, it's again, it's really hard to manage the outputs here. One goal of OpenAI that we have successfully achieved is every year to have at least one result that just feels like a step function better than anything before. I think as long as you see this, you just want to have one really awesome AI feeling thing you'd year. That kind of thing.

31:30I like that. Yeah. That's a good way to tie it back, which is the way we grade that prediction is that you've stopped setting metrics based on outputs. Yeah, exactly. Yeah, yeah. Yeah. But it does feel we're getting really close to something pretty magical. I agree. Thank you. Thank you.

From the publisher

Greg Brockman—OpenAI cofounder and Stripe's first engineer—joins John Collison to talk about research-driven product development, an early moment he thought OpenAI was doomed, S curves in AI advancement, and energy bottlenecks.

Full episode transcript:

https://cheekypint.transistor.fm/1/transcript


Timestamps

(00:00) Intro

(02:51) Was OpenAI the first company to take the scaling hypothesis seriously? 

(04:53) Lessons from Dota about deep learning 

(08:08) What is a good new Turing test?

(08:57) Personalization in AI 

(09:57) Research-driven product development

(10:26) An early moment OpenAI felt doomed 

(15:01) OS limits on AI product development

(17:59) When will AI make novel advancements in math or science?

(20:03) Energy bottlenecks

(22:30) S curves in AI advancement 

(24:00) AI coding 

(26:25) Refactoring as a killer AI use case

(27:26) How OpenAI decides what products to built

(28:53) Growing up in North Dakota

(30:17) How far away is AGI?

More from Cheeky Pint

All 35 episodes
OpenAI cofounder Greg Brockman on the scaling hypothesis and refactoring as a killer AI use caseCheeky Pint · 32 min
Listen in VO