OpenAI’s CPO on what’s coming next: Hardware, GPT-5, Jony Ive, agents, more

10 Jun 2025 · 54 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: OpenAI’s CPO on What’s Coming Next

Episode Overview In this episode of Azeem Azhar's Exponential View, Azeem Azhar interviews Kevin Weil, the Chief Product Officer at OpenAI. The discussion dives into the future of AI, the recent developments at OpenAI, including hardware advancements, the upcoming GPT-5, and the evolving role of AI in society and business.

Key Topics Discussed

  1. Recent OpenAI Launches
  2. Introduction of new connectors for integrating ChatGPT with tools like Google Docs, Gmail, Calendar, and more.
  3. Aim to enhance model utility by accessing personal and enterprise data.
  1. Role of AI in Everyday Life
  2. AI's evolution from a question-answering tool to a proactive assistant that performs tasks.
  3. Discussion on how young people perceive and utilize AI differently, as they have integrated it into their lives more seamlessly.
  1. Addressing Fears about AI
  2. Insights on public apprehension regarding AI technology and suggestions for overcoming those fears by encouraging hands-on experience with AI tools.
  1. Model Evolution and Development
  2. Kevin discusses the unpredictability of AI model progress and the importance of iterative deployment.
  3. Introduction of the concept of model "evals", a way to evaluate AI performance across various tasks.
  1. Prompt Engineering Importance
  2. The significance of crafting effective prompts for AI and how this skill may evolve as models become more advanced.
  1. Defining AI Agents
  2. Kevin defines AI agents as systems that can perform independent tasks, moving beyond simple question-answering roles.
  1. Coding as a Target Use Case
  2. The focus on coding as a primary application for AI, highlighting how it can enhance developer productivity.
  1. OpenAI's Vision for Hardware
  2. Discussion on potential hardware developments to complement AI software, aiming to create more integrated, user-friendly experiences.
  1. Generational Perspectives on AI
  2. Young users approach AI with an "always-on" mentality, differing from older users who are still adjusting.
  1. The Future of AGI
  2. Kevin reflects on the journey toward Artificial General Intelligence and the incremental nature of this progress.

Key Takeaways

  • Integration of AI into Daily Tasks: The aim is to evolve ChatGPT into a tool that not only answers questions but also performs tasks autonomously, enhancing productivity.
  • User Control and Trust: As AI systems take on more responsibilities, maintaining user control is essential for building trust.
  • Rapid Technological Change: The pace of AI development is unprecedented, with capabilities emerging rapidly, making it crucial for users to engage with AI tools early on.
  • Collaborative Development Model: OpenAI's approach combines research and product development, fostering a close collaboration that accelerates innovation.
  • Expanding Opportunities in AI: Coding is a significant focus area due to its broad applicability and potential to democratize technology for a wider audience.

Closing Thoughts Kevin emphasizes the transformative potential of AI across all sectors and the importance of ensuring that AI tools are accessible and beneficial to everyone. As the technology progresses, it is essential to rethink workflows and harness the capabilities that AI brings to both individual lives and organizations.

Links

  • [Kevin Weil's LinkedIn](https://www.linkedin.com/in/kevinweil/)
  • [Azeem Azhar's Substack](https://www.exponentialview.co/)
  • [Azeem Azhar's Website](https://www.azeemazhar.com/)

This episode provides valuable insights into the direction of AI technology and its implications for various sectors, reflecting OpenAI's commitment to pushing the boundaries of what's possible with AI.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00We're in this transition from ChatGPT being a thing that answers questions to a product that actually goes and does tasks for you in the real world. you should be in control of any actions that it takes. The best way to do that is to kind of co-evolve together. Sam has sort of talked a little bit about generational differences. So what is it like if you're in your early 20s, you're using ChatGPT? If you're younger, it's just a core part of the way you operate your life. There's sort of an always-on nature. You realize that you have this super assistant in your pocket that can not just answer any question, but it can teach you anything that you want to learn.

0:33How far out are model capabilities being delivered? Are there things that are being worked on today that might not make their way into models until mid-2026? It's unpredictable. Sometimes things take longer than expected. Other times you see these capabilities that you didn't expect at all that are kind of emergent, and all of a sudden something just works. It's a completely different way of building products. Computers can, like, do things that they couldn't do two months ago, and we're constantly in that state.

1:06Today, I'm so excited to welcome Kevin Weil. He is the Chief Product Officer of OpenAI, coming off the back of a storied career in companies like Twitter and Facebook. So much to talk about. So first of all, Kevin, thank you for making the time. You've been busy today already. It's morning in San Francisco, and you've already launched some things. Yes, we have. And thank you so much for having me. It's my first live Substack, so I'm excited. but hopefully many more to come. You are really keeping us on your toes. And so just an hour before we went live, there was another spate of product launches from OpenAI.

1:43What do they tell us about your vision for ChatGPT and where the product is going? The biggest thing that we launched this morning, we launched like six different things this morning, but I think the most important one kind of for the long-term future of AI is we launched a series of connectors that can connect to your either personal data or if you're an enterprise, your enterprise data. So this is connectors into Google Docs, to Gmail and Calendar, to SharePoint, to OneDrive, Dropbox, Box, Linear, all of these different tools that you use every day to get things done. With the rise of our reasoning models, connecting them into the services and the data that you use helps the models be way more useful.

2:28So it's not just that you can now ask, at work, for example, if you connected in for your Google Docs or SharePoint, you connected into the docs that you use every day, suddenly ChatGPT can get all of this context on the enterprise and on the state of projects and on the latest about any particular thing that's going on. You wouldn't ever have an employee at your company that you didn't give access to docs. It's like where the conversation and the strategy and other things are happening. So now you have ChatGPT that has the ability to do it. And today they're read-only, so they're able to access the information but not create on the other end.

3:08But you can imagine going forward, another big part of this is that ChatGPT should be able to take actions, should be able to help you write a document, or create a presentation, or write tasks to your task management system, and ultimately combine all of these together to begin really working like an employee. Right. So that's the hint, right? That direction of moving ChatGPT from something we interact with, with a challenge and response into something that is more evolved and feels like it's really doing work for us as an employee. We're going to spend some time digging into that. It's such a big vision.

3:44Let me ask about you. You're the chief product officer of what I think is the most important company in the world today. So how does that feel like to you day to day and week to week? I mean, it's a privilege. It's the most exciting place I've ever worked. And I've been very fortunate in my career to work at a bunch of awesome places with really great co-workers. But I think, you know, the opportunity in front of us, the rate at which AI is changing all of our lives, means we have an ability to make a really big impact. And we take that super seriously. So I get to work with amazing colleagues.

4:19I get to kind of, you know, have a front row seat to the way that our models are evolving and that the way this is all impacting our lives. And hopefully we get to build some products that make a difference in your life and in the lives of everybody listening here. Well, nearly a billion people have made at least voted with their mouse finger to use ChatGPT just since November 2022. Technology always changes us. The television did. We got TV dinners and we got water cooler discussions. The internet obviously changed us in different ways. Cars as well. We got big box retail. We got the suburbs.

5:00We got the picket fences, desperate housewives. How might AI products reshape our everyday life? Well, I think one of the other interesting things you get when you have these waves of technology is you start by doing the thing that you were doing with the previous wave, but sort of with the new medium. You know, the first TV advertisements were people standing on a stage reading their radio advertisements. And then people slowly figure out that you can actually, you know, you do what we have as commercials today, which are much more interactive and sort of dynamic. And so, you know, we're still probably in that mode of people, when they look at the impact that AI can have in their lives or in their work, they're kind of like, okay, I have these processes.

5:44How do I, you know, sprinkle AI on top of this to make it better, faster, etc. And like, that's fine. That's all good. As in previous technology transitions, the power comes from completely reimagining the work that you're doing from first principles using the new technology. So mobile wasn't just about a computer in your pocket. It was, you know, you have access to GPS and you have these totally notifications and totally new ways to interact with technology. I think over the next year, we're all going to be in the process of reinventing the way that we do things with AI. And the fun part is the technology is moving so fast.

6:28It's not, I mean, faster than any technology I've ever worked with before. I've ever seen in my career. So even as we're reinventing, the technology is gaining new capabilities. And it's just, it's an exciting time to be alive. I mean, it is that that rate of change is quite something I did. I was looking for something today, which I know would not have been possible four months ago. So I was using O3 to help me find a little portable mobile 5G router for when I travel and I just need to get a good quality signal. And of course, O3 went into all the technical specifications of the radio frequency chipset and said, well, you know, this earlier router has got a slightly better Qualcomm chipset they no longer use, and you'll get one extra signal bar in rural environments in the US, but not in Europe.

7:16And I'm thinking that sitting there going, this is kind of insane. But I'm really curious about those types of behaviors. What are the surprising behaviors that you've seen emerge around ChatGPT from your users. Sam has sort of talked a little bit about generational differences, but give us a flavor of things that you've learned about behavior that are just really hard for product teams to research or figure out a priori. Well, I do think, I mean, Sam has talked about the way that, like I said, the people today, I think a lot of people are sort of sprinkling AI onto their existing workflows. And it's young people that are, you know, they're sort of, this is native to them in a way that it's not native to, you know, those of us who grew up without AI.

8:00And it's, I mean, my kids just, they're like, of course, you can talk to a super powerful AI that like, you know, can customize itself and answer any question you have. And kids that are graduating college these days as like engineers are like, they don't know any other way to write code than using, you know, AI editors like Cursor and windsurf and it's just completely natural to them and so it gives them superpowers in a way and it's one of the reasons that we look a lot at how our youngest users and our young users you know late teens college users things like that are using the product because it teaches us a bunch can you characterize those those differences though so what is it like if you're in your early 20s you're using chat gpt how does it feel different to somebody who might be in their mid 30s or defaulties.

8:53There's just a, there's sort of an always on nature. You just, you realize that you have this super assistant in your pocket that can not just answer any question, but they can, it can teach you anything that you want to learn. And so when you, when you go about life that way, the rest of us are like trying to remember, you know, trying to think about the processes that we go through and how we can reimagine them. If you're younger, you might not have had those processes. And so you built them from the ground up with AI. And so it's just a core part of the way you operate your life. And in some ways, they're ahead and the rest of us are catching up.

9:30A lot of people are scared of this technology. A lot of people are nervous. When I've just come back from Brussels and spoken to people in a different sector, and you do get that sense of fear, what could change in the product to address that? One thing that I think is really important is people ask, you know, what should I do about AI? How should I think about it? And my answer is always just use it. And, you know, of course, I think everybody should try ChatGPT, but whether it's us or any of the others, just start using it because it's the number one way to realize it's not this like super scary thing that you read about.

10:09It just helps you get more done. And you're like, oh, this is great. I now have this like incredible new like thing that can be part of my life and help me get things done and help me like automate a bunch of the like boring work that I have. So the number one thing is just like, just start using it. Also, given the rate that the technology is improving, if you don't start using it now, it's going to be even harder to sort of catch on to it. And if you believe that AI is going to be a big part of our lives, that's not a train you want to miss. So, you know, that's the number one thing. Just start using it.

10:39But then we think a lot, especially as you get towards agents and other things, we think a lot about making sure that the user is in control. So you don't want the, you know, as you're getting to use ChatGPT or any other AI, you don't want it to go off and do a bunch of things for you without you, you know, feeling in control. And so, you know, it's one thing if it's like answering a question for you, reading some docs and summarizing them, things like that. But as we're in this transition from ChatGPT being a thing that answers questions to a product that actually goes and does tasks for you in the real world, and you should be in control of any actions that it takes.

11:20And over time, as the models get better and you begin to trust them more, then sure, you can give it, you know, sort of more leash and trust it to take more actions autonomously. but you should control that every step of the way and I think that's one of the most important ways that we're looking to build trust. I mean, there's a lot that will change as you move to more and more of these agent-based workflows and the models that be able to use tools that I guess we have to co-evolve. Definitely, we should look into those questions. I'm just curious about your own experience. A lot of us have these oh shit moments when we use LLNs, right?

11:54There's something sublime that happens or there's two years of your work that gets returned in five seconds. What was your most recent oh shit moment using your own products? Well, so I'll give you, this one's a little bit more, I don't know, pragmatic, but it was a really meaningful thing for me. We have, we were talking about kids as people were sort of coming on and one of our sons had a minor surgery and it was one of those things where all odds were that it wasn't going to be a big deal, but there's a small chance that it was a really bad thing, you know? And so they do the surgery, and they take the thing, and they go to biopsy it, and you're waiting to hear back.

12:42And as a parent, you're nervous, even though you know logically that the odds that it's anything bad are really small. And at some point, we got a letter in the mail with a bunch of, you know, doctorese on it that looked pretty intimidating, frankly. You know, there were a bunch of words I didn't understand and characterizing what this thing was. And it wasn't, it didn't say like, you should worry about this or you should not worry about this. It just like characterized it and then, you know, finished. And I couldn't get ahold of the doctor. She was in surgery or something, you know. And so there I was like, oh my God, what does this mean?

13:19And so I took a picture of it. I put it in chat GPT. And I said, like, should I be worried? And I said, like, can you explain this to me? Like I'm five. And chat did it and was like, no, this is totally fine. Everything like nothing to worry about. And I actually ended up not being able to get ahold of the doctor for 72 hours because she was just super busy. And that 72 hours would have been a terrible 72 hours for me as a parent if I was just sitting there brooding. and ChatGPT was able to answer. And that's like, that's, you know, us with like access to great healthcare. You think about the impact of this all over the world where people don't have the same access.

13:56It's really powerful. And I think that's sort of an underappreciated part of ChatGPT. I mean, that is a wonderful story. And I'm glad that your son is healthy and comes out of that well. But it also speaks to the power of this particular product, right? It's a really complex, complex product. And there's no way I can talk to you without asking you about how that complex product finds its way in that funny dropdown in the top left of ChatGPT. There must be something that you guys are all so brilliant. There must be some internal joke about, well, how should we order it? What should we call the next one?

14:34What's really going on there? And shouldn't we all just be using O3 and 4.0 if we're in a hurry? It's a totally fair question. And you can also, you can and you should make fun of us for our naming as well. It kind of comes back. So we have this philosophy of iterative deployment, which is these models are, I think AI is going to change all of us. It's going to change the world. It's going to change society. And we believe that the best way to do that is to kind of co-evolve together, is to get these models out there and put them in people's hands, help them understand. And also, you know, they help us discover the capabilities of the models and the weaknesses and other things.

15:16And so we sort of learn together and can iterate really quickly and improve. So that's part of it. The other part of it is we're building a lot of new capabilities as we go. And if we took the time to only, you know, we just had like one model and we just had to build everything into one model, we'd end up moving a lot more slowly. And, you know, things would be simpler for sure. But we'd end up moving a lot more slowly because sometimes it's easier to build a new model with a certain set of capabilities that's great at certain things, not as good at other things. And so together you have sort of a collection of models that can do lots of things.

15:53Each individual model has its strengths and weaknesses. And so basically, we've optimized for going faster and getting more capabilities in people's hands at the expense of a bit of confusion. And then over time, as we sort of gain more control over some of the new functionality, we understand it better, we build it back into the core model. So you have models like GPT-4 that can do a lot of things well. And this is what we're trying to do with our forthcoming GPT-5 is take a lot of the things that we've learned and build more of the capabilities into a single model so that it's easier for people to reason about.

16:33It's like, what model do I use? Just use GPT-5. And in a perfect world, it knows how hard the question you ask it is. And so it knows whether it should give you an answer like this or whether it should think for a while. And, you know, that's what we're shooting for. And, you know, in a way, that's lifting the cognitive load that is currently on the user because I will sit there and I'll think, do I have time? Is that a complicated question? Does it need to go to the reasoning model 03? Does it need a longish prompt for 03? Because it might get the wrong end of the stick if I give it a shortish prompt.

17:06And in a way, what you're talking about is crunching that all and building it into the model itself. Yeah. And, and look, what I don't want to do is over promise and say, Oh, in the future, like, we'll only ever have one model, and it's going to all be simple. Because we, we're also, you know, say we launched GPT five, we're then going to have a bunch of new capabilities beyond that, that we're trying to build an experiment with, and we're going to want to get those out to people and deploy iteratively, and so on. So I expect you always have this sort of phenomenon of there are new models and then you've got sort of your workhorse models and then you've got some of the new ones that have certain like frontier capabilities that we're experimenting with and learning with together.

17:48And then over time, as those mature, they all kind of get merged back into single models. So it seems that a lot of the velocity is actually about getting through the loop, right? The learning loop. And you just run all these horses at the same time so you can gather enough data about what models work against what capabilities. In a way, that helps you develop and deliver on GPT-5. And I hope you'll share the date with us. Down to the minute. Down to the minute, yes. So we can time ourselves. But does that also mean that, you know, how far out are model capabilities being delivered? Are there things that are being worked on today that might not make their way into models until mid-2026?

18:32Yeah, it's a good question. It's one of the most interesting things about working here is you kind of have a sense of what's coming. And, you know, I'm not in the research team, so I'm getting this, you know, as I collaborate and work with the researchers. I would say like, you know, on the product side, we have a decent sense of what's coming in the next, say, three months, maybe a hazy sense over the next six months. And beyond that, it's harder to say. You know, you have a certain set of capabilities, you know, where like, you know a bit, but you see things coming through the haze a little bit.

19:08And sometimes, sometimes capabilities are, you know, it's research, right? So it's not like you just, you have the formula and you just turn the crank. We're uncovering new things and it's unpredictable. So sometimes things take longer than expected. Other times you see these capabilities that you didn't expect at all that are kind of emergent. And all of a sudden something just works. Can you give an example of one that just worked that you weren't expecting? Well, I mean, deep research is an interesting example where for a while there were a handful of researchers thinking about the, it was like, okay, we could probably make the model able to do this like iterative kind of research where, you know, with deep research, you give the model a arbitrarily complex query to go research something that would probably take you a week.

19:56And it will go off and it'll go off and do like 100 searches, but not all at once. It'll do, it'll do three or four or five. And then it'll reason about the results that it gets back, try and understand how they pertain to what you asked and what gaps are still there. And then we'll go off and do some more searches and maybe think again. Maybe it'll write some code for a little while as it thinks and then go do some more, you know. And so it's this sort of iterative, I mean, it's what you would do if you were, if someone made you write a very complex research report, you wouldn't, you go do some research.

20:27Sorry, Kevin, I don't do that anymore. I just go to deep research and I actually can't remember how to do it on my own. But yes, I understand, right? You have to, you sort of inch your way across and you figure out exploration strategies and go down dead ends and come back. Yeah, and so to your question, that was something that some folks were like, okay, this is coming together, but it's not clear exactly when it will come together. And so there was a small team of researchers that just believed in this and were working to make this real. And for a while, it wasn't good enough, it wasn't good enough.

21:01And then there was some advances and all of a sudden you're like, okay, this is getting good enough. And somewhere in that timeframe, we put also a product and engineering team working on it with them. And then you have the thing that I think really is the magical part of OpenAI. When you get a research team and a product and engineering team just in the same room, all bringing their unique skills to bear and you understand the problem you're trying to solve. And so you're bringing back use cases, creating evals and benchmarks for how you measure whether you're successful against those use cases.

21:36The research teams are taking that and using that to improve the model itself. And you get this tight loop of the model improving towards a particular product. And I think our best products are the ones that we build that way. And deep research is a good example. That's a novel way of thinking about product development. I mean, if I think about the history of product development, quite often it was before the 90s, before the consumer internet, was run by engineers and they'd say, we've got a new chip that can do this. And you then try to figure out some software that can do useful things on that chip.

22:10I think the big breakthrough of the consumer internet was to put product managers at the heart of product development. And we talked about lean and iteration and being very data-driven and user-centric. And now you're getting to this new model, which I would characterize as not at all a return to the pre-internet product engineering led. but something that is quite novel because the researchers discover something that's a little bit a new capability and then you have to have a very rapid discussion about how can that capability be productized and then this word you used, eval, which I guess it means how can it be measured to see whether it's actually doing its tricks.

22:53So is it really a new discipline that is evolving at this point? I think it's a completely different way of building products. It's certainly different than anything I've ever done in my career. And, you know, within research, there's kind of a there's a spectrum, right? There are there are parts of our research team that are that are just like deep research. It's almost academic in nature, because they're trying to just like they're looking for new breakthroughs, they're trying to, to find things that nobody has ever, you know, figure out things that nobody's ever figured out before. And those kinds of things, you don't want to be, you know, you don't want to be product driven at all, because you just you want to give a lot of room for exploration and fundamental breakthroughs.

23:31And then there's kind of the other end of research. It's more in the post-training side where you really are trying to teach the models to do specific things very well. And those teams tend to be much more like, you know, partnered with product and engineering teams with a common goal. And then it's kind of a spectrum in between. I think the right way for us to be is not, we certainly don't want to be entirely product-led. That's not the magic of this place. It's not maybe entirely research-led either, because it's good to know feedback about what problems you can solve for people and how we can make the biggest impact in the world.

24:11It's really sort of a combination of both, with research really at the core, though. And I've loved it. It's the most fun in the world. and move super fast. Computers can do things that they couldn't do two months ago and we're constantly in that state. But when you're working with research, then it's more than just looking at a scaling law and saying, oh, we've just, Sarah Fryer has just signed off another 100 ,000 GPUs. Therefore, it will be able to do this in six months when the trading run is done. There's more to it than that. But one of the things I find fascinating is how do you map those capabilities against products and you talked about evals, which I guess are evaluations.

24:52What is the structure of an eval? And is it what replaces what was in an old product requirements document that we might have had 15 years ago? Yeah, sort of. I think in some ways for understanding where the model is good and where it's not. If you think of the model as sort of an intelligence of some sort, intelligence is so multifaceted. People are smart in a million different ways. And one smart person is better than another in certain areas and worse than another. So one way to think of evals is a way to measure capabilities and intelligence of models on different dimensions. So you can have evals around how good it is at solving USAMO Math Olympiad-style problems.

25:40And another around how good it is at chemistry. And another one about how good it is at creative writing. But are you using the public benchmarks for that, the sort of RKGI and AIME and GPQA as your way of measuring? Some. And then also when we're building specific products, I think one of the most effective ways to build products is to take the skill that you want the model to have in order to meet the product need. Turn that into an eval so you can actually understand whether, you know, how good you are at it and also how you're getting better over time. But one of the fascinating things is like the evals that we all used, you know, a year ago to measure models They're all very kind of cut and dried like you're testing against math And with math, there's a right answer You can talk about creative writing evals though and with creative writing.

Read the full transcript

26:30There's no there's no answer So how do you grade that right? That's what that's one problem The other is like as you start to take on more complex tasks. You're not just answering questions you're actually trying to automate some multi-step workflow, there may be ambiguity in the right way to do that. If I'm an AI booking a flight for you, there's not a single way to grade which correct flight. You also get into these really interesting, challenging, subjective ways of how do we actually grade this particular task? And part of having an eval, if you want to at least automate it, is you need to also have a grader for it so that you can very quickly understand how you're doing on that eval.

27:11So it is interesting. It's one of the skills that I think is going to be more and more important for PMs over time is the ability to actually create evals for the products that you're building. Yes, I mean, that's one. And the other one is the prompt at the front end. Because what we are starting to see with these leaks, and I don't know how real they are or not, but there are a number of X accounts that say I've just had the leaked system prompt of and then insert your favorite foundation model or coding tool. And the system prompts, which are the sort of structured instructions that go out with every query, are really, really quite complex.

27:49I mean, they're product in of themselves now, right? They run to thousands of words. They're highly structured. There's clearly strategies being applied as they get put together. So how important is that skill and capability when you're shipping products to people like me? I mean, actually, more than people realize, I think. I would love to make it, over time, less of a thing. And I think over time it is. Like, if you go back a year or two, everybody was talking about prompt engineering and it was going to be the skill that everybody had to master in order to do anything with AI. You don't hear it talked about quite as much like that.

28:27And I think that's a good thing. You know, ideally, it matters less and less that for any particular user, if they have a question, they want an AI to do something for them, you shouldn't need to get into like Arcana around, did I use the exact right word? And did I give my exam? You know, it should just work. I think that's part of increasing intelligence is the model can understand what you're looking to do and do a good job of it without you having to work super hard at it. That said, prompts still do matter and the models are very controllable with prompts. And so, you know, we still find we'll launch something and we'll find that it's not behaving in certain ways the way we want it to.

29:08And we can adjust it with a prompt a lot of times. You don't need to go back and like retrain the model. So it's both that I want to make it less necessary over time and that it is still a powerful vector. Well, these are two of the vectors of the direction of the product. But the third one is the idea of agents and what agents will bring to us. I think probably the first agent product that OpenAI launched that I used was deep research. And the word agent is being thrown about quite a lot. I mean, I'll use the word agent. And what I mean is I've strung a bunch of prompts together through your API.

29:43And there's a bit of logic to move a document through a series of steps to the other end. What does an agent mean to you? We think of an agent as something that can do independent work. So it's not just a quick, you know, you ask a question, you get an answer, but it's actually off doing tasks for you in the real world. So another, I think deep research is a great one where it's off doing, you know, hundreds of searches and putting together a complex report for you that might have taken you a week. I think another is Codex, which is our software engineering agent that we just launched. You know, what you can do is if you have a code base that you're operating against, you're building a new feature in a code base or debugging something, you can just give this agent the prompt like, hey, I need you to fix this thing.

30:28I want you to do this to the background of my web page. I want you to build this new feature. And it will go and look through your entire code base, understand all of the context. If you're fixing a bug, it'll go try and figure out where that bug exists. And then it will write new code for you and create a pull request like a diff. you know, here's the set of changes we need to make to the code. And then you can go review the code and the agent did the whole thing. And so you know, I used to be an engineer, I still write code a little bit in my spare time, but I haven't written a single, you know, line of code for open AI.

31:03But with Codex, I was this was like, you know, a few days before it launched, I was, you know, it was like 11 at night or something, I was doing a bunch of work that I had to do before I could go to bed. And I was like, you know what, I bet I could fix a bug right now. And so went and found a bug that looked relatively simple, and just, you know, pasted the context into codex said, can you go off and fix this bug? By the way, it was in a language that I had actually never worked with in my life. So it would have taken me even more time if I had to do it myself. And 10 minutes later, I had a pull request, it looked reasonable, I submitted it, an actual legitimate engineer looked at it and said, yeah, this looks right.

31:43And now there's a few lines of code shipping today that came from me using Codex. It speaks to the power of this thing when you can just have this software agent off actually solving real world tasks for you. And in the meantime, I was writing email and following up on Slack and doing all the things that I do in my day job. So it was just purely additive, which I think is really cool. Yes, because the codex process takes a little bit of time. It has a lot of material it has to read and understand and then make the changes. And I'm curious about, this is a question that everyone who's built a product that does code automation or developer augmentation is asked is, so today what portion of the OpenAI code base is in the first instance produced by codex rather than by a human engineer?

32:35Yeah, it's pretty meaningful and it's increasing quickly. Right, okay. Somewhere in the meaningful, I'll go up and ask 03 what meaningful means in percentage terms. And I'll get a good distribution there. The cool thing is you can fire off 10 of these tasks at once, right? So we try and actually give you the value of all this parallelism where it's not just you can do one thing. But if you have a codex agent working for you, why not have 10 codex agents working for you on 10 different tasks? And by the way, just to connect it to the previous topic on evals, there's a really important kind of subtlety to them too, where they have to be tailored to the product that you're trying to build and the problem that you're trying to solve.

33:19where coding isn't one thing. Coding is a small vertical of the entire world. But even within coding, you can be good at lots of different kinds of coding. And with Codex, that was a great example of going and saying, okay, what kinds of coding really matter to us? What kinds of tasks? And all the tasks that a developer does, what kinds of tasks do we really want to be good at? And we created evals for those. And then we made sure to monitor as we trained the model, is it getting better and better and better at these? and you go and accumulate tasks and examples for the model to learn from, but you do it against a specific set of evals that correspond to a specific set of problems you want to solve.

33:58It's very capability-driven in that respect, right? And then it speaks to how do you actually do enough testing both to make sure that you are getting to the right level of score that you want, but also making sure that it doesn't go off the rails, right? And I think that as these agents get more and more complex, given more complex tasks. That's something to bear in mind. I mean, in one of my workflows, which was a very, very simple one where I wanted an agent to go through and grab some data from a series of web pages and populate an Excel spreadsheet. And I was using some third-party agent framework.

34:38And it was so diligent, Kevin, that it said, I must check my work, which it did about 400 times and left me with a$75 bill and had got it correct the first time, right? Got it right and got stuck in this strange loop. So that's one of the things I think that I hear when I go around people saying, well, how are we going to be able to control these things? Not from a humanity out of control measure, but from a enterprise reliability. How can I make sure that this isn't like the sorcerer's apprentice and the thing runs out of control when I simply asked it to book one flight to Italy and it's booked me 200.

35:19And how do you test for all of that? Yeah, I think part of this is about making sure that like we talked about earlier that the user is in control here. So you should be able to at some point be like, hey, you know what, you've checked enough, like, you're good. And the other interesting thing in all of this is the technology is evolving so quickly, like much more quickly than I think we're used to with technology. We're used to things taking like decades to deploy and to really achieve scale. One of the phenomenons you see with AI technology is there'll be some benchmark, some eval that AI just can't crack.

36:01And people are like, oh, AI just can't do that. And then one day, somebody ships a model that gets like 5 % on that eval. Still mostly can't do the job, but just like begins to get it. And then what you inevitably find is like two months later, there's a model that's at 30 on that eval. And then four months later, there's a model that's at 60. And then, you know, within six months, it's completely saturated and like models are great at that new skill and will forever be. And so you go very quickly from like, proof of existence to like, oh yeah, of course AI models can do that. That like rate of development is still, I think, something that we're not totally used to.

36:47It's that first one or two percent, right, that becomes hard, that proves it can be done. It's the Kitty Hawk flyer. And then within 30 years, we're moving large numbers of passengers across the Atlantic. But in this case, it's within 30 days. So I want to ask about coding and coding agents. So if you look at the growth of Gen AI applications and SimilarWeb had some data come out a few weeks ago, the baseline was that the generalized chatbots are growing at 25 % per quarter. That's the chat GPTs and so on. Virtually every other product category, image generation, video generation, sound generation, is growing slower than that or declining in size.

37:29And I view that as the black hole that is the capability of your core models. The one category where growth was faster, 75 % a quarter according to SimilarWeb, was in coding. I was curious, have you selected coding from a kind of commercial perspective because you can really see the demand and developers are always willing to experiment? Or do you select coding because being a testable, structured, verifiable set of outputs, it's a slightly easier challenge than the sort of fuzzy amorphous tasks that occupy the rest of the world? Yeah, it's a really good question. And actually, coding is this vertical that kind of hits all of these things.

38:11You know, for one, it's really important to us because if we can speed up coding, if we can make, you know, every engineer more effective, we also make ourselves more effective. And so we can build even faster and we can bring AGI to the world faster. So it's interesting to us from that perspective it's a clear kind of milestone or step on the way to AGI itself, because it's a very sort of general purpose reasoning. It's also a relatively gradable task. Like you can tell like, like in math or other things, if you get the answer right, it's also something that our engineers are familiar with. So it's a problem space that they understand and have good intuition for.

38:50It's also a huge market, as you were saying. It's also a market full of early adopters. technologists leaning into this, it's also relatively sort of open and unregulated. It's not like trying to go into health or something where there are all kinds of other things you have to do. And so it's this aggregation of all of these interesting things that make coding a really interesting market. And I haven't seen that data, but I totally believe it. And would you say that within coding, are you already seeing signs of serving people like you who are not technically engineers anymore. In other words, we're seeing that expansion of the market through these tools.

39:31Oh, yeah. And I think there's going to be so much value in democratizing coding out to the world. There's like, what, 30 million developers or something worldwide, depending on how you define it, which is great. That's a lot of people. But imagine if a billion people can write code. I was talking to somebody the other day who was just telling me they were, during COVID, They were working for their local county, trying to get, you know, vaccinations and stuff out to people. And they were trying to put together a website to track so people could like sign up and just do basic stuff. And the whole world was busy and they just, they couldn't do it.

40:09They couldn't create a website. They didn't have the skills to do it. And as a result, they were, you know, managing things less efficiently, doing a bunch of manual work at a time when everybody was slammed. And he was just saying, like, can you imagine if I had these tools, we would have been able to create a website overnight. It would have just worked. And, you know, they would have been able to do their work more effectively. And you have that times a million as you look around the world. And so, I mean, that is actually the other thing about coding that I think is super interesting. That's maybe like the, you know, ninth reason why coding is a good, it's such a general purpose technology.

40:46If you can create code, then you can create all kinds of things. And so there's something really powerful to the idea that a billion people might be able to write code. But it also speaks, though, I think, to how this may fundamentally change the software industry in the way that the internet changes the software industry, not just because of packaging and distribution, but the way we interacted with social technologies, right? My Microsoft Word, when it was on floppy disk, never allowed me to sort of exchange notes with somebody else on like Google Docs. One of the big questions that's out there is, as a platform company that is building the most performant models out there, how much space do you leave for startups?

41:34I remember when Microsoft introduced disk compression in Windows 95 or 97 or something, and there were a whole load of third-party companies that offered disk compression that essentially went out of business there and then. And this is something that happens on X every time you release a new foundation model. It's like every time Kevin tweets, another 50 startups die. Where is their space, right? Where is their space in the software world, in the AI software world, which is safe from the increasing capabilities of the foundation models that you and other companies are building? So Stephen Sanofsky told me an interesting story about this one time.

42:12Stephen used to run Windows at Microsoft and Office and, you know, everything. and he was telling me this story about like the transition from windows 93 to windows or whatever it was called back then to windows 95 windows 3.1 maybe where it was just like the beginning of the internet and and so most people weren't using the internet and if you were going to actually get on the internet with windows 3.1 you had to like go to some you know university of Oregon professor's website and download a TCP IP stack, compile it yourself, and like, you know, install some device drivers, and then you could actually go on the internet.

42:51And then in Windows 95, of course, the internet was happening. And they were like, okay, we need to ship this stuff with Windows. And so they did. And it was like you were saying, there was this, you know, there were a bunch of people who were like, hey, now that, you know, you've just like, put that university of whatever a professor, he did all this work and now you just shipped it. Come on. And Stephen's point was that you would never want to live in a world where today you still had to go to like some professor's website and download a TCP IP stack and compile it yourself to get it going. You just want to use the internet.

43:23And basically the expectations of a platform, the consumer expectations of a platform are an increasing function of time. And if the platform can provide more of the technology that, you know, if you see something where in order to build the actual thing that people want to build, you've got 10 different companies having to go build the exact same piece of like foundational infrastructure, you should probably just provide that. And then those 10 companies can go do like more interesting stuff. And I've always remembered that story. It's really stuck with me because I think the fact that people are just going to expect more and more out of these platforms is very real.

44:02But the upside is all for third parties in this, for developers in this world, because if the platform provides more of the building blocks, then they can spend less time re-implementing the wheel on these building blocks and more time doing the thing that they actually uniquely add value in. And AI is going to change absolutely everything in our life. Any industry, any vertical, any geography that you can imagine, AI is going to touch. And so there's so much opportunity for developers to reinvent and reimagine. I think anything that we can do on the platform side to help accelerate that by making more of the building blocks easy, we should be doing.

44:43So let's imagine my son, 20 years old, and he wants to build a product on top of OpenAI. Where is a good place for him to go and build it? I mean, almost everywhere. There's so much opportunity. Sam said this one time and it stuck with me. He said, if you're building a company and you're building at the frontier of the model capabilities, if you're building something that really just barely works and you can't wait for our next model because you know it's going to make your product sing, then you're probably building in the right place because you're introducing something new to the world. You're making something possible that wasn't possible before.

45:28and that's where you want to be. If instead you're building like some sort of scaffolding around that covers up the weaknesses of a current model and you're actually afraid of our next model because it might not have those same weaknesses, that's a bad place to be building because on average, models are going to improve really fast and what's a weakness from one model will not be a weakness of the next. So like the, I think the thing to be building is like what we talked about at the beginning, re-imagining use cases from first principle building them from the ground up with AI. And if you're in a place where you're excited about the next model that comes out, because that's what's going to make your product sing, that's a great place to be.

46:07That's a fantastic heuristic, which is actually if you're a founder out there, think about something that you want to build that the models are not yet capable of, but will be capable in a little bit of time and you can build on top of that capability. We can't talk about products without talking about your new product buddy, Johnny Ive. So tell us a little bit about what the mood in the office was when that lovely black and white photo was released. Oh, people are incredibly excited. I mean, how could you not be? I use products that Johnny designed my entire day. He's been a part of building some of the most sort of cherished products and hardware that we use every day.

46:45How could you not want to work with him? And getting to know him through the process and other things, he's also such a lovely human being. For somebody who's accomplished so much, he is so humble, thoughtful, kind, soft-spoken. And then, you know, he'll say something sometime and you'll be like, oh my God, like that is a completely different way of looking at the world. And that opens my eyes to something that I just had never thought about. And so is this combination of like genius and also wonderful human being, how could you not be excited to work with him? And of course, he's a Brit. How is he going to work, his group and your group going to work?

47:27How are you guys going to interface? Well, I mean, he's coming in focusing on these sort of consumer hardware products. And then also over time, I think will play a very significant role in design as a whole at OpenAI. And again, I'm very excited about that. How could you not be? Johnny Ive coming in to take charge of a lot of your design, you know? you know he'll tackle the the drop down hopefully at uh at some point it's like we've got the eye the eye touch there and and so so i think that it's interesting to talk a little bit about hardware and how that interacts just in the last couple of minutes with the overarching vision so the overarching vision in a way you start at the beginning you talked about ai systems that will act a little bit like employees um i guess for people in their domestic lives that's more like like helpers.

48:18We don't generally have many. We think of it as a super assistant. Like a super assistant. And so what is the relationship between that and the need to have a hardware device alongside? I already have a hardware device. It's pretty good. I'm talking to it right now. It's more the opportunity. As we've said a few times, AI is going to touch every part of our lives and every part of our days in every part of the world. And that means, I think that there's an opportunity to reinvent and reimagine a lot of the services and the products that we use every day. You know, in some cases, there are a lot of products that I use every day that are great.

49:00They probably need to fundamentally change and they should with AI. And if they're not, they're not taking advantage of all these amazing new capabilities that we have, especially where those capabilities are going to be in 12 months, 24 months, 36 months. So I think there's an opportunity to reinvent and reimagine here. And that's true both on the software side and on the hardware side. So, you know, we have some thoughts about how that might occur. Obviously, Johnny has thought deeply about this, and we're excited to see what we can build together. And I'm sure there will be lots of others building in the space.

49:32And that's one of the reasons that we have, like, we put so much care and attention to our APIs and our developer platform. Because, The world is not just open AI. There's going to be incredible startups and incumbents and everybody else building really cool things using AI. And we'd like to power it in any case. Some of these will be first-party products that we build, and some will be other products that others build that leverage our models. And both of those things are really important to us. I mean, I hear that. I hear what you're saying as well, because I've already started to realize the limitations of the phone as the form factor for working with the models.

50:11You can't really put in a longish prompt to O3. I'm really reliant on talking to it. And if I'm in a noisy place, that won't work. The idea of having an ambient intelligence around me, I always have an AI model listening into my meetings and I'm talking to them regularly to do my work. So you start to see the limitations of something that's got the power draw of the phone and the size of the phone and does other things. So that will be a really exciting opportunity. And please sign me up for the alpha test. well before GA. We've got a couple of minutes. I just want to throw out some questions.

50:44How far behind are the top Chinese AI firms in core foundation model capability? Not as far behind as they used to be. And I think as US AI labs, we need to be very cognizant of that. I think it's really important that the leading models, the models that we all use, are models that are built off of democratic principles, not authoritarian ones. And we take that really seriously. Is there an AI app out there, whether it's in China or coming somewhere else, that's not built by OpenAI that you quite like and you like to use and play around with? I mean, I think a lot of the video apps are super fun.

51:19I also think, I also find Waymo magical. That's my go-to example of the way AI is touching our lives. You know, again, like, self-driving was like two years off for 10 years. And now suddenly it's here and it works and it's going to change a lot. It's absolutely magical. You are a keen runner, and I'm curious about whether you have a Garmin or a Suunto and what you would want from your exercise tracker that AI could bring that it isn't today. Ooh, that's a good question. Actually, I have an Apple Watch that I mostly use, and then if I'm doing like a 100-mile race or something, this doesn't quite have the battery, so I'll use a Garmin.

52:00What do I want? I think actually one of the things that I would love is better coaching. just like a little bit. And I think the AI is totally capable of doing it. I think Strava has some things that they're working on around this. But I would love to see better coaching and like AI analyze workouts and things like that. So that I got the, I think it's possible to get the kind of analysis that you would get from like a professional coach today from an AI for most users. And I feel like that's the kind of thing that, you know, five years from now, we're going to be like, oh yeah, I can't even imagine when that didn't exist.

52:34But it's only sort of peeking through right now. Just a little bit. But I do hear you. I think the possibility of having that personalized coaching would be quite sublime. So last question, when are you going to ship AGI? We're working on it. Every day we get a little closer. Every day. I mean, when will we know? Will we know? I think, look, I think it's one of those things that we talked about intelligence being multifaceted. There are already a bunch of places today where AI is way better than a human. And there are places where AI is like laughably worse than a human. But every, you know, every so, every like month or whatever, when there's new models, the baseline creeps up.

53:16And more and more things, AI is becoming superhuman at more and more things. And at some point, it's going to be superhuman at, you know, substantially everything. And we're going to call it, but it doesn't happen all at once. I think it's not like we go to bed one night and there's no AGI and we wake up in the morning and there's AGI. It's an incremental process of AI, you know, getting better and better at more and more things. Well, with that thought, Kevin, you keep climbing that hill. Thank you so much this morning for giving us your time. It's great to have you. Thank you for having me.

From the publisher

This week, I'm speaking with Kevin Weil, Chief Product Officer at OpenAI, who is steering product development at what might be the world's most important company right now.

We talk about:

(00:00) Episode trailer

(01:37) OpenAI's latest launches

(03:43) What it's like being CPO of OpenAI

(04:34) How AI will reshape our lives

(07:23) How young people use AI differently

(09:29) Addressing fears about AI

(11:47) Kevin's "Oh sh!t" moment

(14:11) Why have so many models within ChatGPT?

(18:19) The unpredictability of AI product progress

(24:47) Understanding model “evals”

(27:21) How important is prompt engineering?

(29:18) Defining “AI agent”

(37:00) Why OpenAI views coding as a prime target use-case

(41:24) The "next model test” for any AI startup

(46:06) Jony Ive's role at OpenAI

(47:50) OpenAI's hardware vision

(50:41) Quickfire questions

(52:43) When will we get AGI?

Kevin's links:

LinkedIn: https://www.linkedin.com/in/kevinweil/

Twitter/X: @kevinweil

Azeem's links:

Substack: https://www.exponentialview.co/

Website: https://www.azeemazhar.com/

LinkedIn: https://www.linkedin.com/in/azhar

Twitter/X: https://x.com/azeem

Our new show:

This was originally recorded for "Friday with Azeem Azhar", a new show that takes place every Friday at 9am PT and 12pm ET. You can tune in through Exponential View on Substack.

Produced by supermix.io and EPIIPLUS1 Ltd.


Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

More from Azeem Azhar's Exponential View

All 44 episodes
OpenAI’s CPO on what’s coming next: Hardware, GPT-5, Jony Ive, agents, moreAzeem Azhar's Exponential View · 54 min
Listen in VO