Inside OpenAI Enterprise: Forward Deployed Engineering, GPT-5, and More | BG2 Guest Interview

11 Sep 2025 · 1 h 9 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Summary: Inside OpenAI Enterprise: Forward Deployed Engineering, GPT-5, and More

Podcast Title

BG2Pod with Brad Gerstner and Bill Gurley

Description

A bi-weekly discussion on technology, markets, investing, and capitalism featuring Brad Gerstner and Bill Gurley with guest host Apoorv Agrawal.

Episode Title

Inside OpenAI Enterprise: Forward Deployed Engineering, GPT-5, and More

Description

This episode explores how OpenAI is transforming enterprise sectors through various applications of AI, featuring insights from Sherwin Wu and Olivier Godement of OpenAI.

---

Key Highlights

  1. OpenAI’s Enterprise Mission: Beyond ChatGPT
  2. OpenAI's mission revolves around building AGI (Artificial General Intelligence) and distributing its benefits globally.
  3. The enterprise platform includes APIs and partnerships with various industries, emphasizing B2B applications.
  1. Case Studies of OpenAI's AI Implementations
  2. T-Mobile:
  3. Partnership focused on automating customer support through voice and text interactions.
  4. OpenAI models are embedded in T-Mobile’s systems to improve service efficiency.
  • Amgen:
  • AI deployment aimed at accelerating drug development processes.
  • The goal is to streamline R&D tasks and administrative processes in healthcare.
  • Los Alamos National Laboratory:
  • Involves a secure deployment of a specialized AI model on a supercomputer for national security research.
  • Highlights the bespoke nature of AI deployment in sensitive environments.
  1. Challenges in AI Deployments
  2. Discussion on why 95% of AI deployments fail, emphasizing the importance of:
  3. Top-down buy-in and establishing a dedicated team (referred to as a "tiger team").
  4. Clear definitions of success metrics (e.g., effective evaluation metrics).
  5. The necessity of infrastructure and scaffolding for AI models to function properly.
  1. The Next Generation of AI: GPT-5
  2. The release of GPT-5 and its performance benchmarks.
  3. Focus on model behavior, instruction-following capabilities, and reduction of hallucinations.
  4. Discussion on reinforcement fine-tuning, enabling customized AI models tailored to specific use cases.
  1. Multimodality in AI
  2. Progress in integrating text, voice, and video functionalities into AI models.
  3. The Real-time API development aims to provide seamless, low-latency interactions for customer support.

---

Insights and Discussions

Physical vs. Digital Autonomy

  • Comparison of advancements in self-driving cars (physical autonomy) versus AI agents (digital autonomy).
  • Noted that physical systems have more established infrastructure, making them more advanced compared to digital environments.

Key Takeaways from GPT-5 Feedback

  • Positive reception regarding instruction-following capabilities, though some challenges remain in ensuring responses are not overly concise.
  • Ongoing adjustments in performance to balance reasoning time and response quality.

Long/Short Game Insights

  • Long Bets:
  • Healthcare: Predicted to benefit significantly from AI advancements.
  • eSports: Considered undervalued and expected to grow.
  • Short Bets:
  • General AI tooling and frameworks viewed skeptically due to rapid market evolution.

---

Conclusion The episode emphasizes OpenAI's transformative impact across industries through tailored AI solutions. The discussions highlight the challenges and successes in deploying AI at scale, particularly in enterprise environments, and the innovations brought forth by GPT-5. The ongoing evolution of AI technologies promises significant advancements in various fields, particularly in healthcare and customer service.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00We literally had to bring the weights of the model of physically into their supercomputer. When San Francisco, you could take a car from one part of SF to the other fully autonomously. As opposed to the digital world, I can't book a ticket online right now. Physical autonomy is ahead of digital autonomy in 2025. Yeah, agents are like really in day one. Here, like Chagy PT only came out in 2022. And the slope, I think, is incredibly steep. I actually do think self -driving cars have a good amount of scaffolding in the world. You have roads, roads exist, they're pretty standardized. Stoplights. AI agents are just kind of dropped in the middle of nowhere.

0:33We'll start with long short game. I'm short on the entire category of like tooling eVAL's products. Healthcare is probably the industry that will benefit the most from AI. I think I'm a JPL. You're definitely a JPL. The first one was the realization in 2023 that I would never need to code manually in a cable ever again.

1:01Hey folks, I'm a Poole Vagrival and today at the OpenAI Office we had a wide -ranging conversation about OpenAI's work in Enterprise. I have with me the head of engineering and head of product of the OpenAI platform, Sherwin Wu and Olivia Goedment. OpenAI is well known as a creator of Shia GPT, which is a product that billions across the world have come to love and enjoy. But today we dive into the other side of the business which is OpenAI's work in Enterprise. We go deep into their work with specific customers and how OpenAI is transforming large and important industries like healthcare, telecommunications, and national security research.

1:33We also talk about Sherwin and Olivia's outlook on what's next in AI, what's next in technology, and their picks both on the long and short side. This is a lot of fun to do. I hope you really enjoy it. Well, two world -class builders, two people who make look building easy. Sherwin, my palancer 2013 classmate, tennis buddy, with two stops at Korra and OpenDorth through the IPO before joining OpenAI. Before chat GPT, you've now been here for three years and lead engineering for all OpenAI platform. Olivier, former entrepreneur, winner of the Golden Lama at Stripe, where you were for just under a decade, and now lead all of the product at OpenAI platform.

2:15That's right. Thanks for doing it. Thank you. Thanks for having us. You know, as a shareholder, as a thought partner, kick an ideas back and forth, I always learn a lot from you guys. And so it's a treat. It's a real treat to be do this for everybody. You know, I'll open with people of no open AI as the firm that build chat GPT, the product that they have in their pocket that comes with them every day to work, to personal lives. But the focus for today is opening up for enterprise. You guys lead opening up platform. Tell us about it. What's underneath the opening up platform for B2B for enterprise?

2:49Yeah, so this is actually a really interesting question too, because when I joined OpenAI around three years ago to work on the API, it was actually the only product that we had. I think a lot of people actually forget this, where the original product from OpenAI actually was not ChatGPT, it was a B2B product, it was the API we were catering towards developers. And so I've actually seen the launch of ChatGPT and everything downstream from that. But at its core, I actually think the reason why we have a platform and why we started with an API is it kind of comes back to the OpenAI mission. So our mission obviously is to build AGI, which is pretty hard in and of itself, but also to distribute the benefits of it to everyone in the world, to all of humanity.

3:31And it's pretty clear right now to see Chaggit BT doing that because my mom, maybe even your parents are using Chaggit BT. But we actually view our platform, and especially our API, and how we work with our customers, our enterprise customers, as our way of getting the benefits of AGI of AI to as many people as possible to everyone in every corner of the world. ChagyPetit obviously is really, really, really big now. It's, I think, like the fifth largest website in the world. But we actually, by working through developers, using our API, we're actually able to reach even more people in every corner of the world in every different use case that you might have.

4:05And especially with some of our enterprise customers, we're able to reach even use cases within businesses and users of those businesses as well. And so we actually view the platform as kind of our way of fully expressing our mission of getting the benefits of AGI to everyone. And so, concretely, though, what the platform actually includes today, the biggest product that we have is obviously our developer platform, which is our API. Many developers, the majority of the startup ecosystem builds on top of this, as well as a lot of digital natives, Fortune 500 enterprises at this point. We also have a product that we sell to governments, as well in the public sector.

4:44That's all part of this as well. And also an emerging product line for us in the platform is our enterprise product. So what we actually might sell directly to enterprises beyond just a core API offering. Fascinating. And maybe to double down, I think B2B is actually quite core to the open -air mission. What we mean by distributing edge -eyed benefits is, I want to live in a world where there are 10 next -small medicines going out every year. I want to live in a world about education, public service, civil service, increasingly optimised to everyone. There are large categories of use cases that only go through B2B, unless you enable the enterprises.

5:29We talked about the parent tier. I think that's probably the same fees that parent tier is like, those are the businesses who are actually making stuff happen in the real world. If you do enable them, if you do accelerate them, that's how essentially you benefit if it's newly distributed, AGI. Yeah. Well, maybe we can double click into that Olivia. The reach for chat is obviously wide billions of users. But for enterprise, it's maybe tell us about it. Maybe we go deep into a customer example or two. And what is an organization that we have helped transform, maybe, and at what layers? So if I were to step back, like we started our B2B efforts with the API like a few years ago, Initially, the customers were startups, developers, indie hackers, extremely technically sophisticated people like who are building cool new stuff essentially and taking massive market technical risk.

6:22So we still have a bunch of customers in our category and we love them and we keep building with them. On top of that, over the past couple of years, we've been working one more with traditional enterprises and also like digital natives. Essentially, I think basically everyone woke up like with TGPT and like those models are working. There is a turn on value and they could see a centrily many use cases in their enterprise. Couple of examples we tell like the most. One which is very both fresh and you know it's quite cool. We'd be working a lot with T -Mobile. So T -Mobile leading like US teleco operator, T -Mobile has like you know a massive customer support load.

7:01Like you know people asking like you know hey I was charged like that amount of money was going on or you know my I say, I'll phone like, not working anymore. A massive, like, you know, share of that load is like, you know, voice calls, like people want to talk to someone. And so for them, like, you know, to be able to essentially automate like more and more and, you know, to help like people like self -service in a way, like, you know, debug their subscription was pretty big. And so we've been working with Tim Mobile pretty much for the past year at that point to basically automate like, not only like text support, but also voice support.

7:32And so today, like, you know, there are like features that can the T -Mobile app that if you call actually handled by open -air models behind the scenes, and it does sound like super natural human sounding latency -quality -wise. That one was really fun. A second one, which is very interesting. Just on that, can I ask you a follow -up question? We've got text models, we've got voice models, maybe even video models someday that are deployed at T -Mobile. But what above the models or adjacent to the models might we have helped T -Mobile with, for example? Yeah, there is a tongue we're doing. The first one is, you know, you have to put yourself in the shoes of an enterprise buyer.

8:09Like, their goal is to automate, you know, reduce, like, you know, optimize customer support. And, you know, going from, like, a model, like tokens in tokens out to that case, it's hard. And so, you know, first, like, there is a lot of design, like, you know, system design. We do have actually now forward deployed engineers who are helping us quite a bit. For the foreign engineers. Yeah, that's familiar to the far -of -the -term from Palantir. Yeah, it's a great term. Well, you at the ease at Penentio? I was not an FD, I was on, I think they called it the DevSide, right, it's like side -frightering.

8:37I was also only an internet at Palantir, but yeah, it's a great term. I think it accurately describes what we're asking folks to do, which is like embed very deeply with customers and honestly like build things specific to their systems. They're deployed onto these customers. But yeah, we are obviously growing and hiring that team quite a bit because they've been very effective, like a team mobile. And you're just with my life. Yeah, yeah, yeah, yeah, forward deployed. Yeah, but go ahead so forward to the ploy of engineering on the plant engineers and the sort of like systems and like integration is that doing is you know first like you know You have to orchestrate those models like those models are not just you know Those models like no nothing about like you know the CRM like you know and like what's going on And so you have to plug the model to like many many foreign digital tools Many of those like tools like in the enterprise to not even have like API is like clean interfaces, right?

9:23It's the first time they're being exposed like you know to a third party a system And so there is a lot of you know standing up like you know like get ways like tools connecting Then you have to essentially like define what good looks like, you know Again like to pray a new exercise for everyone like you know defining like a golden set of evils Is you know easier than sound? Harder than it sounds And so we're spending like a bunch of time with them evils are important. You've got a super important Especially like audio evils and audio evils are like extra hard to grow and get right and but like a bulk of the use case here is actually audio and have like five minute call transfer power, you actually know that the right thing happened.

10:01It's a pretty tough one. Yeah, it's pretty tough. And then, you know, actually nailing down like the quality of the customer experience, like, you know, until it feels unnatural. And here latency and interruptions, they're really like, you know, important part. We shipped in GA and API, built -in API. I think it was that's right. A couple of weeks ago. Yeah, that's a nice thing. I think it is like a beautiful work of engineering. There was a really cracked team behind the scenes, which basically allows us to get the most natural sounding, like voice experience without having these weird interruptions on your lag, where you can feel that essentially the thing is off.

10:41So yeah, cobbling with that together, and you get that really good experience. Yeah, that's a lot more than just models. Yeah. Yeah, I was going to say, one actually really great thing that I think we've gotten from the T -Mobile experience is actually working with them to improve our models themselves. So for example, the real -time GA last week, we obviously released a new snapshot of the GA snapshot. And a lot of the improvements that we actually got into the model came out of the learnings that we have from T -Mobile. It brings in a lot of other change from other customers, but because we were so deeply embedded into T -Mobile and we were able to understand what good looks like for them, we were able to bring that to some of our models.

11:16That makes sense. And so this is a large customer with tens of millions of users, if not hundreds of millions. And the before and after is on the support side, both tech support internally and then their customer support. Yeah. Makes sense. Is there another one that you guys can share? I like a lot. I'm Jen. I'm Jen, the health care business. I'm Jen, yeah. So we are working quite a bit with healthcare companies. I'm Jen is one of the leading healthcare companies. we specialize into drugs for cancer or inflammatory diseases that are based out of LA. And we've been working with AppGen to essentially speed up the drug development and the conversation process.

11:57So the sort of the North Star is pretty bold. And it's really interesting. When you similarly, we embedded pretty deeply with AppGen to understand what other needs. And it's really interesting, like when I look at those healthcare companies, I feel like there are two big buckets of needs. What is like pure R &D? It's like you're seeing like a massive amount of data and like you have super smart scientists who are trying to, you know, combine, test out things, you know, so that's one bucket. A second bucket is like, you know, much more like, you know, common across other industries. It's like pure, like, you know, admin, document, authoring, document, script, you know, which is, you know, So by the time your R &D team has essentially locked the recipe of a medication, getting that medication to market is a ton of work.

12:43You have to submit to various regulatory bodies, get a ton of reviews. And when we looked at essentially those problems, what we knew, what models were capable of, we saw a ton of benefits, a ton of opportunities to automate and augment essentially the work of those teams. And so yeah, I'm John, there's been like a top customer of DP5 for instance. Wow. I mean, this could be hundreds of millions of lives if a new drug is developed faster. Yeah. Exactly. Huge impact. So that's, you know, that's I think one good example of like a kind of impact on which you need to enable enterprises like to do it.

13:17Right. And so I think we're going to do more and more of those. And yeah, frankly, like, you know, on the personal level, like it's a delight, you know. Yeah. I can play like a tiny role essentially, like doubling the kind of medication that people get in the real world that shows a pretty good achievement. Huge. Huge, huge. I know you had one. Yeah. So one of my favorite deployments that we've done more recently actually is with the Les Alamos National Labs. So this is the government national research lab that the US government is running in Les Alamos, New Mexico. It's also where the Manhattan Project happened back in the 40s and 50s.

13:52Back when it was a secret project. So after that, they ended up formalizing it as a city and a program and then now it's a pretty sizable national Laboratory. This one is very interesting because one just the depth of impact here is like unimaginable for me It's like on the scale of of Amgen and some of these other larger companies, but you know, obviously they're doing a lot of actual new research there So a lot of new science they're doing a lot of stuff with our defense department and defense use cases as well So very intense, you know, very intense stuff But the other thing I was actually very interesting about this one was that it's also a story of a very bespoke and new type of deployment that we've done.

14:31So because they are so their government lab, they're so restrictive and high security and high clearance with a lot of their things, we couldn't just do a normal deployment with them. They couldn't, you know, you can't have people doing national security research just hitting our RAPIs. And so we actually did a custom on -prem deployment with them onto one of their supercomputers called Banado. And so this actually involves a bunch of, you know, very bespoke work with some FDs, also with a lot of our developer team, to actually bring one of our reasoning models, 03 into their laboratory, into an air -gapped, you know, supercomputer of Banado, and actually deploy it, and get it installed to work on their hardware, on their networking stack, and actually run it in this particular environment.

15:15And so it was actually very interesting because we literally had to bring the weights of the model physically into their supercomputer, in an environment by the way, where you're not allowed to have, you know, it's very locked down for a good reason. You're not allowed to have like cell phones or like electronics with you as well. And so I think that was a very, very unique challenge. And then the other interesting thing about this point, I mean, it's just how it's being used, right? So the interesting thing is because it's so locked down and on -prem, we actually do not have much visibility to exactly what they're doing with it, but we do have, you know, they give us feedback.

15:48Yeah, yeah. They actually do have some telemetry, but it's within their own systems. But we do know that it's being used for a bunch of different things is being used for aiding them in terms of speeding up their experiments. They have a lot of data analysis use cases, a lot of notebooks that they're running with rings of data that they're trying to process. us. They're actually using it as a thought partner, which is something that's pretty interesting to me. 03 is like pretty smart as a model. And a lot of these people are tackling really tough, you know, novel research problems. And a lot of times they're kind of using 03 and going back and forth with it on their experiment design on like what they actually should be using it for, which is something that we can really say about our older models.

16:29And so yeah, it's just being used by, for a lot of different use cases for the, for the and the other cool thing is it's actually being shared between Los Alamos and some of the other labs Lawrence Livermore, San Dia as well, because it's the supercomputer setup where they can all kind of connect with it remotely. Fascinating. I mean, we've just gone through three pretty large scale enterprise deployments, right, which might touch tens if not hundreds of millions of people. But there was this on the other side of this is the MIT report that came out a couple of weeks ago. So 95 % of AI deployments don't work, a bunch of scary headlines that even shook the markets for a couple of days.

17:09Like, put this in perspective, like for every deployment that works, there's presumably a bunch that don't work. So maybe we can, maybe talk about that, like what does it take to build a successful enterprise deployment, a successful customer deployment, and the counterfactual based on all your experience, serving all these large enterprises? I think at that point I may have worked with a couple of hundreds. I think enterprises. So, okay, I'm going to a pattern match. What I've seen being clear leading indicator of success. Number one is the interesting combination of top -down, buy -in, and enabling very clear group of a tiger team, essentially, an enterprise, which sometimes makes up open AI, an enterprise employee.

17:56So, typically, you take the top leadership points that you can get in the body, like a priority, but then letting the team organize and be like, if you want to start small, start small, and then you can scale it up, essentially. So that would be part number one. So top down by an and a bottom called a tiger team. Tiger team, people, a mix of technical skills and people who just have the organizational knowledge, institutional knowledge. It's really funny. In the enterprise, customer support, particularly example, what we found is that the vast majority of the knowledge is in people's heads. It's probably like a thing that we have these in general.

18:37But you take customer support, you think that everything is perfectly documented, like GRI, et cetera. The realties, the standard operating procedures like the SOPs are largely in people's heads. And so unless you have that tag your team, mix of technical and subject value expert, really hard to get something at the ground. That would be one. Two would be eVals first. Like whenever we define that good eVals, that gives a clear clear common goal for people to hit, whenever the customer fails to come up with good eVals, it's a moving target, essentially if you've made it or not. And eVals are much harder than what it looks to get done.

19:17And eVals also oftentimes need to come up bottom up, right? Because all of these things are kind of in people's heads in the actual operator's heads. Like it's actually very hard to have a top down mandate of like, you got like, this is how the eVals should look. A lot of it needs the bottoms of adoption. Right. Yeah. Yeah. Yeah. And so we've been building quite a bit of tooling on eVals. We have like a Niveals product and you know, we're working on more to essentially solve like, you know, that problem or you know, make it as easy as we can. The last thing is, you know, you want to help time essentially.

19:44You have your eVals, the goal is to get to 99%. And you started 46, how do you get there? And here, frankly, I think, oftentimes, a mix of, I will say almost wisdom from people who've done it before. A lot of that is art, sometimes more than science, or knowing the course of the model, the behavior. Sometimes we even need to find tune ourselves to models when there are some clear limitation, and being patient, getting your way up there, and then ship. Can we go under the hood a little bit? You know, one of the things that we think about a lot is autonomy more broadly, right? What is the makeup of autonomy?

20:25On the one side, you know, in San Francisco, you could take a car from one part of herself to the other fully autonomously. No humans involved, nope. You press about it. No humans. Right? They've done billions of rides. I think it was like what, three and a half billion rides. It has this on the Tesla FS. I think Wem was done like millions, tens of millions of rides. That's a lot of autonomy. In the physical world, as opposed to the digital world, I can't book a ticket online right now. There's all sorts of problems that happen if I have my operator try to book a ticket. And it's very counterintuitive because the bar for physical safety is so much higher.

21:02The bar for physical safety is higher than the human's capability because lives are at stake. The bar for digital safety, not that high because all you're going to lose is money. Nobody's life is at stake, but yet, physical autonomy is ahead of digital autonomy in 2025. What seems counterintuitive? Like, why is that the case at, you know, at a technical level? Why is it that what should sound easier is actually a lot harder? Yeah, so I think there are kind of two things at play here. And I really like the analogy with self -driving cars because they've been, they've actually been like one of the best applications of AI.

21:37I think that of, yeah, of users recently. But I think there are two things in play. One of them is honestly just the timelines. Like we've been working on self -driving cars for so long. I remember when I, you know, back in like 2014, it was kind of like the advent of this and everyone was like, oh, it's happening in like five years. Turns out it took like, I don't know, 10, 15 years or so for this time. So there's been a long time for the technology to really mature. And I think there's probably like dark ages, you know, back in like 2015 or 2018 or something where it felt like it wasn't gonna happen.

22:05I'm sure off of delusionment. Yes, yes, yeah. And then now we're finally seeing it get deployed, which is really exciting. But it has been like, I don't know, 10 years, maybe even 20 years from the very beginning of the research. Whereas I think AI agents are like really in day one. Here, like Chatchy PT only came out in 2022. So around three, less than three years ago. But I actually think what we think about with AI agents and all that really, I think, started with the reasoning paradigm that when we released the Owen preview model back in late last year, I think. I actually think this whole reasoning paradigm with AI agents and the robustness that those bring has only really unfolded for like a year, less than a year, really.

22:48I know you had a chart in your blog post, which I really like, which the slope is very meaningfully different now. Self -driving started very, very early. Slope seems to be a little bit slower, but now it's reaching the promise land. But man, we started super recently with AI agents, and the slope, I think, is incredibly steep, and we'll probably see a crossover at some point. Yeah. But we really have only had like a year really to explore these things. Do you think we haven't crossed over already when you look at like the coding work in particular? Yeah, it's a good point. It's like, you know, your chart actually shows AI agents as below self -driving.

23:19But like, you know, it's like, what is the why access? Like by some measures, like I would not be surprised actually if, you know, AI products are AI agents, products are making more revenue than WAMO at this point. Like WAMO is making a lot, but like just look at all the startups coming up. That's a point. you know, chat GPT and how many subscriptions are happening there and all of that. And so maybe we have actually crossed and you know, a couple years from now, it's going to look very, very different. Yeah, the Vi -axis is tangible, felt autonomy. Yeah, yeah. Don't make the objective. I don't see.

23:47How do I feel about the actually? We've lived more than revenue. But revenue is a good one. We should probably redo that with revenue. The, there's a second thing I wanted to mention on this as well, which is the scaffolding and the environment and which these things operate in. So I actually remember in the early days of self -driving, a lot of the like researchers around self -driving, we're saying that the roads themselves will have to like, change to accommodate self -driving. I think it might be like sensors everywhere so that the self -driving car is going to interact with it. Which I think is like, you know, retrospect overkill.

24:16But I actually do think self -driving cars have a good amount of scaffolding in the world for them to operate in. It's like not completely like unlimited. You have roads, roads exist. They're pretty standardized. You have stoplights. It's people generally operate in like pretty normal ways, and there are all these traffic laws that you can learn. Whereas AI agents are just kind of dropped in the middle of nowhere, and they kind of have to feel around for them. And I actually think, you know, going off of what Olivier just said too, my hunch is some of the enterprise points that don't actually work out, likely don't have the scaffolding or infrastructure for these agents to interact with as well.

Read the full transcript

24:52A lot of the like really successful deployments that we've made, a lot of what our FDs and up doing with some of these customers is to create almost like a platform or some type of scaffolding connectors organizing the data so that the models have something that they can interact with in a more standardized way. And so my sense of self -driving cars actually have had this in some degree with roads over the last, you know, over the course of their deployment, but I actually think it's still very early in the AI agents. Space and I would not be surprised of a lot of these, a lot of enterprises, a lot of companies just don't really have the scaffolding ready.

25:23So if you drop an AI agent in there, It kind of doesn't really know what to do and it's impact will be limited. And so I think once this scaffolding gets built out across some of these companies, I think that a plane would also speed up. But again, to our point earlier, I think there's no slowdown, there's no, you know, things are still moving very fast. That's great. Well, you know, I've thought about autonomy as a three -part structure. You've got perception, you've got the reasoning, the brain, and you've got the scaffolding, the last mile of making things work. Maybe we can dive into the second part, which is a reasoning, which is the juice that you guys are building with GPT -5, most recently, huge endeavor, congrats.

26:02The first time you guys have launched a full system, not a model or a set of models, but a full system. Talk about that. I mean, the full arc of that development, what was your focus? I mean, honestly, the benchmarks all seem so saturated, like, clearly it was more than just benchmarks that you were focused on. And so what was a North Star like tell us about GP5 soup to nuts? It's been the work of love of many people for a long time And to your point, I think GP5 is amazingly intelligent you look at the benchmark like you know the sweet bench and the likes It is going pretty high But I think to me equally important and impactful was I would say the craft like the style the tone the behavior of the model So you know capabilities intelligence and you know behavior of the model On the behavior of the model, I think it's the first model, like large model release, for which we have worked so closely with a bunch of customers for like a month and month.

26:59Essentially, to better understand like what are the concrete locks? Like what are the concrete blockers of the model? And often, it's not about having a model which is way more intelligent, a model which is faster, a model that better follows instruction, a model that is more likely to say no, when he doesn't know about something. And so that super close customer feedback loop on DP5 was pretty impressive to see. And I think all the love that DP5 has been getting in the past couple of weeks, I think people are starting to sell that essentially, the builders. And once you see it, it's really hard to come back to a model which is extreme intelligent, but an exclusively academic, essentially way.

27:43Yeah, yeah. Are there trade -offs that you made as you were going through it? Like maybe what are the hardest trade -offs you made as you were building GPD -5? I actually think a very clear trade -off, which I think I honestly think we are still iterating on is the trade -off between the reasoning tokens and how long it thinks versus performance. Yeah. Yeah. And I honestly, this is something that I think we've been working on with our customers since the launch of the reasoning models, which is these models are so, so smart. especially if you give it all this like thinking time. I think the feedback I've been seeing around GPT -5 Pro has been pretty crazy too.

28:17It's like, you know, these 100 - and 3 .9. Yeah, I saw that Sam read it. But like, these like unsolved problems that none of the other models could handle. You throw to GPT -5 Pro and it just like one shots it is pretty crazy. But the trade -off here is you're waiting for 10 minutes. It's quite a long time. And so these things just get so smart with more inference time. But on the product builder on the API side for some of these business use cases, I think it's pretty tough to manage that trade off. And for us, it's been difficult to figure out where we want to fall on that spectrum. So we've had to make some tradeoffs on how much of the model think versus how intelligent it should get.

28:54Because as a product builder, there's a real latency tradeoff that you have to deal with, where your user might not be happy waiting 10 minutes for the best answer in the world. might be more okay with the substandard answer and like, no, it at all. Yeah, I mean, even between GPD five and GPD five thinking, I have to toggle it now because sometimes I'm so impatient, I just want to ASAP. Yeah, I think there's a bit of a ability to skip, right? And it's a great way to do it. That's right. I'm impatient, I just want to add a more simple answer. That's right, that's right. Well, four weeks in, GPD five, how's the feedback?

29:24Yeah, I think the feedback has been very positive, especially on the platform side, which has been really great to see. I think a lot of the things that Olivier mentioned have been come up in feedback from customers. The model is extremely good at coding, extremely good at reasoning through different tasks, but especially for coding use cases, especially when it thinks for a while, it will usually solve problems that no other models can solve. I think that's been a big positive point of feedback. The robustness and the reduction hallucinations has been a really big positive feedback. I think there's an evil that showed that a hallucination is basically went to zero for a lot of this.

30:03It's not perfect. There's a lot of work to be done, but that's a big one. I think because of the reasoning in there too, it just makes the model more likely to say no, less likely to hallucinate answer. So that's been something that people have really liked as well. Other bit of feedback has been around instruction following. So it's a really good instruction following. This almost bleeds into the constructive feedback that we're working on. We're afraid that it's so good at construction following, that instruction following, that people need to tweak their prompts or it's almost like too literal.

30:30That's one of the interesting title of actually. Because when you ask people, developers, like what do you want? You want them all to follow instructions, of course. But once you have a model that is extremely little, essentially, then essentially forces you to express extremely clearly what you want. Otherwise, they might make your sideways. And so that was an interesting feedback. It's almost like the monkey power, where it's developers and platform customers ask for better instruction following. Yes, we'll give you really good instruction following. but it's like, you know, follows it almost to a T.

30:59And so, it's obviously something that the team is actually working through. I think a good example of this, by the way, is some customers would have these prompts. I remember when we were testing GPD -5, one of the negative feedback that we got was the model was too concise. We were like, what's going on? Why is the model so concise? Interesting. And then we realized it was because they were reusing their old prompts from other models. And with the other models, they have to like, you have to like, really beg the model to be concise. So they're like, 10 lines of like, be concise, really be concise.

31:26Also, keep your answer short. And it turned out when you give that to GPD -5, it's like, oh my gosh, this person really wants it to be concise. And so the response would be like one sentence, which is too turs. And so just by removing the extra prompts around being concise, the model behaved in a much better way, and much closer to what they actually animal. Yeah. Turns out, writing the right prompt is still important. Yes. Yes, yeah. Prompt engineering is still very, very important. Yeah. On constructive field effort, GPD -5, there's actually been a good amount as well, which we're all working through.

31:55One of them that I think is, I'm really excited for the next snapshot to come out to fix on this, is code quality and small code, like paradigms or idioms that they might use. I think they're feedback around the types of code and the patterns in which it was using, which I think we're working through as well. Then the other bit of feedback, which I think we've already made good progress on internally, is around the trade off of the reasoning tokens and thinking and latency around intelligence, I think, especially for the simpler problems, you don't usually need a lot of thinking, the thinking should ideally be a little bit more dynamic.

32:31And of course, we're always trying to squeeze as much reasoning and performance into as little reasoning tokens as possible. So I had to mention that that could kind of go down as well. Yeah. Well, huge congrats. I mean, it's been, I know it's a work in motion for a bunch of our companies they've had incredible outcomes with GPD five, one of them's expo, cybersecurity business, just like a huge difference. from that, it was pretty crazy. Huge, huge upgrade from whatever they were using prior to that. And I think they're going to need a new e -vow. See, that's right. They're going to need a new e -vow.

32:59It's all about e -vails. On the multi -modality side of it, obviously, you guys announced the real -time API last week. So a team mobile was one of the featured customers on there. Talk about that, how obviously the text models are leading the pack. But then we got audio and we got video. Talk about the progress on the multi -model models. when should we expect to have the next big unlock, and what would that look like? It's a good question. The teams have been making amazing posts on multimedia. On voice, image, video, frankly, the last generation models have been unlocking quite a few cold -use cases.

33:34One of the feedback that we've received is, because like text was so much like the back on intelligence, like people felt like in virtual on voice, that the mall was somewhat a little less intelligent. And until you actually see it, it does feel weird, like you know, to have a better answer like on text versus voice. And so that's pretty much a focus that we have at the moment. I think we like filled like part of that gap, but not the full gap for sure. So I think you know, catching up, I would say you know, with the text like you know, would be one. A second one, you know, which is absolutely fascinating is the model is like excellent at the moment on like, you know, Easy casual conversation, talk to your coach, your therapist.

34:18And we basically had to teach the model to speak essentially better, like in actual work economically valuable setups. Give an example, the model has to be able to understand what a necessity is and what it's meant to spell a SSN. And if one digit is actually fuzzy, it should actually have to repeat versus guess. There are lots of intuitions like that that someone, of course, has, of our voice, that we are currently teaching the model and that's like an ongoing work actually with our customers. Until we actually confront the model to actual customer support calls, actual sales call, it's really hard to get a fill for those gaps.

34:56So that's a top priority as well. This is a completely off -scrap but an interesting question that comes up in voice models, particularly the real -time API is previously people were doing, they were taking a speech input, convert that to text, then have some layer of intelligence. Then you would have a text to speech model that would sort of play it back. And this would be a stitch of these three parts. Well, the real -time API, you guys have integrated all of that. Yes. And how does it happen? Because a lot of the logic is written in text. A lot of the Boolean logic, or any call it any function calling, is written in text.

35:34How does it work with the real -time API? Is that... That's an excellent question. So the reason why we should do a real -time API is that we saw that for the stitch model. The stitch model. Yeah, the real -time stitch. Yeah, the stitch model. Like a stitch together. A stitch together. Like, stitch two text, thinking, text to speech. Yeah. Like, we saw essentially a couple of issues. One, slowness, like, you know, my pops, essentially. Two, lots of signal, like, you know, a cross -search model. Like, the speech text model is less intelligent. Yeah, you release the motion. Exactly. Exactly, accent.

36:08Exactly, right, buzzes. And when you are doing actual voice, phone calls, essentially those signals are so important, 40 % to system, one of the challenges that we have is what you mentioned, which means a slight difference architecture, essentially, for text versus voice. And so that's something that we're actively working on. But I think it was the right call to start essentially with let's make the voice experience like natural sounding to a point -wise century, you're feeling comfortable, like putting in production, and then working backward, like to unify the sort of the orchestration logic essentially, a customer that is.

36:46And to be clear, a lot of customers still stitch these together. It's kind of what worked in the last generation, but what we're increasing seeing is more and more customers moving towards the real -time approach, because of how natural it sounds, how much lower latency. It is especially as we up -level the intelligence of the model. But also even taking that back. I will say it's like pretty mind -blowing to me that it works. Like the fact that, like, I think it's mind -blowing that these LLM's work at all, or you just train it on a bunch of texts, and it's just, you know, autoregressively coming up with the next token, and it sounds super intelligent.

37:14That's like mind -blowing in and of itself, but I think it's actually even more mind -blowing that this speech -to -speech set of actually works perfectly, because you're literally taking the audio bits from it, from someone's speaking, or putting it into the model, and then it's generating audio bits back. And so, to me, it's actually crazy that it's works at all. Let alone the fact that it can understand accents and tone and pauses and things like that and then also be intelligent enough to handle a support call or something. If you've gone from text in text out to voice in voice out, that's pretty crazy.

37:43We have a bunch of companies in a portfolio that are using these models, you know, Parlo on the customer support side, live kit on the infrastructure side and you know, there's a bunch of use cases we're starting to see that that a speech speech model could could address. Obviously a lot of the Carter ones still still running on what you're calling the stitch model. Yeah. But I hope the day is not far when it's all on real -time API. It's going to happen at some point. Right. And actually maybe that's a good segue into talking about model customization because I suspect that you have such a wide variety of enterprise customers.

38:16I think you mentioned what hundreds of customers or maybe more. Each of them has a different use case, a different problem set, a different call it envelope of parameters that they're working in, maybe latency, maybe power, maybe others. How do you handle that? Talk about what OpenAI offers enterprises who need a customized version of a great model to make it great for them. Yeah, so model customization is actually even something that we've invested very deeply in on the API platform since the very beginning. So even pre -chatchyBT days, we actually had a supervised fine -tuning API available, and people were actually using it to great effect.

38:52The most exciting thing actually, say around model customization, it obviously resonates quite well with customers because they want to be able to bring in your own custom data and create your own custom version of, you know, 03 or 04 mini or something or GPD -5 even suited to their own needs. It's very attractive, but the most recent development I think is very exciting has been the introduction of reinforcement fine tuning. There's something we announced late last year, I think in the 12 days of Christmas,

39:24we've Yeah, so it's called, it's actually funny. I think we made up the term reinforcement fine tuning. It's like not a real thing until we... It's stuck, no. I can't read all the time. I remember we were discussing it, and I was like, I don't know about our... You were not getting, you were not getting. Yeah, yeah. So reinforcement fine tuning. So it really, it's introducing reinforcement learning into the fine tuning process. So the original fine tuning API does something called supervised fine tuning, call it SFT. It is not using reinforcement learning. It is, you know, it's using supervised programming.

39:57And so what that usually means is you need a bunch of data, a bunch of prompt completion pairs. You need to really supervise and tell exactly the model how to how it should be acting. And then when you train it on our fine tuning API, it moves it closer in that direction. Reinforcement fine tuning introduces like RL or reinforcement learning to the Slope. Way more complex, way more finicky, but in order of magnitude more powerful. And so that's actually what's really resonated with a lot of our customers. it allows you to, if you use RFT, the discussion is less of creating a custom model that's specific to your own use case.

40:28It is, you can actually use your own data and actually crank the RL, yeah, turn the crank on RL to actually create a best -in -class model for your own particular use case. And so that's kind of the main difference here. With RFT, the data set looks a little bit different instead of, you know, prompt completion pairs, you really need a set of tasks that are very greatable. You need a greater, that is very objective, that you can use here as well. And so that's actually been something that we've invested a lot in over the last year. And we've actually seen a couple, a number of customers get really good results on this.

40:59We've talked about a couple of them across different verticals, so RoGO, which is a startup in the financial services space. They have a very sophisticated AI team. I think they hire some folks in deep mind to run their AI program. And they've been using RFT to get best in class results on parsing through financial documents. I have three questions around it and doing tasks around that as well. There's another startup called the Cordance that's doing this in the tax space. I think they've been targeting an evalicle tax bench, which looks at CPA -style tasks as well. Because they're able to turn it into a very greatable setup, they're actually able to turn the RFT crank and also get, I think, Soda results on a tax bench just using our RFT product as well.

41:45So it has kind of shifted the discussion away from just customizing something for your own use case to really leveraging your own data to create a best in class, maybe best in the world model for something that you care about for your business. Yeah, if you're like the the base models are getting so good at instruction for a wing that for you know, behavior like steering like you don't need to find your end at point, you can describe what you want and the model is pretty good at it. But pushing the frontier on actual capabilities, my herchies that our FT will pretty much become the norm. Like, you know, issues are actually pushing in your field, like, you know, intelligence, like, you know, to a pretty high point.

42:22Like at some point, like, you know, you need to aisle, I'll always send you to with custom environments. Yeah, fascinating. And even going back to the point earlier around, like, top -down risk bottoms up for some of these enterprises, a lot of the data that you end up needing for RFT require, like, very intricate knowledge about the exact task that you're doing and understanding how to grade it. And so a lot of that actually comes from bottoms up. Like I know a lot of these stars will work with experts in their field to try and get the right tasks and get the right feedback to craft some of these data sets.

42:51Without further ado, we're gonna jump into my favorite section, which is a rapid fire question. We had a lot of great friends of ours send in some questions for you guys. We'll start with the ultimateers favorite game, which is a long short game. Pick a business and idea, start up that you're long and the same short that you would bet against that there's more hype than there's reality. Whoever's ready to go first, long short. My long is actually not in the AI space, so this is going to be slightly different. There we go. My short is though in the AI space. So I'm actually extremely long eSports.

43:27And so what I mean by eSports is the entire professional gaming industry that's just imagining around video games. Very near and dear to my heart, I play a lot of video games, and so I watch a lot of this. So obviously, I'm pretty in the weeds on this. But I actually think there's incredible untapped potential in esports and incredible growth to be had in this area. So concretely what I mean are like, you know, a really big one is League of Legends, all of the games at Riot Games puts out. They actually have their own professional leagues. They actually have professional tournaments, believe it or not.

43:57They rent out stadiums, actually now. But I just think it's like, if you look at kind of what the youth and like what younger kids are looking and where their time is going, It's predominantly going towards these things. They spend a lot of time on video games. They watch Moe's bolts and the Excel cube as kid bolts, you know. Yeah, yeah, yeah, yeah, a growing number of these two. I've actually been to some of these events and it's very interesting. Yeah, yeah, I'm extremely long stuff. And so they're booking out stadiums where people to go watch electronic sports. Yeah, yeah, yeah. It's uh, I literally went to Oracle Arena, the old where your stadium to watch one of these, I think before COVID.

44:35And then the so it's just before COVID wow that's five years ago. Six years ago. So I actually I've been following this for a while and I just think it had a really big moment in COVID like everyone is playing video games. And I think it's kind of like come back down. So I think it's like undervalued you know it's like I think no one's really appreciating now but it has all the elements to like really really take off. And so the youth are doing it. The other thing I'd say is it is huge in Asia like absolutely massive in Asia. It is absolutely big in Korea and China as well. Like, you know, we rented out Oracle Arena, I think, or like the event I went to was an Oracle Arena.

45:08My sense is in Asia, they rent out like the entire stadiums, like the soccer stadiums, and the players are ready like celebrities. So anyways, you know, as like, you know, the, I know like Korean culture is really making its way into the US as well. I think that's another tailwind for this whole thing, but anyways, esports, I think, is something you should keep an eye out on, because there's a lot of room for growth. Very unexpected. Yeah. Good to hear. Short. My short, my short is a little spicy, which is I'm short on the entire category of like tooling around AI products. And so this encapsulates a lot of different things, kind of cheating because some of these, you know, I think are starting to play out already, but I think like two years ago it was maybe like Eval's products or like frameworks or vector stores.

45:51I'm pretty short those. I think nowadays there's a lot of additional excitement around other tooling around AI models. So RL environments, I think are really big right now as well. Unfortunately, I'm very short on those. Not really, I don't really see a lot of potential there. See a lot of potential and reinforcement learning and applying it, but I think the startup space around RL environments, I think is really tough. Main thing is one, it's just a very competitive space. There's just a lot of people kind of operating in. And then two, if the last two years have shown us anything, the space is evolving so quickly.

46:26And it's so difficult to try and adapt and understand what the exact stack is that will really carry through to the next generation of models. I think that just makes it very difficult when you're in the tooling space, because today's really hot framework or really hot tool might just not get used in the next generation of models. So I've been noticing the same pattern, which is the teams that build like Braytia with Lex Thought -ups. in AI are extremely pragmatic. They are not super intellectual, but the perfect world, etc. It's funny because our generation has basically started in tech in a very stable moment.

47:05Technology has been building up for years and years with access, cloud, etc. We were raised in a very stable moment where it makes sense at that point to design very good abstractions and touldings because you have a sense what's going. but it's so different today. Like the white point knows what's going to happen next year or two, so it's almost impossible like to define like the perfect tuning platform. Right, right, right, right. Well, that's, there's a lot of that going around right now. Yes. Spicy, a lot of homework there. Olivia, over to you sir. Long short. Long short. I've been thinking a lot about education from the past month in the context of kids.

47:45I'm pretty short on any education, which basically emphasizes human memorization at that point. And I say that having mostly been through the education myself, but I learn so much on like, you know, history facts, like, you know, legal things that are. Some of it, like, does shape your way of thinking. A lot of it, frankly, is just like, you know,

48:13in the I'm quite short on that. That's why you only need memory when strategy is bionic. You can just think about straight into your head. What am I long at? Frankly, I think healthcare is probably the industry that will benefit the most from AI in the next lecture or two. I will say more. I think all the ingredients are here for a perfect storm. Mm -hmm. A huge amount of structure and structure data, you know. It's basically the heart of the far -away companies, the MOA's excellent at digesting processing that can of data. A huge amount of admin, heavy documents, heavy culture. But at the same time, companies which are very technical, very are in different companies, who's technology in a way that the heart, what they do.

49:06And so I'm pretty bullish on their health. This is like life sciences. So you mean life sciences, research organizations that are producing drugs. Exactly. Yeah. Yeah. It's almost like, you know, over the last 20, 30 years, these, these like, pharma or like biotech companies have, have basically, if you look at the work that they're doing, like only a small amount of it is, is actual research and so much of it ends up being admin and like, you know, documents and things like that. And that area is just so ripe for, you know, something to happen with AI. And I think that's what we're seeing with Amgen and some of these other customers.

49:36Exactly. And it's also like not what they want to do. It's, I think it's good that we have some regulations there, obviously. but just means that they have like rims and rims of things they kind of go through. And so, you know, when you have a technology that's able to really help bring down the cost of something like that, I think it'll just tear right through it. And I think once governments and institutions are going to realize that, like if you step back, it is probably one of the biggest buttons to human progress, right? You step back in the past decade, like how many true drugs have there been, not that many.

50:08like, you know, how life would be different, like if you double that rate, essentially. So once you realize what it takes, yeah, my hunch is that we're going to see quite a bit of momentum in that space. Wow. All right. Lots of homework there as well. Yeah. Next one. Favorite underrated AI tool? Other than chat GPT maybe. I love granola. Oh, man. I know. It's my idea. Oh my god. Like two votes for granola. There is something like, hey, what about chat GPT record? I like to do physical as well, but there are some features of the granular which I think are really done well Like the whole like you know integration with your Google calendar is excellent.

50:45Yeah And just you know the quality of like the transcription and like the summaries pretty good Do you just have it on because I know your calendar is back to back you just have journal on so the thing is that I don't use granular Intensity I don't have my personal life mostly. Yeah, I see on dates I'm joking. I'll say, yeah, granola is actually going to be mine. So two votes for granola. I was going to say that easy answer for me is codex, as a software engineer. It's just like, so it's gone so good recently. Codex, Cli, especially with GPT -5. Especially for me, I tend to be less time sensitive about like, you know, the iteration loop with coding.

51:22And so, leaning in a GPT -5 on codex, I think, has been really... What about codex has changed? Because, you know, codex has also been through a journey. Codex has been an out for a bit. I remember like it's been launched for like a more than over a year ago. It's like what's what's changed about codex? Yeah, I was actually going to say it's like last has been around for a bit So it's been less than a year for codex. It's the time violation is so crazy in this field. It feels like it's been around for a year ago with GP4 Oh, like you know, I mean that demo like that feels like ages ago one and even come out yet Probably there's a Christmas that happened yet the voice demo is a naming thing.

51:54Okay, but anyway. Yeah. Oh, there was a codex model That's what I'm thinking. There was a codex model. Yeah. Yeah. Yeah. We are we are You're not too blame for that confusion. Also, I think the GitHub thing was called Codex as well. Yes, yes, that's right. But I'm talking about our coding product within ChatGPT, which is the Codex Cloud offering, and then also Codex Cli. So actually, maybe if I were to narrow my answer, a little bit more to Codex Cli, which I've really, really liked. I liked the local environment set up. The thing that's actually made it really useful in the last, I'd say, month or so, is one, I think the team has done a really good job of just getting rid of all the paper cuts, like the small product polish and paper cut things.

52:28It just, it kind of feels like a joy to use now. I feel more reactive. And then the second thing honestly is GPD5. Like I just think GPD5 really allows the product to shine. It's, you know, at the end of the day, this is kind of a, there's a product that really is dependent on the underlying model. And when you have to, you know, like iterate and go back and forth with model, like four or five times to get it right, to get it to like, you know, do the change that you want. Versus having it think a little bit longer and it just like one shot and does exactly what you want to do. You get this weird, like, bionic feeling where you're like, I feel so mind -melded with the model right now, and perfectly understands what I'm doing.

53:04And so getting that dopamine head and feedback loop constantly with Codex has made it kind of like an indispensable thing that I really, really like. And the other thing I'd say Codex is just really good for me is, so I use it for personal projects. I also use it to help me understand code bases like as an engineering manager, now I'm not as in the weeds on the actual code. And so you're actually able to use codecs to really understand what's happening with the code base, have it like ask questions and have an answer about things and really catch up to speed on things as well. So like even the non -coding use cases are really useful with codecs, Clay.

53:39Fast and having Sam had this tweet about codecs usage ripping. I think like yesterday, so I wonder what's going on there, but you're not alone. Yeah, I think I think I'm not alone just judging from the Twitter feedback. I think people are really realizing how great of a Com motion codex client GPT -5R. Yeah, I know that team is undergoing a lot of scaling challenges, but I mean it the system hasn't gone down for me So props to them But we are energy people crunch that we'll see how you know how long that goes awesome awesome All right the next one Will there be more software engineers and 10 years or less?

54:17There's about 40 -50 million. Who's the most efficient software engineers? That's what you mean, the cool time, the scheduled jobs. Yeah, because this is a hard one. Cause like, I think without a doubt, there's gonna be a lot more software engineering going on. Yes, of course. There's actually a really great post that was shared, I think in our internal, it was like a Reddit post recently. I actually think that highlights this. Is a really touching story. It was a Reddit post about someone who has a brother who's nonverbal, and I actually don't know if you saw this, it was just posted. It's a person on Reddit posted, they have a nonverbal brother who they have to take care of.

54:47The brother, they tried all these types of things to help the brother interact with the world use computers, but vision tracking didn't work because I think his vision wasn't good. All the tools didn't work. And then this brother ended up using ChatGBT. I don't think he used Codex, but he used ChatGBT and basically taught himself how to create a set of tools that were tailor -made to his nonverbal brother. Basically, a custom software application just for them. And because of that, he now has this custom setup that was written by his brother and allows him to browse the internet. I think the video was like him watching the Simpsons or something like that, which is really touching.

55:21But I think that's actually what we'll see a lot more of like this guy's not a professional software engineer, his title is not software engineer. But he did a lot of software engineering, but probably pretty good enough definitely for his brother to use. So the amount of code, the amount of like building that'll happen I think is just gonna go through an incredible transformation. I'm not sure what that means for software engineers, like myself, maybe there's equivalent or maybe there's... Of course more showing. Yeah, more of me. More means this is a bit... We've all been... That's right. But definitely a lot more software engineering and a lot of...

55:50Yeah. I like that completely. I like completely the thesis that there is a massive software shortage. Yeah. Like in the world. Like we've been sort of an... Accepting it. You know, for the past 20 years, like the goal of software was never to be that super rigid, super hard to build, you know, artifact. It was to be like, you know, customized, like, malleable. and so I expect that we'll see way more, sort of a reconfiguration of people's job in scale set where way more people code, I expect that product managers are going to code more and more, for instance. But you made your PM's code recently.

56:25Oh yeah, we did that, that was very fun. We started not doing PRDs, product requirements documents. Classic PM thing, you write like five pages, like my product does, etc. And, you know, PMs have been basically like coding prototypes. And one is pretty fast with GTI pipes and like codecs. Yeah, just a couple hours, I think. Reaking fast. And second, it sort of conveys like so much more information. I don't recommend. Yeah, yeah. You get a feel essentially for the feature like is it just a write or not? So yeah, I expect that sort of, you know, behavior we're going to see more and more. Yeah. Instead of writing English, you can actually now write the actual thing you want.

57:02Yeah. And that's amazing. Advice for high school students who were just starting out their career. My advice is, I don't know, maybe it's evergreen, like, prioritize critical thinking above anything else. If you go in a field, which requires like extremely high critical thinking, like, you know, skills, I don't know, math, physics, or, you know, maybe trophies in that bucket, you will be fine, regardless. If you go in a field that sort of turns out that thing, and again, guess back to the memorization, like you know, parlant matching, I think you will probably be less fit to prove. Yeah. As a, you know, what's a good way to sharpen critical thinking?

57:43Um, use trashy BT and have it test you. That's true. Having like, you know, workplace tutor who essentially knows how to put the bar like 20 % but what you can do all the time, you know. He's actually probably like a really good way to do it. Yeah. Nice. Anything from you, sir? Mine is, I think it's just, I think we're actually in such an interesting, like, unique time period where the, like, younger, so, like, maybe there's a more general advice for not just, like, high school students, but just, like, the younger generation, even, like, college students. It's, like, I think the advice would be, don't underestimate how much of an advantage you have relative to the rest of the world right now because of how AI native you might be or how interesting.

58:30like, you know, in the weas of the tools you are. My hunch is like, high schoolers, college students, when they come into the workplace, they're gonna have actually a huge leg up on how to use AI tools, how to actually transform the workplace. And my push for like some of the younger, I guess high school students is like, one, like, just really immerse yourself in this thing. And then two, just like, really take advantage of the fact that you're in a unique time where like, no one else in the workforce really understands these tools as deeply probably as you do. A good example of this is actually we had our first intern class recently at OpenAI, a lot of software interns.

59:04And some of them were just like the most incredible cursor power users of like ever seen. They were so productive. Yeah, I was shocked, for the way. Yeah, I was like, yeah, I know we can get good interns, but like I don't know they'd be like this good. And I think part of it's just like they've grown up using these tools for better or worse in college. Yeah. But I think the meta level point is they're so like AI native. And even like, I don't know, me and Olivier, We were like kind of AI and native, we were open AI, but we haven't been steeped in this and kind of grown up in this. And so the advice here would just be like, yeah, leverage that, don't be afraid to kind of go in and spread this knowledge and take advantage of it in the workplace because it is a pretty big advantage for them.

59:43Yeah, I can't remember who said this to us at Bounder, but every intern class was just getting faster, smarter, like laptops, like smarter every generation. You sure didn't peak in 2013 when I was a natural. That's right. That's 113, yeah. Two guys like you. That's right. That's right. That's right. That's right. Yeah. Well, lots happened. You're lots happened, since you guys joined OpenAI, right? With three years and almost three years. In your OpenAI journey, what has been the rose moment, your favorite moment, the bud moment, where you're most excited about something, but still opportunity ahead, and the torn toughest moment of your three -year journey.

1:00:23The tone is easy for me. What people do, blip, which is the crew of the board. That was a really tough moment. It's funny because after the fact, it actually reunited quite a bit the company. There was a feeling, open air had a pretty strong culture before, but there was a feeling of camera -adjury, that was even stronger. But sure, like tough on the day off. It's very rare to see that anti -fragility. Most orcs after something like that break apart, but I feel like opening, I got stronger, opening, I came back. It's a good point. I feel it made to put an extra on for real now, essentially, when they look after the fact, when they look at other news, like different shows, whatever, like, you know, bad news, essentially, I feel the company has built, like, you know, a thicker skin, and, you know, an ability to like recover, like, wake weaker.

1:01:12And she think part, I think it's definitely right. Part of it, too, I think is also just the culture. I also think this is why it was such a low point for a lot of people. So many people just at opening a care so deeply about what we're doing, which is why they work so hard. You just care a lot about the working, almost feels like your life's worked. Like it's a very audacious mission and thing that you're doing, which is why I think the blip was like so tough on a lot of people, but also is what I think help bring people back together and we were able to hold together and get that thick skin as well.

1:01:39Yeah. I have a separate worst moment, which was the big outage that we had in December of last year, if you remember. Yeah, you remember. I do. It was like a multi -hour outage, really highlights to us how essential of almost like a utility the API was. So the background is I think we had like a three, four hour outage some time in November, December, last year. Really brutal, pure set of zero. No one could hit tragedy, but you know, it hit the APIs. It was really rough. That was just really tough just from a like, you know, customer trust perspective. I remember we like talked to a lot of our customers, we kind of like post -mortem them on what happened and kind of our plan moving forward.

1:02:20Thankfully, we haven't had anything close to that since then. And I've been actually really happy with all the investments we've made in reliability over the last six months. But in that moment, I think it was really tough. On the happy side, on the roses, I think I have two of them. The first one would be GPT -5 was really You would like to sprint up to GPT -5. I think really showed the best of OpenAI, having like cutting edge, like science research, like extreme customer focus, extreme infrastructure and inference talent. The fact that we were able to shape such a big model and excaly to many, many, many, many tokens per minute, almost immediately, I think speaks to it.

1:03:11So that one I really with no addages. You know it or just yeah really good reliability. Yeah. Like I remember when we shipped like deep for turbo like a year ago and half ago we were terrified by you know like the in -spec of traffic and I felt we've really got on like much better at you know shipping like those you know massive updates. The second rose like you know happy moment for me would be the first deaf day was really fun. It felt like coming over. It's like an open AI. We are embracing that we have a huge community developers. We are going to ship models and products. I remember basically seeing all my favorite people open AI on the earth.

1:03:50You know, essentially, nirthing out on what are you building, what's coming up next. It felt really a special moment in time. That was actually going to be mine as well. I'll just pay you back off of that, which is the very first DevDay 2020 -3 November. I actually, I remember it, so I mean, obviously a lot of good things have happened since then. There's just a very, I don't know why, for me it was a very memorable moment, which was one, it was actually quite a rush up to DevDay, we shipped a lot, so our team was just really, really sprinting, so it was like this high stress environment kind of going up.

1:04:24To add to that, of course, because we're opening a eye, we did a live demo on Sam's keynote of all the stuff that we shipped. And I just remember being in the back of the audience, sitting with the team, and waiting for the demo to happen. Once it finished happening, we all just let out a huge sigh of relief. They're like, oh my God, thank you. And so there's just a lot of build up to it. For me, the most memorable thing was I remember right after DevDay, all the demos worked well, all the talks worked well. We had the after party, and then I was just in a way mode, driving home at night with the music playing.

1:04:55It was just such a great end to the DevDay. That was what I remember. That was my rose for the last - I love it. That's awesome. I assume you guys are, but please tell me if you're a GIPL, yes or no. And if so, what was the moment that got you there? What was your aha moment? When did you feel the AGI? I think I made J -Pield. I think I made J -Pield. You're definitely a J -Pield. I am, OK. I've had a couple of them. The first one was the realization in 2020 that I would never need to code manually in a cable ever again. I'm not the best color, frankly. I chose my job for a reason. But realizing that what I thought was a given, that we humans would have to write basically machine language forever is actually not a given.

1:05:43And that the pace of crisis huge, feeling the AGI, the second field the AGI moment for me was maybe the progress on voice and multi -modality. And you know, text, like at some point you get used to it. Like, okay, you know, the machine can write pretty good text. Yeah, voice makes it real. But once you start actually talking, like, you know, to something that you really understand, or tune, like, you know, understand my accent, like in French, it felt like sort of a, a totally moment, like, okay, machines are going beyond, like, cold, mechanical, deterministic, like, you know, like logic, to something like much more, like, emotional, and like you know, tangible.

1:06:25Yeah, that's a great one. Yeah, mine are, so I do think I am a GI pill, like probably gradually became a GI pill over the last couple of years. I think there are two, and for me, yeah, I think I actually get more shocked from the text models. I know the multimodal ones are really great as well. For me, I think they actually line up with two general breakthroughs. So the first one was right when I joined the company in September, 2022. We is pre -chatch GPT. Yeah, two months ago. About the time GPT -4 already existed internally. And I think we were trying to figure out how to deploy. I think Nick really talked about this a lot earlier as a chatch GPT.

1:07:03But it was the first time I talked to GPT -4. And it was like going from nothing to GPT -4 was just the most mind -blowing experience for me. I think for the rest of the world, maybe going from nothing to GPT -3 .5 in chat was maybe the big one and then going from 3 .5 to 4. But for me, and I think for a lot of maybe some other people who join around that time going from nothing to or or not Nothing but like what was publicly available at the time going from that to GPD 4 was just incredible Like I just remember asking throwing so many things out I was like there's no way this thing is gonna be able to give an intelligible answer and it's like knocks it out of the park It was it was absolutely incredible GPT4 was insane.

1:07:38I remember GPT4 came out when I was in interview with open AI And I was still on the fan should I join so that thing I was like, okay, I'm in I mean guys, there is no way I can walk on it if he else. That's cool. Yeah, yeah. Yeah, so GBD, yeah, GBD4 was just crazy. And then the other one was, like, is the other breakthrough, which is like the reasoning paradigm. I actually think the purest representation of that for me was deep research and throwing, like, like asking it to like really look up things that I didn't think it would be able to know and seeing it like think through all of it, like be really persistent with the search, get really detailed with the writeup and all of that.

1:08:16that was pretty, pretty crazy. I don't remember the exact query that I threw it, but I just remember, I feel like the field age, I moments for me are like, I'll throw something at the model that I was like, there's no way this thing we'll really get. And then it just like knocks out of the park. Like that is kind of the field age, I moment. I definitely had that with deep research with some of the things that I was asking. Yeah. Well, this has been great. Thank you so much, folks. You guys are building the future. You guys are inspiring us every day and appreciate the conversation. Yeah, thank you so much.

1:08:43Thanks for having us.

1:08:52As a reminder to everybody, just our opinions, not investment advice.

From the publisher

Open Source bi-weekly convo w/ Bill Gurley and Brad Gerstner on all things tech, markets, investing & capitalism. This week, guest host Altimeter’s Apoorv Agrawal explores how OpenAI is reshaping enterprise with Sherwin Wu, Head of Engineering OpenAI Platform, and Olivier Godement, Head of Product OpenAI Platform. From T-Mobile’s AI voice support to Amgen’s drug breakthroughs to Los Alamos’ air-gapped supercomputer—this episode dives into the real world of AI at scale. Enjoy another episode of BG2!


Timestamps:

(00:00) Intro

(01:50) OpenAI’s Enterprise Mission: Beyond ChatGPT

(06:00) Case Study: T-Mobile <> Voice & Support 

(11:30) Case Study: Amgen <> Accelerating Drug Development

(13:45) Case Study: Los Alamos National Lab 

(17:00) Why 95% of AI Deployments Fail?

(20:30) Physical vs Digital Autonomy: Scaffolding & Infrastructure

(26:00) GPT-5: Release, Benchmarks vs Behavior

(30:00) GPT-5 Feedback: Instruction Following, Hallucinations, Code Quality

(33:00) Multimodality: Text, Voice, and Video

(35:30) Audio: Realtime API vs Stitched Audio

(38:00) Model Customization & Reinforcement Fine-Tuning (RFT)

(43:00) Rapid Fire: Long/Short Picks 

(1:03:00) Highlights and Lowlights @ OpenAI

Show Notes:

T-Mobile Partnership: https://www.t-mobile.com/news/business/t-mobile-launches-intentcx-with-openai

Amgen Partnership: https://openai.com/index/gpt-5-amgen/

Los Alamos Partnership: https://www.lanl.gov/media/news/0130-open-ai

MIT AI Report: https://mlq.ai/media/quarterly_decks/v0.1_State_of_AI_in_Business_2025_Report.pdf

Produced by Dan Shevchuk

Music by Yung Spielberg

Available on Apple, Spotify, www.bg2pod.com

Follow:

Brad Gerstner @altcap https://x.com/altcap

Bill Gurley @bgurley https://x.com/bgurley

BG2 Pod @bg2pod https://x.com/BG2Pod

Apoorv Agrawal @apoorv03 https://x.com/apoorv03

Sherwin Wu @sherwinwu https://x.com/sherwinwu
Olivier Godement @oliviergodement  https://x.com/oliviergodement

More from BG2Pod with Brad Gerstner and Bill Gurley

All 44 episodes
Inside OpenAI Enterprise: Forward Deployed Engineering, GPT-5, and MoreBG2Pod with Brad Gerstner and Bill Gurley · 1 h 9 min
Listen in VO