Does OpenAI’s new model deliver on the hype? Inside GPT-5 with Jeremy Kahn

9 Aug 2025 · 23 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Pioneers of AI Podcast Episode Notes

Episode Overview

Title

Does OpenAI’s new model deliver on the hype? Inside GPT-5 with Jeremy Kahn

  • Host: Rana el Kaliouby
  • Guest: Jeremy Kahn, AI Editor at Fortune
  • Release Date: [Date not specified in the transcript]
  • Focus: Analyzing the capabilities and impact of OpenAI's newly released GPT-5 model.

Key Topics Discussed Introduction to GPT-5

  • OpenAI has released GPT-5 after significant anticipation.
  • Discussion on whether GPT-5 lives up to the hype and its implications for OpenAI's future.

First Impressions of GPT-5

  • Jeremy Kahn's Opinion: GPT-5 is a good model but does not represent a radical transformation towards Artificial General Intelligence (AGI).
  • Improvement noted over previous iterations, but not a "massive leap forward."

Interface and Functional Changes

  • New Features:
  • Users no longer have to select between reasoning and faster responses; GPT-5 autonomously determines the best mode based on the query.
  • A "router system" that assesses prompts and directs them to the most suitable model variant.
  • Personality Types:
  • Introduction of four selectable personality types for responses (e.g., "Cynic," "Listener," "Robot").

Training Process and Data

  • Details on the training process remain unclear; however:
  • Involves a family of models with various sizes and capabilities.
  • Special emphasis on coding and English writing improvements.

Performance Metrics

  • Coding Capabilities:
  • Claims of 8%-10% improvement on coding benchmarks over previous models.
  • Hallucination Rates:
  • Hallucination rates reduced, but still present; roughly 10% in general mode and around 4.8% in reasoning mode.
  • Comparison with Competitors:
  • Generally on par with other models like Claude 4 and Gemini 2.5 Pro.

Market Position and Business Implications

  • OpenAI's Strategy:
  • Aggressive pricing to maintain competitiveness with Google and Anthropic.
  • Encouraging user retention by enhancing existing subscriptions.
  • User Base Growth:
  • Aiming to attract more users, especially those utilizing the free version.

Employee Retention and Corporate Strategy

  • Discussion about OpenAI enabling share sales to retain talent amid competition from companies like Meta.
  • Mental Health Advisory Board:
  • Acknowledgment of the chatbot's use in mental health support and implementing guardrails for safety.

Health-Related Performance

  • Claims that GPT-5 performs better on health-related queries, despite OpenAI's caution against using it as a medical resource.

Continuous Learning and AGI Discussion

  • Sam Altman asserts that GPT-5 is a step toward AGI, although the model does not learn continuously from user interactions.
  • Discussion on the undefined nature of AGI and what constitutes significant progress.

Concluding Thoughts

  • Jeremy's Final Assessment:
  • GPT-5 is a significant evolution but not a revolutionary leap.
  • The model's improvements are valuable but represent incremental advancements rather than a transformative breakthrough.

Key Takeaways

  • Evolution vs. Revolution: GPT-5 enhances previous models but does not redefine the AI landscape.
  • User Engagement: By improving features and capabilities, OpenAI aims to bolster user loyalty and attract new users.
  • Caution in Applications: OpenAI's role in mental health and health-related advice is complex, emphasizing the need for responsible use.

Contact Information

  • For more updates, visit [Pioneers of AI](http://pioneersof.ai/)
  • Engagement opportunities available via voicemail at 601-633-2424.

Production Credits

  • Executive Producer: Eve Trow
  • Producer: Rachel Ishikawa
  • Associate Producer: Jordan Smart
  • Mixing and Mastering: Brian Pugh
  • Original Music: Brian Holiday
  • Head of Podcasts: Lital Moulad

---

This document provides a comprehensive summary of the discussions surrounding GPT-5 in the "Pioneers of AI" podcast, offering insights into its capabilities, implications for the AI industry, and reflections on the future of artificial intelligence.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:03For weeks now, we've been hearing about OpenAI's newest AI model. And finally, on Thursday, OpenAI released the much-anticipated GPT-5. It's a big deal, so big that we're dropping this bonus episode in your feeds to talk about it. So today on the podcast, we're giving you a GPT-5 update. How revolutionary is it? What's different about it? And what does it mean for the future of OpenAI? I'm Rana El-Khalyubi, and this is Pioneers of AI, A podcast taking you behind the scenes of the AI revolution.

0:46And here to help us unpack all things GPT-5 is Fortune's AI editor, Jeremy Kahn. Thank you so much for joining us again on the show. And I know you're on vacation, so we really appreciate you making time for us. Thanks for having me on again, Rana. All right. So right off the bat, what are your first impressions of GPT-5? Yeah, I mean, I think it's a very good model. I mean, there were people who were expecting, you know, this would be the AGI kind of moment. And I don't think it's quite that. It's a very good model. It definitely shows some improvement over what was available before. It's a leap forward, but not maybe a massive leap forward.

1:24So for our listeners who have not yet had a chance to play with it, is there anything different or noteworthy about the interface? Yeah. So this is the first time OpenAI has had a model where you don't have to pick whether you want to use reasoning capabilities or use something that's a faster response from where it draws simply from the pre-training data. Before, the interface required you to select which model you wanted to answer your query. Now you can just ask the question, the model itself decides how to answer that question, whether it should use reasoning, whether it shouldn't use reasoning, how long it should think, you know, quote unquote, think about the answer.

2:05And all that, you know, wasn't the case before. So part of the way GPT-5 works is a kind of router system. It's actually a family of models. And it's going to determine which of those models to send your prompt to based on the content of the prompt. This is something that some of the competing models of providers had already built into systems. So this was true with Gemini 2.5 Pro from Google, and it was also true with the Cloud 4 series from Anthropic. Those models also had this ability to decide, based on your query, how to answer that prompt. This was not true before with OpenAI's ChatGPT, and now it is.

2:46So that's the biggest difference in the interface. Do you have a sense of how different was the training process? Like did they train on just more data or different data? Like what's different in the back end? Yeah, I think we don't entirely know. They have said that it is a family of models of various sizes. Some of them are very fast. And usually if it's very fast, often those models are fairly lightweight. They're fairly small. And they've been trained on highly curated data. And it may be for certain tasks, for instance, coding or math tasks, that they've created very fine-tuned models that will handle those particular queries.

3:25And then maybe for some other types of queries, if it's a complicated logic problem where you do need a lot of reasoning, they may have used models that are quite a bit larger and take a lot more time to think. We don't know exactly how they trained it. We know it performs particularly well on coding. So they almost certainly gave it highly curated data sets on coding. We know it does slightly better at performing some English writing tasks. So they may have also created some refined data sets to evaluate and kind of post-train the model on those tasks. So there's clearly something they did in terms of training that's different.

4:02What exactly that is, they're not telling us. It's, yeah, black box for now. Is there anything different about the personality or the tone? Yeah. So one of the things that they've done is actually allow you to select four different personality types, particularly this is the case if you use the voice mode. And you can get it to answer in different ways. One of them is like I think they call it the cynic, which is going to be more skeptical. It's going to be more sarcastic in how it answers. One is called the listener. It's more empathetic. There's one called the robot, which they thought might be better for what OpenAI said is maybe it's better for business use cases where you just want to kind of a very concise answer, very factual.

4:43And so you can adjust the tone based on these sort of four different personality types. That's another difference in how the interface works with GPT-5. Yeah. You know, there's been a lot of hype about GPT-5 and a lot of anticipation the past several weeks. There are some critics that are saying that the hype is basically a marketing ploy to drive engagement. What do you think? Yeah, well, like I said, I think it is a very good model. But is it orders of magnitude different than anything we've seen from competitors? And the answer is no. So to some extent, there is a bit of hype here. And I think it's interesting because you had someone like Sam Altman saying, oh, this is the first model where it really feels like you're talking with a PhD-level researcher on any topic.

5:31And I don't know. In my experimentation with it so far, I'd say, yeah, it's good. But is it really – do I really feel like I'm talking to a PhD-level researcher on every topic? I'm not so sure. One of the analogies that Altman used in the press conference where they announced this model was to say it was a bit like when iPhones started having retina display cameras, retina level cameras instead of the old sort of pixelated cameras. Yeah, and that was a big leap forward in photography, but it wasn't as big a leap forward, I'd say, as like going from the pre-iPhone era to the iPhone era. And I think some people were expecting, given the hype around this, that this would be a similar kind of leap from going from like the pre-smartphone era to having smartphones.

6:14And it's clearly not that. This is an upgrade in features. It's a difference of degree, but not a difference of kind. Yeah. I love that. Difference in degree, but not a difference of kind. Okay, so let's unpack how it's doing better on some of these benchmarks. And I want to start with the model's coding capabilities. You know, how much better is it? And also, what does that mean for some of these platforms like Cursor and Winstaff and Lovable that are basically packaging some of these coding capabilities and have done really well as companies? Yeah. So the coding capabilities are definitely better.

6:49They are about 8 % to 10 % better on a lot of coding benchmarks than the previous O3 model from OpenAI. They did not provide benchmarking against competing models from other providers, but we're starting to see some of those evaluations come out from third parties. And it looks like it's pretty good, but maybe not quite as good as Claude 4 still at coding. So it's sort of in the range. It's definitely better than what OpenAI had available before. So they have one benchmark that's called the SWE benchmark, which stands for Software Engineer. according to OpenAI's own evaluations on that, it is better, it is the best model out there.

7:33It gets 75 % on that benchmark, which is very good. It is like about 8 % better than anything else out there on those questions. But again, in some third-party evaluations, people are saying they don't prefer the answers as much as what Claude Forrest produced. What this means for those companies that are sort of like Cursor, those companies are kind of wrappers around various models. And so for them, it's great. I think they get to move these capabilities into their system. It may require, as every new model release does for these companies, require them to do some work on the back end themselves because they use a lot of kind of meta-prompting strategies where there are prompts that you may not be aware of as a user that are taking place in the background.

8:17And every time these models are upgraded, the way those prompts or answer changes slightly. So I'm sure they have a little bit of work to do. But in general, it's good for them because it just gives them a more capable system to play around with. It should be good for the users of all of those kind of wrapper products. Okay, so let's talk about the lower hallucination rates. Is there any data to back this up? Yeah, so they, again, did evaluations of the models on hallucination rates that's in the system card for GPT-5. But again, it's compared to their own previous models. It's not compared in the system card to any competing models from different providers.

8:56Now, according to their data, it is significantly better in terms of factual accuracy. But, you know, what you should note, though, is it's still GPT-5 in kind of the main mode still has a hallucination rate where it's about 10%. And even in what they have called the thinking mode, where the model uses its reasoning abilities and takes a long time to provide responses, in general, that usually produces better, more factually accurate results. But even there, you're talking about they've gotten the hallucination rate down to, I think, 4.8 or 4.9 percent. But that still means about one out of every 20 times you're going to have an inaccuracy.

9:37So it does seem from some of the third-party evaluations that have been done and some of the third-party comparison of benchmarks, various people have published self-evaluations, that it's about on par with Gemini 2.5 Pro in terms of hallucination rate. And again, maybe slightly better than Claude 4, but it's close. So yeah, that's what we know so far. Yeah. How do you think this affects OpenAI's business against some of its competitors? I mean, it sounds to me that they're still neck to neck, you know? Yeah, I think they're very close. So over the last basically six months, it seemed like they had actually kind of lost their pole position.

10:18It was, you know, on a lot of things, Claude was better. And then Gemini 2.5 Pro was a very good model. It came out on top on a lot of benchmarks. So this allows them to claim, I think, that they're back in the lead by a little bit. So they caught back up and pulled maybe slightly ahead. But they're not ahead by miles on a lot of this. But I think what it does is if you were already a heavy user of Chep GPT and you had a pro subscription, I mean, they've basically given you a reason not to switch to somebody else. If you were considering, you know, abandoning it and go, you know, thinking, oh, I think, you know, Anthropics Claw is so much better now or Gemini is really good now.

10:55Maybe I should consider dropping my OpenAI subscription. They've basically given you a reason not to do that. And meanwhile, the other thing to note about OpenAI is they're very focused on the consumer business. And they've rolled GPT-5 out to everyone, including the people who just used the free version of the service. And I think for a lot of those people who will not have even maybe played around with the reasoning models before because they weren't available to the free users prior to this, they may be kind of blown away by how much better this is compared to what they had before. And it may bring in even more consumer users.

11:27They're now, I've said, they have about 700 million active weekly users. That's a big number. And they may even gain more users now that this model is available. So I think in terms of their consumer business, this is good. Also, in terms of their enterprise business for the people who use their API, they've priced this pretty aggressively. They have matched the pricing that Google offers for Gemini 2.5. And they have way undercut, actually, what Anthropic charges for Claude 4.1, Opus. So I think they may find that a lot of businesses, you know, are going to switch to using them or swap them in.

12:07And it may be very good for their enterprise business. Yeah. While we're at it, so OpenAI is allegedly selling shares held by current and former employees at a price that values the company at$500 billion. dollars. And one theory is that a little liquidity will stop staff basically from jumping ship to Meta or X or Anthropic. What do you think of that? Yeah, I mean, I'm sure that's exactly what they're doing. You know, I'm sure that the reason for the secondary share sale, they've been losing staff to Meta. Meta has been on this huge campaign of poaching staff from other AI companies for this new super intelligence unit that Meta has set up.

12:47And they've literally been offering$100 million dollar signing bonuses and in some cases even more than that to lure top researchers. There was even some reports that they were offering one particular researcher shares that might have been worth a billion dollars. I mean, it's incredible amounts of money. And we know that when they hired Alex Wang from Scale, they basically did this deal where they made a huge investment in Scale for$14.7 billion. So they're spending tons of money to poach talent. And OpenAI has been trying to counter those offers to keep staff. And I think they've been struggling a little bit.

13:26And this secondary share sale would definitely give a reason for a lot of their employees to stick around. Yeah, it's an incredible time to be an AI talent, for sure. Yeah, exactly. So in other news around OpenAI, they appear to be acknowledging that they kind of need to address the issue that a lot of people are using ChatGPT for mental health support. And so they've set up a mental health advisory board. They've also implemented some new guardrails to prevent users from kind of viewing and using the chatbot as a therapist. Can you talk more about that? I think this is really, this is actually very important, not just for kind of the raw open AI models, but for a lot of the companies that sit on top of it that are addressing mental health and addiction or suicide prevention, etc.

14:18Yeah, I mean, it does seem like it's one of the primary, you know, it's a huge use case for these chatbots is people using them essentially as kind of therapists and using them for in ways that the users often feel like is improving their mental health. But there has been some documented cases also of people who have, you know, had detrimental effects on their mental health from conversations with these chatbots, particularly sort of going down rabbit holes on conspiracy theories and those sorts of things. So I think they're very much a double-edged sword. OpenHead keeps saying, look, we didn't train this to be a therapist.

14:51It's not a doctor. It doesn't, you know, shouldn't necessarily be used in that way. But I think users are doing it anyway. They've done some things to try to provide disclaimers to users. But it's not – the models will – they haven't created a situation where the model will refuse to provide sort of therapy-like advice. They have tried to make that advice safer. So the model should not tell you that it's okay to self-harm or anything like that. And if you show suicidal thoughts, they now have the model sort of recommend that you seek outside help. So these are good things. You know, I think some people would like to like them to go further.

15:31But it's I guess it's a it's a fine line between sort of being an empathetic listener and acting as a kind of friendly companion and acting as a therapist. And I think they're finding it hard to kind of to straddle that divide. And I think also, they don't necessarily want to I mean, it's such a popular use case that I think actually, they don't want to discourage users too much from doing this. You know, I think actually they're quite happy. It makes the model very engaging in a lot of ways. People who are using it as a therapist are not likely to abandon using the model. So I don't know. I think they're a little bit two-faced sometimes on this stuff when it comes to this issue of, you know, should people be using chatbots as kind of mental wellness tools?

16:18Yeah. You know, what's interesting is they also tout GPT-5 as being much better on health-related questions. I certainly use chat GPT to answer health-related questions. I upload our medical, you know, our medical, like, blood tests and whatnot. What are they claiming around that? Yeah. I mean, this is another, this is sort of an example of what I was sort of alluding to is that on the one hand, they're very eager to say, oh, it's not a doctor. It's not, you know, it doesn't have medical training. It certainly is not supposed to be a therapist. And yet, you know, they know that people are using in this way and they're now touting very explicitly how good it is at answering these health questions.

16:59So I feel like, again, it's a little bit of doublespeak there on some of this. You know, they keep saying, look, there's no patient-doctor confidentiality here, and this is not necessarily, you know, their back end is not HIPAA compliant. But anyway, getting back to what OpenAI said about the health performance of these models, they did evaluations on how it does, you know, with health questions. And they're saying it does much better than any previous model out there. And they have claimed it's better than competing models at, you know, essentially making diagnoses. Which is interesting. Again, this is an area where they know users like to use chatbots in this way, find it very helpful.

17:38And they're very much leaning into this as a kind of consumer-facing product, and this is part of that. You know, there's been this image that's been circulating on social media, and it's basically a graph showing the amount of tokens processed on OpenAI, which is, of course, representative of usage. Have you seen this? And what's your take? They're basically like, it's a lot of usage. And then on June 6th, it really drops. Right. Yeah, I've seen this. So yeah, a lot of people have used this to say, aha, you know, the people who are really using ChatGPT are students, both in high school and university level.

18:14And it dropped then because that's, by then it's when sort of most schools were letting out. And kids are just using these models to cheat rampantly. And that's the primary use case of ChatGPT. I mean, to the extent that we know that the token usage graph is accurate, I mean, I do think students are a big user group. It is true that students are using these models a lot, and they're using it a lot for school work. Now, some of that's cheating, and some of that may actually be quite legitimate use cases. One of the things that OpenAI did before actually releasing GPT-5 is the week before they had rolled out this new study mode for ChatGPT where you can specifically ask it to act as a tutor, take on the role of a tutor.

19:01And it won't then give away the answers directly to you. It kind of leads you through Socratic questioning towards the answer yourself. It can literally create a whole curriculum for you, walk you through topics. It's actually quite powerful. My fortune colleague, Sharon Goldman, wrote a column about this. She said she'd been terrible at high school algebra and had always been really ashamed of it. And then with study mode, she literally had to create an algebra curriculum for her and walked her through it. And I think, you know, this is one of the great use cases of these things. And I think things like study mode are fantastic.

19:37So I want to come back to something Sam Altman said during a press briefing about GPT-5. And he basically said this is a significant step along the path towards artificial general intelligence. Why do you think he said so? Oh, well, you know, they've been hyping this model for a long time. And we know that their goal as a company is to explicitly, you know, achieve AGI and make sure that its benefits are, you know, accrued widely to all of humanity. So saying that's a significant step towards AGI, you know, I think is a way of saying, look, we're still on this mission. We're still in the leadership pole position in this race to get to AGI.

20:17And I think this whole thing about AGI, it's a very undefined term, as you know. I don't think anybody really knows what that means anymore. And we don't have a good benchmark just for AGI. So it's unclear how close are we really? Are we closer than we were yesterday? Yeah, probably a little bit. But how much is very hard to tell. Yeah. You know, Altman said, and I love this, this is not a model that continuously learns as it's deployed from new things it finds, which is something that to me feels like it should be part of AGI. I totally agree with that. Because, again, for some of our listeners who may not be familiar with how this works, they're basically training a new model, deploying it, and then it's this model that's in production.

21:01It's not updating itself on the go. And I do feel like that ought to be part of AGI. No, I totally agree. Like, yeah, you know, continuous learning, I think, would definitely need to be part of AGI. You know, something about learning efficiency, too, is not really taken into account. Humans can learn from very few examples. It's not totally clear that this is, you know, what happens with these models. So I do think there's several bits of this that are still not there. All righty. So to close us out, is GPT-5 revolutionary? And does it tell us about where AI is heading? I would say GP5 is, again, it's a good model, but it's not revolutionary.

21:40It's very much an evolutionary step in the development of these models. It is not some, you know, great discontinuous leap forward. It tells you actually how bad the evaluations are because a lot, I think OpenAI even said this, it's like, this model has good vibes. I mean, that's like what a lot of people, it's like, it feels better than the other models. It's like, there is sort of something to that, but it's also like a really unscientific way to evaluate things. But, you know, look, OpenAI still has the largest user base of any of these competing products. And putting this capability in the hands of so many people is significant.

22:15Amazing. Well, thank you, Jeremy, for joining us. This was so helpful. Yeah, thanks for having me, Rana. Pioneers of AI is a Wait What original. Our executive producer is Eve Trow. Our producer is Rachel Ishikawa. And our associate producer is Jordan Smart. Our senior talent executive is Stephanie Stern. Mixing and mastering by Brian Pugh. Original music by Brian Holiday. And our head of podcasts is Lital Moulad. You can join the conversation on LinkedIn, Instagram, TikTok, YouTube, and X. Just search for at Pioneers of AI. Thanks so much for listening.

23:03Thank you.

From the publisher

OpenAI has released the latest iteration of its flagship product, GPT-5. It arrives after much anticipation, but does it live up to the hype? Jeremy Khan, AI Editor at Fortune, returns to the show to break down what this new model can do. We’ll explore how it performs compared to past iterations, what OpenAI decided to improve and why, and what GPT-5 means for the AI giant’s future.

Learn more about Pioneers of AI: http://pioneersof.ai/

Follow Pioneers of AI on all channels: https://linktr.ee/pioneersofai

At the center of AI is people, so we want to hear from you! Share your experiences with AI — or ask us a burning question — by leaving a voicemail at 601-633-2424. Your voice could be featured in a future episode!

See Privacy Policy at https://art19.com/privacy and California Privacy Notice at https://art19.com/privacy#do-not-sell-my-info.

More from Pioneers of AI

All 125 episodes
Does OpenAI’s new model deliver on the hype? Inside GPT-5 with Jeremy KahnPioneers of AI · 23 min
Listen in VO