916: The 5 Key GPT-5 Takeaways

22 Aug 2025 · 10 min · 4 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Jon Krohn summarizes “The 5 Key GPT-5 Takeaways” from OpenAI’s GPT-5 release (Aug 7), arguing that GPT-5 progress is on a predictable improvement curve and that it meaningfully reduces hallucinations and deception while consolidating multiple model modes into one.

Guests

No guests are interviewed or named in the transcript. Mentions include METR (Model Evaluation and Threat Research) and Dr. Andre Berkov (via a viral LinkedIn post).

Key claims

GPT-5 fits the same 50% software-task success curve (doubling time ~213 days), consolidates reasoning/speed/quality into one model, hallucinations drop from ~5% (O3) to ~1% or less, and deception rates drop from up to ~50–90% (O3) to ~2–16% (GPT-5 with thinking). Benchmarks: GPT-5 is “on par” with Claude Opus 4 on SWE-bench.

Notable examples

Charts for software-task time-to-50% success; hallucination-rate comparisons (O3 vs GPT-5); deception-evaluation reductions; projected timelines (8-hour tasks ~50% in ~400 days; 90–95% in ~18–24 months).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Key Takeaways from GPT-5

0:14 to 4:14

Explore the 5 most important insights regarding the GPT-5 model's capabilities and improvements.

“My first big takeaway is that unlike the leap from GPT-3 to GPT-4, the transition from GPT-4 to GPT-5 may not feel is groundbreaking.”

Hallucinations and Safety Improvements

4:15 to 6:41

Understand the reductions in hallucination rates and safety advancements with GPT-5.

“My second big takeaway is that GPT-5 consolidates several different LLM capabilities into a single model experience.”

Deception Rates in AI

6:41 to 8:07

Review the significant decrease in deceptive behavior in GPT-5 compared to previous models.

“from kind of 5 % hallucination rates to less than 1%.”

Future Potential of LLMs

8:07 to 8:52

Discover the exciting possibilities of LLMs in current applications and future innovations.

“And my fifth and final takeaway is, as the great Dr.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00This is episode number 916 on the 5 Most Important GPT-5 Takeaways.

0:08Welcome back to the Super Data Science Podcast. I'm your host, Jon Krohn. In today's episode, I'm providing you with the 5 Most Important Takeaways from the release of OpenAI's long-anticipated GPT-5 model. My first big takeaway is that unlike the leap from GPT-3 to GPT-4, the transition from GPT-4

0:28Jon Krohn:to GPT-5 may not feel is groundbreaking. And certainly a lot of folks out there have been expressing underwhelm about the model. This underwhelm, however, is misplaced as evaluations by METR, an organization called Model Evaluation and Threat Research, as the research clearly illustrates. And if you watch the YouTube version of today's episode, I've actually got a chart showing this progress, which, and I've got a link in the show notes if you want to see the chart, if you're just listening in an audio only format. And so, yeah, so this chart shows an exponential rate on the vertical axis. And what it's showing is 50 % accuracy on software development tasks.

1:22Jon Krohn:And what it's showing is the time to complete software development tasks at a 50 % accuracy rate. So when GPT-2 was released in 2019, it could only replace a human task that would take a couple of seconds. GPT-3 that came out in 2020 was able to replace humans on or do the same kind of task as a human at a 50 % success rate on tasks that would take about 10 seconds. GPT-3.5, we were at about 30 seconds. That's a big jump. GPT-4, that was a really big jump. It was about five minutes there. So we're going from kind of 30 seconds with GPT-3.5 to a five-minute software development task being handled at a 50 % success rate by GPT-4.

2:13Jon Krohn:That was in 2023. And so, you know, that big jump from GPT-3 in 2020 to GPT-4 in 2023 from, you know, just a handful of seconds to several minutes, that feels like a huge jump. But actually, following forward through all the kind of model releases, GPT-40, O3, Grok 4, and now with GPT-5, we're following this same curve, this 50 % success rate on software development tasks. It has a doubling time, an average doubling time of 213 days. And GPT-5 fits perfectly onto this curve. In fact, it's doing a little bit better than you would anticipate it would do, than a cutting edge model in the summer of 2025 would do.

3:03And so this is really exciting because it means that we're on this trajectory. It means that in just a couple of years, we will have, you know, in about 400 days from now,

3:15Jon Krohn:actually, that's not much more than a year, we can anticipate that we'll have models that can handle an eight-hour task, a full workday at about a 50 % success rate for these kinds of well-defined problems like mathematical problems, software development problems, and so on. So it's not all kinds of problems. It's only a 50 % success rate. But following on behind this curve, maybe by 18 months or 24 months is a 90 % or 95 % success rate. So this means that we're going to be able to have more and more really complex, long tasks be handled successfully by LLMs, which means more and more opportunity to be inserting these kinds of models confidently into, you know, important processes in enterprises, personally, in other organizations.

4:02It's a really, really big deal. and GPT-5 fits perfectly on this curve. So underwhelmed maybe, but we are exactly where we should be at this time. All right. So that's my first big takeaway. My second big takeaway is that GPT-5 consolidates several different LLM capabilities into a single model experience. So prior to GPT-5's release on August 7th, you might use GPT-4.0. If you were prioritizing speed, You might pick GPT 4.5 for high quality creative writing, and you might pick O3 Pro for challenging mathematical or coding tasks. And so that was kind of weird to have to guess what the right model is for a particular kind of task.

4:44It's something that you would just kind of get used to. Now with GPT 5, it's just one model experience. And so it figures out what the right kind of reasoning approach is, you know, how kind of heavy a lift in terms of reasoning abilities prior to opening responses required. This is convenient for sure, but this isn't actually, OpenAI isn't the leader in this because Anthropic has had this capability for at least several months. I've gotten used to this with Claude Sonnet 4 and Opus 4 in the Claude user interface for some time now. All right, my third big takeaway is that since the advent of LLMs, naysayers have been complaining that because of hallucinations, generative and agentic AI applications have limited viability in serious commercial or industrial use cases.

5:36I've already found that the hallucination rates since GPT-4, but particularly since agentic approaches like OpenAI's deep research were released, that hallucination rates are negligible already for most use cases. Well, GPT-5 continues to make big, big strides here. So again, a chart that you can read about in the, well, that you can see in the GPT-5 report that OpenAI released on August 7th, or that I'm showing in video versions of this podcast. It shows that hallucination rates have plummeted relative to OpenAI's O3 model. So O3 had about 5 % hallucination rates

6:25Jon Krohn:on particular kinds of prompt situations. And this has now dropped to 1 % or less with GPT-5. As long as you have the thinking, as long as thinking gets engaged by GPT-5, which you might as well use. So that is a big drop from kind of 5 % hallucination rates to less than 1%. That is a big deal, definitely going in the right direction there, and meaning that more and more use cases are now relatively safe to be using with GPT-5, with LLMs in general, the cutting-edge ones. Speaking of safety, my fourth key takeaway, and as I reported on in detail in episode number 908, LLMs are prone to dangerous deception, especially when their objectives are threatened.

7:11Like hallucinations, this is another area where GPT-5 makes huge strides, making GPT-5 much safer to use within agentic applications than OpenAI's predecessor models. So again, I've got a pretty dramatic chart in the video version of today's episode that you can check out.

7:30Jon Krohn:But basically, the deception rates on different kinds of deception evaluations with OpenAI 03 were as high as 50%, in one case, 90%. But GPT-5 with thinking brings that down to between 2 % and 16%, depending on the specific deception evaluation. So again, huge, huge step change in, you know, reducing deception with GPT-5, just like we saw a big, big, big step change reduction in hallucination. All right. So that's pretty exciting. And my fifth and final takeaway is, as the great Dr. Andre Berkov already recently pointed out in a viral LinkedIn post, which I've linked to in the show notes, with GBT5 performing only on par with other existing proprietary models such as Claude Opus 4 on key benchmarks like Sweebench, the time to get super excited about what the next cutting edge LLM will be able to do is passed.

8:34now is the time to get super excited about what you can be building and accomplishing with llms the tools available to ai practitioners are extraordinary what process can you now automate to a high degree of accuracy what new capability can you improve society with if you're not sure as i've said many times on this show an llm conversation to ideate on what you could be doing with cutting edge ai tech is merely a browser click away all right that's it for today's episode. I'm John Krohn, and you've been listening to the Super Data Science Podcast. If you enjoyed today's episode or know someone who might consider sharing this episode with them, leave a review of the show on your favorite podcasting platform, tag me in a LinkedIn post with your thoughts, and if you aren't already, obviously subscribe to the show.

9:19The most important thing to me, however, is that you just keep on listening. Until next time, keep on rocking it out there, and I'm looking forward to enjoying another round of the Super Data Science Podcast with you very soon.

9:33Thank you.

From the publisher

GPT-5 has just been released, but with not very much fanfare. In this Five-Minute Friday, Jon Krohn asks if GPT-5 deserves the community’s underwhelmed response to its release. He outlines five features of the model and explains why people might be feeling less than enthusiastic in the broader context of LLM development. Which LLMs are leading the way, and which are still playing the game of catch-up?

Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/916⁠⁠⁠

Interested in sponsoring a SuperDataScience Podcast episode? Email natalie@superdatascience.com for sponsorship information.

More from Super Data Science: ML & AI Podcast with Jon Krohn

All 130 episodes
916: The 5 Key GPT-5 TakeawaysSuper Data Science: ML & AI Podcast with Jon Krohn · 10 min
Listen in VO