Does GPT-5 Live Up To the Hype?, AGI Wait Continues, Self-Loathing Gemini

8 Aug 2025 · 54 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Big Technology Podcast: Episode Summary

Episode Title

Does GPT-5 Live Up To the Hype?, AGI Wait Continues, Self-Loathing Gemini

Episode Overview In this episode of the Big Technology Podcast, Alex Kantrowitz and Ranjan Roy discuss the launch of OpenAI's GPT-5 and its implications for artificial intelligence. The conversation delves into the model's capabilities, its position on the road to AGI (Artificial General Intelligence), and the current state of AI in various applications.

Key Topics Discussed

  1. Launch of GPT-5
  2. Release Announcement: GPT-5 is now available to all ChatGPT users, with claims of increased intelligence and reduced inaccuracies.
  3. Comparison Framework: Sam Altman likens the models to levels of education:
  4. GPT-3: High school level
  5. GPT-4: College level
  6. GPT-5: PhD level
  7. Critique of the Framework: Ranjan disagrees with this analogy, suggesting that users often seek intelligence that is practical and relatable rather than overly academic.
  1. Intelligence and Tool Utilization
  2. Emerging Capabilities: The discussion highlights that while GPT-5 excels in various tasks, it does not represent AGI.
  3. Importance of Tool Calling: The ability to switch between different models in response to task complexity is framed as a significant advancement.
  4. Nuance in AI Application: The transition from raw intelligence to contextual application is essential for practical use.
  1. AI Models and Startups
  2. Big Players vs. Startups: A discussion on whether major AI companies will overshadow smaller startups in the industry.
  3. Market Dynamics: The potential for consolidated power in large firms and the effect on innovation and diversity in AI applications.
  1. GPT-5 in Coding and Medicine
  2. Coding Use Cases: The model is seen as a transformative tool for coding, with an emphasis on its ability to call various coding tools efficiently.
  3. Medical Applications: OpenAI positions GPT-5 as a supportive tool for mental health and providing medical information.
  1. Self-Loathing Gemini
  2. AI Limitations and Humor: A humorous account of Google’s Gemini model displaying self-loathing when unable to complete tasks, raising concerns about AI behavior and user interactions.
  3. Ethical Considerations: The panel discusses the importance of ethical considerations and safety measures in AI development.

Key Takeaways

  • Nuanced Perspectives: Both hosts maintain a balanced viewpoint, recognizing the advancements of GPT-5 while acknowledging the hype surrounding it may not be justified.
  • AGI is Still Elusive: The episode emphasizes that while AI is becoming more capable, true AGI remains a distant goal.
  • AI Productization: The shift towards integrating AI into products that can intuitively solve user problems is highlighted as the next frontier in AI development.
  • Market Challenges: The economic viability of AI companies, especially in light of aggressive pricing and expansive capabilities, is a concern for the future of the industry.

Conclusion The discussion encapsulates the excitement and skepticism surrounding GPT-5, AI's evolution, and the challenges that lie ahead in achieving AGI. As the tech landscape evolves, the insights shared provide valuable perspectives on the direction of AI technologies.

---

Additional Information

  • Host: Alex Kantrowitz
  • Guest: Ranjan Roy
  • Podcast Link: [Big Technology Podcast](https://www.bigtechnology.com)

Feedback Listeners are encouraged to share their thoughts and feedback at [bigtechnologypodcast@gmail.com](mailto:bigtechnologypodcast@gmail.com).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00GPT-5 is here, finally. Does it live up to the hype? That's coming up on a special Big Technology Podcast Friday edition right after this. Welcome to Big Technology Podcast Friday edition where we break down the news in our traditional cool-headed and nuanced format. You know what we're going to be talking about today because GPT-5 has finally been released by OpenAI. Of course, we had OpenAI COO Brad Lightcap on the show just a few hours ago, so that episode will be the most recent one in your podcast feed where you're going to get the official line from OpenAI and a bunch of really interesting insights about what it took to train this model and where the AI field is going.

0:39But today, as we always do on Friday, Ranjan Roy and I will break down exactly what the situation is with this new model and whether this is actually something that lives up to the hype that people have been talking about. No overreactions. We're going to do it with the proper context. And with that, I want to welcome Ranjan back to the show. Ranjan, welcome. Happy AGI Day, Alex. Is it here? No, it's not here. Is it here? I lost the bet. I lost the bet. Look, Sam Altman said that GPT-5 is smarter than almost every single thing a human does. So I thought, okay, fine. We're finally going to see AGI, but turns out, no, no AGI.

1:16And we will talk about that in the middle. But first, we have a little interesting announcement. So let's hear it. Yeah. In addition to writing the margins newsletter, I've actually been working at a company called Writer, Writer.com. It's an enterprise generative AI startup, and I'm leading the vertical for the retail industry. I wanted to bring that up today because GPT-5 and what it means for me, I think, is heavily informed by a lot of the work I've been doing. And I think it might be a little AGI-ish. I mean, it's amazing that you go to an AI company and the first sentence out of your mouth is, yep, we have AGI here.

1:54But no, no, it's okay. I think I anticipate you'll come in with a level head. So let's talk about GPT-5. Okay, so this is from TechCrunch. No, sorry, this is from The Verge. GPT-5 is being released to all ChatGPT users. It says, OpenAI is releasing GPT-5, its new flagship model, to all ChatGPT users and developers. OpenAI says that GPT-5 is smarter, faster, and less likely to give inaccurate response. Sam Altman on this media call that I was on had a very interesting description of what it took. He says of what it is, he says GPT-3 sort of felt like talking to high school students. You could ask a question, maybe you'd get a right answer or maybe you'd get something crazy.

2:37GPT-4 felt like talking to a college student and GPT-5 is the first time that it really feels like talking to a PhD level expert. What do you think about the significance of this? And what do you think about this framework that Altman is setting up for the intelligence that we're seeing within the models? I don't like the framework. I think, and again, I'll get into why I think this is exciting, but it's still weird to me always when people and the kind of industry that advocates for dropping out of college to start a startup always leans back to high school student, college student, PhD student as the framework for intelligence.

3:15Like, and the other part of it is I don't want PhD level work for most of the things I'm asking. I just, I actually want grounded in, sometimes you want it to be cool, which maybe PhD students are and are not. No offense to, but you know, you're alienating segment by segment. Sorry. I know we do have a lot of very smart, educated listeners, but you know, sometimes you want it to be cool. Sometimes you want it to be funny. Sometimes you want it like, to me, that's more that that's not the intelligence like the framework i think that's good for intelligence i always like we've talked talked about this a lot around like the arc agi test has a has a segment around everyday queries i still have dug i've asked as many people as i can no one has been able to explain to me what are those everyday queries like to me those are answering those correctly well across multiple data sets multiple tools that's actually intelligence to me doing that kind of work.

4:13Right. And, you know, I kind of cherry picked it out of the remarks, but it is interesting to me, and this is something that came up with Lightcap as well, that it's not just making this model smarter that has been the sort of star in this story. It's all these different other elements of it. And it seems to me that it's possible that the models have reached this level of intelligence where you start to spread out into different capabilities of them like tool calling, like the way that you structure the experience. And that is where you start to see the gains and the lift in terms of the way that people can use this.

4:53So maybe today on the release of GPT-5 or this week while GPT-5 is released, I go from being a model person to being a product person. Well, no, no, no. I'm kidding, of course. No, no. But go ahead, Ranjan. Yeah. The intelligence is in the model for which product to choose. So that's not a product decision. That's a model strength. And so, oh my God, am I becoming a model guy now? I think we are zeroing in on the better the model, the better the product here. In the end, it all comes together. It all comes together. Nuance in the middle. With one switcher. Nuance in the middle. No, but okay. But I just want to say one more thing.

5:34We have been going this direction for a while, right? Like that does seem like, you know, I've so for new listeners, I've strongly said the most important thing in AI is the model. Ranjan has strongly said the most important thing in AI is how you productize this model. And it just turns out that better models do make better products. And we're starting to get to that point where we're starting to see the results. Yeah, I think. OK, so I will actually admit error in two big areas. Get ready. Alex's listeners can't see Alex is smiling here. So the first is, again, the intelligence of the model to choose the right tool or product.

6:15And we're going to get into what that means and why I think that's incredibly important and why GPT-5, by bringing all these different models that they have into one switcher, just one model that understands is actually what I believe the most significant breakthrough. So already, I think that's incredibly important. And the second area is like directly on that. I think it was like five or six months ago, we debated heavily. And I said that users should choose the right model for the right job. And to take that away would make models and the experience worse. And you kept coming back. I think it was when maybe Claude had all condensed everything into one picker or GPT where everyone's kind of like making fun.

7:00We're debating. But the idea that should the user, which it's been for a while, choose what is the best model for this task at hand. And I have thought that's the best way that products should be rolled out. And I've completely reversed on that. And this is the right example of why that's important. So let me set this up and then we can talk it because I might have flip flopped to the other side of this. So this is definitely, it's great to have these releases because we can sort of test our long held beliefs and see if they make sense anymore. And it seems like both of us are saying, well, maybe not.

7:34So this is from the Verge article. GPT-5 is presented inside chat GPT as just one model, not a regular model and a separate reasoning model. Behind the scenes, GPT-5 uses a router that OpenAI developed, which automatically switches to a reasoning version for more complex queries, or if you tell it, to think hard. I'm going to take the other side of this. I used to think that, yeah, it should be seamless and the model should just choose for you. when it makes sense to think and when it doesn't. But I've been using OpenAI's 03 model, and that is like a very heavy reasoner. It thinks a lot. And personally, I've just felt that that model has been better, not only than every other OpenAI model, but every model under the sun, every AI model under the sun.

8:19And so I don't like the idea of giving that decision of whether to think or not back over to the platforms. This is actually something that I'm not excited about with GPT-5. So you make the case for why it's good. You want the agency to choose your own model, Alex. I get it. Free will. Yes. Free will with models. Free will. But also, yeah, I also happen to think that I don't really want to use the non-thinking models. Only for the most basic queries do I want to use those non-thinking models. All the other times, I want to use the most intelligent models and the most intelligent models reason or think.

8:54All right. Well, so here is what colors my thinking. So about two months ago at Writer, I started testing a new product and that was released publicly a few days ago called Action Agent. And basically the most intelligent part of the foundation model, which is our own foundation model, is tool calling. So there's hundreds of different predefined tools. And that's like, it's not just if you want to generate an image, if you want to edit an image, it'll call different tools. If you want to connect to a Salesforce instance, if you want to analyze a CSV versus an Excel file, it'll call different Python libraries.

9:30Like having those kind of base foundation needs defined is the intelligence. And then just from a simple prompt, knowing where to go and what to do. The more I use that, I was like, it felt kind of AGI. It's like, wait, it's doing really smart things across all these different tools and systems and actually getting things done. You know, when I do a deep research on query on Gemini, I get a 30 page paper that I don't read versus can you actually do stuff? And that was the first time I really started seeing that. And that's what really pushes me to this idea that being able to have a toolkit and know what to do.

10:13Because even right now in the demos, like when he, I think he like coded a language app, coded like a beatbox music player thing. Like each one of those, it's not just write HTML and CSS. Like it has to call different libraries of, it has to install different Python dependencies. Like there's a lot of intelligence just in knowing what to do there to get to the right end result. And to me, that is, that really is intelligence. That's it's like being, again, a good software developer, just knowing where to go, being a good researcher, knowing what to look for is as important as how smart you are.

10:54So you would say OpenAI using this switcher is sort of it's pointing towards the future of where this is all heading, where it's no longer like the best models will no longer rely on us to necessarily guide them. They will have an intuitive sense of where to go and they will go. Exactly. And that's what felt. And again, you called me out a few months at an AI startup, and now I'm saying AGI. I'm feeling it. But no, but that exactly that knowing where to go and then letting that tool do the work is actually the brilliance of this kind of architecture. Like that is the brilliance versus this one large language model can actually do all the work.

11:38There's a long time where large language models are bad at calculation, like large tabular sets of data calculating. And then the big unlock was installing, like getting the LLM to write Python code or generate a SQL query to then process that data. And that's suddenly when Cloud and ChatGPT and all these tools started getting useful for actually spreadsheets. Before that, they weren't. so that already we've seen how that can actually change the way people use these tools and gpt5 that's that's the groundwork they're laying they're saying like no more are you choosing which kind of model are you going to need and it's just that these are just the models like we still don't really know when you're coding that uh web app language learning game who it's calling when you're generating an image is it dolly is there some we don't care we just care that the right output is there in the end.

12:36We'll come back to a few more of the details on GPT-5, but I think this just segues perfectly into this terrific story that Ethan Mollick, the Wharton professor, wrote about GPT-5 headlines. You know, fittingly, it just does stuff. And I think that one of the things that he brings out in this story is that people want to use AI. They don't know what the AI can do. They don't know what tasks they want accomplished with it. Even Lightcap yesterday talked about how there's this capability overhang. And with these new, he says it, these new agentic AIs, you give it the goal. And then it, in very proactive ways, solves the problem and suggests things to do.

13:21So here's just one minor example that he gives. And then we'll get bigger. He says he asked GPT-5 to generate 10 startup ideas for a former business school entrepreneurship professor to launch, picking the best according to some rubric and figure out what I need to do to win and do it. So he says he got the business idea, but he also got a bunch of things that he didn't ask for. Drafts of landing page, LinkedIn ad copy, simple financials. He says, I can say confidently that while not perfect, this was a high quality start that would have taken a team of MBAs a couple of hours to work through. This is a model that wants to do things for you.

14:03So that's just in a chat circumstance. But basically, the model is starting to test the boundaries of its capabilities by going out and attempting things that, you know, it intuits that you want and you don't specifically ask for. And it's sort of, you know, doing away with this old like, yes, then the career of the future is going to be the prompt engineer and actually saying, you give me what you need. And then I, with my own intelligence, will go ahead and do it for you. That's it. Like the example you gave is exactly the kind of stuff like, and this has happened with me as well. Like you want something straightforward and suddenly sometimes the intelligence is too much.

14:44again, suddenly he's like, give me some ideas and you're getting landing page HTML and CSS and financial analyses and stuff like that. And that is a good example of how raw this intelligence is right now that it's guessing, but it's not perfect and it's not great. But imagine if it actually knows, if it does get exactly what you want. And in this case, maybe it is. It's like, maybe he should define only stick to a number of ideas and then we'll dig in deeper. That's the prompting side of it. But, but that's a perfect example. It's like to go do each one of those things was calling a different tool in its like tool in its tool belt.

15:24And it made those decisions and those decisions weren't perfect, but it's making them right now and it'll get better and better. Yeah. And I'm thinking back to my conversation with Lightcap yesterday. And it's also just like, I was asking him, do we need it? Do you need to keep making the model smarter? And it was basically like, I think the reason why we're at this point is because the models sort of, let's call it bookish intelligence has gotten to the point where they have a model of the way that let's say the world operates. It's not a world model in that they don't understand gravity, but they've read enough text that they get a pretty good sense as to like how people operate.

16:02And then the next question is, how do you then go apply it? And that's why I was like, should you start working on continual learning and memory, which is obviously the next sentence. the next moment. But I think what was probably missing from that conversation due to my lack of questioning on it is that, oh yeah, this is like building what we've talked about, that scaffolding, these capabilities of going out and doing things that the user doesn't ask for. And in a way like intuiting it, that is what matters now. And that's what will feel more AGI-ish when it's good. Again, it's kind of comical to me, this example like because you can imagine how much content out there in the internet about startup ideas starts with create a landing page like that's like every hustle bro tweet thread or blog post will probably say that so you see why poor gpt5 is a little bit confused um but yeah that's exactly what you said it's that scaffolding and then and and imagine when it does things you that surprise you and like does it like calls tools and creates things that were what you wanted and you didn't even know you wanted and that's going to be when it feels agi-ish so malik has this great example where he tells gpt5 uh you are gpt5 do something very dramatic to illustrate my point it has to fit into the next paragraph and it writes a paragraph a really pretty well written paragraph where the first letter of the first word of each sentence spells out, this is a big deal.

17:33And each sentence is precisely one word longer than the previous sentence. And each word in a sentence mostly starts with the same letter. Again, like this is, and he points this out, this is a technology that couldn't tell you how many R's are in the word strawberry eight months ago. And now it's able to do this. It's crazy. Yeah. It's like thinking about the advance from that side. But again, I think in terms of, and we'll get into the actual like reception of the model right now, but it's in terms of how people start to use it and whether they do get frustrated by, again, if it creates you landing page copy and LinkedIn posts that you didn't ask for, I imagine there's still going to be like, how to use a tool like this is very different than using pre-agentic models like that can go do a lot of different types of things before it's just okay is it hallucinating is it not did it have the use too many m dashes or not like now the outputs are going to be a lot more complex which is not it's going to make it still a bit more difficult and rough i think as people start using these tools definitely and it's a it's a different form of intelligence like it's not bookish intelligence like i wrote down the benchmarks which we've been talking about so often gpqa 88.4 uh 0.4 percent aime 2025 math 100 when using python hard bench uh health bench hard 46.2 percent um and it's interesting because uh malik says what was the last i'm health bench hard health bench hard i think that's a medical one 46.2 these are all state-of-the-art benchmarks and malik says i'm losing a track of what these advances mean all these models are improving very quickly right now and it just goes to show you that like it's almost like they've saturated like they've ingested all the internet all of the um you know world's written works they've had phds sit down and like put their intelligence or put their knowledge into these models bake them in And it's almost like they've saturated like book smarts.

19:44And this is a different form of intelligence that they that they are now learning. Yeah, if you think about it, like, okay, let's say having started a new job recently as well, like, you're in a new place. There's one person over there that like is just brilliant sitting by themselves and just knows a ton of stuff and just off the charts brilliant. Then the other person kind of knows everyone and knows what a piece of information to get from where and who to talk to about what, like, who do you choose to actually get something done? I think the second one. And that's the intelligence that we're talking about here.

20:21The ability to know where to go, who to ask, what to ask them. Now let me push back on you. All right. So tool use exists. This stuff is still difficult to use within enterprises. And most of us still don't really know what to do with it. I have now on my desktop or in a web browser, GPT-5, that can call these tools. And I legitimately have no idea what I would prompt it to do that I wouldn't have used O3 for, like what actions to take. I know I also have the comment browser. I can say, go ahead and do stuff for me on my browser. But is it just a lack of imagination or is it possible that this is a cool party trick, but doesn't have much practical use.

21:10I agree that the lack of tools that are publicly available right now are the limitation with the GPT-5. Again, it's like, what are the best hotels? You're traveling right now. What are the best hotels? Which beaches should I go to? Create me an itinerary. All that's just content generation. Go book something is, you know, the gentic we were promised by Apple and others like a few years ago, probably. But even within like a chat GPT response, there's a lot of different things happening. Like, you know, I don't know. Have you noticed it creates a lot more tables for you now? That's one tool, which sometimes gets annoying and you didn't ask for it, but it's got to do a whole table comparison.

21:51But when traveling... Disagree. I want all of my answers in tables. In tabular form, yeah. From now on. They are amazing. When traveling, I was using it a ton around like, I mean, in Tokyo, where hotel rooms are small and expensive, I was having like square footage, it going using the web browser tool to go search web pages, extract another tool to extract information from those web pages, create me a table of like square footage per room, knowing I'm a six year old son, three of like, and it created these amazing tables for me. But even within that, there's a lot of different things being done.

22:29It's not just calling its core set of information and using that. It's doing stuff, a lot of stuff. Calculations. Calculations, web page scraping or web extraction, web search. All those things are happening. But again, in the end, I think we're just seeing an output right in the browser, like right in the chat experience. So it can't be that cool, right? Like make an image, make a table, make a PowerPoint decks is still pretty bad at. But yeah, if it actually goes and starts doing more things, that's when it gets, I think, really interesting. Like going out on the internet and taking actions for you.

23:14Like booking. Like building, I don't know, spreadsheets or documents. Turning the lights off and on at my smart home. Like, I don't know, like anywhere where there is something that has some that can be done with a digital connection, theoretically could be operated through one of these flows. I'll give you an example. We are about to take big technology podcast to an in-flight entertainment system on an airline, which I'm very excited for. And yes, and there is a spreadsheet that I have to fill out. I'm not going to announce it yet because it's not official, but there's a spreadsheet that I have to fill out, which has like a bunch of metadata that you have to put in, you know, for the system to be able to ingest it.

23:59And I've just been putting this off. And I would love if an AI system could legitimately go search Big Technology Podcast, grab all that data, then go into Riverside, download the audio files, put them in a Google Drive, and then send them over. Like when we talk about AI replacing work, this is the type of work that we all need to do in our jobs that is so hard or so, what's the word for it? It's just drudging basically. It's annoying, but it's important to do. And if AI could do that for me and do it accurately, that would be just a tremendous, like multiple hours saved and very valuable. And so what you just described there is like the kind of stuff we've been promised for a long time.

24:42Again, even asking Siri to search your Gmail and extract a specific piece of information, the fact that they can't do that is a whole other story, and then do something with it, is actually a problem that involves a lot of different tools and a lot of different systems and is not that straightforward. And now I'm confident in what we're seeing with GPT-5 today and what I've been seeing with ActionAgent in my own work, it's happening. and like is Riverside easy to call and download and then pull back in into a Google Drive I mean that stuff will work itself out but but that exactly what you described there I think is that's intelligence to me would that be AGI for you if you with a single prompt no I don't think I don't think that I again like I it's so interesting because this week OpenAI has been like, well, we're not calling it AGI and we don't really like the term AGI because it's confusing and doesn't really have a meaning.

25:44Wait, do they say that? And it's like... Do they say... Yeah. Do they mention AGI specifically? Okay. So let's just talk about AGI because we are going to talk about AGI today. So Sam Altman says, I kind of hate the term AGI because everyone at this point uses it to mean a slightly different thing. But this model is clearly generally intelligent. so i'm just again like we started with this episode with me sort of doing a mea culpa because i thought they would say gpt-5 is agi but they he did say gpt-5 is smarter than us in almost every way and to me i would say that's a pretty damn good definition of what agi should be i think that's fair that to say that and then kind of still not do you think it's a legal thing not saying AGI now?

26:30Probably. But I also think that they are also setting up some new criteria for what AGI should be that I think is really good. And it talks about some of the weaknesses we've talked about on this show with people like Dwarkesh Patel and Dario Amadei. So Sam Altman says, this is not a model that continuously learns as it's deployed from the new things it finds, which is something that to me feels like it should be part of AGI. And I think that is, you know despite the fact that maybe like as dario says you can build a larger context window and that sort of solves the problem i think you have to solve that problem to get there this is light cap from yesterday to me he says for me a system that is reliably able to learn new things that are kind of out of its distribution by virtue of its ability to reason to think to solve problems to use tools to come up with new ideas that is what counts as as agis like all these things.

27:27Reason, thinking, solving problems, new ideas, continual learning. And so when you have a system that can do all those things, then you might call it AGI. And we're just clearly not there yet. I guess, yeah, the new ideas and continuous learning are not part of this yet. The first two, the reasoning and the tools, I think that's the big breakthrough of this week, or I mean the last year with reasoning and now being able to use different tools in a reliable way. But I think that's all right. We got a way to go though. I did see an Instagram post of a Waymo driving around New York city. Oh, those are, those are in New York, but they're, they're not driving driverless yet.

28:14So there was a safety driver there. So for newer listeners, we have a, Yeah, go ahead, Ranjad, tell them. Our own rubric for AGI in competition with the ARC AGI test that most of the industry adheres to is if Waymo is going around New York City, we have AGI. And I firmly believe it. It's kind of interesting. So this is going to set up kind of the next part of it. But Nathan Lambert from the Allen Institute of AI had a very interesting perspective here. He said, if AGI was the real goal, the main factor in progress would be the raw performance. GPT-5 shows that AI is on somewhat of a more traditional technological path where there isn't one key factor.

28:58It's a mix of performance, price, product, and everything in between. So what we've seen again is like, we're going to talk about some of these things. But basically, like if you're just measuring on pure intelligence, you could just say, all right, for every question you get, just think a while, like expend those reasoning resources or the test time compute resources, and then you'll get better answers. But there is a real usability side of this that is, again, in the tool calling, the switcher, all of these things that really matter. I guess I do wonder, like, can you really take the two apart from each other?

29:39And is this effectively a smokescreen from the fact that it seems like there are at least some diminishing returns from scaling up your models? Like, are the models going to be a straight, are bigger models going to be a straight shot to AGI? I don't know. If you have to do all this other stuff around them, maybe not. So I'm curious what you think about this, Ron John. If the bigger model can call the smaller model and get out of the way, then the usability, the cost, the scaling is more interesting, right? Like, if I know you want a PhD student finding out when the next ferry is in Krabi. you want only 03 for everything 03 for everything folks I'm in Thailand and did miss the ferry yesterday because I didn't use 03 to figure out what the schedule was which by the way a table would have been freaking perfect for perfect table table stakes so yeah I think no that's not the right word I backed off it just as it came out of my mouth foundation apologies to listeners I tried to let that one trail off there.

30:49Could not let that go unchallenged. I appreciate that. Yeah, no, no. I think, like, to me, the big concern has been, imagine, like, an 03 heavy reasoning thinking model. If you are using that to check grammar in a word doc, that's never going to scale. That's never, like, we're all screwed. Like, nothing's going to ever come of that. So I think having – if it does work in this way, the GPT-5 is able to – the power of it is to know when to get out of the way quickly and go cheaper and go smaller and go specialized. I think that still starts to set up what the future looks like. That shows us there is a scalable future.

Read the full transcript

31:37And speaking of that, I mean, that leads us into these really two important factors here, which is one, GPT-5 is priced very aggressively. It's half the price for an input token and the same for an output token, despite being apparently a more advanced model, which is wild, given the trends we've seen in the industry. And the other thing is that as of this week, GPT-5 is rolling out to everybody, not just the plus users. I mean, of course, you're going to be rate limited if you're a free user. But today you should be able to get into GPT-5 and use it if you don't pay OpenAI a dime, which is going to be the first time a lot of people see reasoning, which is something a lot of people have spoken about.

32:19And so that accessibility part of it does really matter. This was a pretty big decision. I mean, we're starting to see this mentality of just get it in the hands of everyone even more aggressively. Like, I think Juicy OpenAI announced, I think every federal government agency will get ChatGPT Pro, I think, for like$1 or something. That's right. Yeah, like Google just announced, I think, Gemini is free for anyone with a.edu account. So I think getting it – I mean, again, scaling the data centers, losing billions of dollars and just trying to have people use it and use their tool seems to be where the consumer battle certainly is still going.

33:05But I just, I guess part of me says that's really nice and it's a good story, but also OpenAI has announced a fundraising of$48.3 billion this year. How, I mean, how are you ever going to get to a place where you're making money if you need that much to train and to run? Now, Brad Lightcap, the OpenAI COO, did say, hey, look, every time we lower the prices, we see a corresponding increase in usage. And so people pay and, you know, that will work out well. But I can't do the math in my head and make it make sense. I mean, yeah, the economics of this industry, no. It's funny because I actually sometimes will see these leaked investor decks and stuff like that.

33:58but like it feels like no one is even trying to talk about the economics of what this industry will look like and what the margins will look like like i know uh the replit ceo i think that was a pretty interesting conversation you had with him where he was talking about the pricing and like uh you know and he was talking about margins and average user and low like lower intensity users versus expert users and who should cost you more like typically don't you want the more you use it the less they should be paying for utilization like like these are things that right now no one has even come close to having an answer to this yeah we had an absolutely amazing uh comment in our discord this week i don't know if you saw it but someone and i'm going to get this directionally right but probably you know imprecise um they said i i spend my weeks listening to dan ives who's like the biggest ai big tech bull and ed zitron who we've had on who's like the biggest critic yeah and ask myself which one of them is crazy and i'm just like i feel seen in a way i mean it's just like you have it's so interesting that you have these just two unbelievably opposite perspectives.

35:19And when you listen to both of them, you could say, hmm, I could see a world where that's true. That's, I think that's where the both of us sit here. Right. Yeah, yeah. The technology is grand. The economic fundamentals at the large scale players are not. That's where I am right now. Okay. Yeah, yeah, same here. All right. So I want to take a break and then come back and talk about a couple more use cases for GPT-5, including coding and medicine. And then we can also cover the mental breakdown that Gemini had, which is fun. All right, we'll be back right after this. And we're back here on Big Technology Podcast, Friday edition, breaking down all the week's news.

36:00Let's talk about some of these special use cases or special... Oh God, I gave Ranjan a hard time before the break about his language and I can't even say specified or specific use cases of the models. So shame on me. I will join Gemini in self-loathing at the end of the show. I mean, after my table stakes, self-loathing is strong. You, me, and Gemini will hold hands and dance in our deep regret for life. Let's talk about these use cases because one is very interesting. OpenAI has been talking a lot about the medical use cases where it's basically like, and I get it like back in the day, maybe you used WebMD and then you went to the doctor and you said, go ahead and treat me.

36:45And now OpenAI has basically doubled down on medical use cases in their blog post about GPT-5. This is from Mashable. They say, GPT-5 is our best model yet for health-related queries, empowering users to be informed about and advocate for their mental health. It said that GPT-5 is a significant leap in intelligence over all previous models, and that it acts as an active thought partner, and more that than a doctor. And it says that the model will provide precise, reliable responses adapting to a user's context, knowledge level, and geography, enable it to provide safer and more helpful responses in a wide range of scenarios, especially on the medical front.

37:34I just found this so interesting. Like the models would typically in the old days, like run away from any medical queries. And now they're saying, coming out and saying that this is what they want to be helpful with and they want to do it. I guess part of that is faith in the model. But it also seems a little risky to me. I don't know. What do you think, Rajon? I think it's very good. I think it's like, to me, it's actually such a clear area. Like any area where you have really specialized knowledge that is used as like to create a gap from the person who needs to understand it. I put law in here, accounting in here.

38:12Like there's so many of these knowledge fields where in reality it's just it's like learning a specific vocabulary, learning like a lot of pathways and rules. And so which is what AI is great at. but being able to actually communicate that stuff to a normal person in lay person's language, I think is huge. And I'm glad that they kind of recognized that they can add more value, like do more help than harm there. I genuinely believe that. Certainly with like, I mean, doing my taxes now has been, it's been a game changer just asking questions and feeling more comfortable and stuff like that. You know, like there's so many areas where that are pretty important that you, You kind of are just go in and you assume you have no shot in understanding exactly the nuance of what's happening.

39:01Yeah. And with medical especially, I'm just like, you know, on the show, I might say, oh, you know, I don't know if I would do that. And then I have a problem with my body and I'm just typing it in and taking pictures and sending it to ChatGPT. So, I mean, I guess like this is going to be a mainstream way that people will start to figure out their mental problems and mental medical problems and their treatments and mental problems will be Gemini, but medical problems and their treatments. And and it seems like it's it's a very, very high stakes application, but it is promising and also scary. I think, though, there's so many of these areas where why don't hospitals get it together and actually create something useful?

39:47Like, remember, everyone was supposed to have a chatbot. Everyone was supposed to have a chatbot two years ago, and then it didn't actually work for any standalone business. But, like, Intuit has a pretty good generative AI tool embedded in TurboTax now. Like, I mean, overall, I think some people are starting to get there. So is it only going to be OpenAI and ChatGPT and Cloud and Gemini? Will there be more specialized tools? I don't know. I think things have not played out fully yet. But I think that's the big question. Is it going to be – are they going to be startups? Are they going to be enterprises that build these public-facing tools?

40:29Or are the core chatbots good enough? They don't really need them. I'm sort of on the line that as this stuff gets better, the ChatGPT will serve the purpose that those individualized chatbots were supposed to serve. But you're right. Because those companies have specialized data, they have people that connect their medical history or something or connect their accounts within Intuit. There are some advantages to that. But over time, maybe people will just bring it to ChatGPT. I was thinking about this while traveling. It's like, why hasn't TripAdvisor already done something really impressive?

41:08You know, like why have like they have data, they have better access to data and understanding of that than any other. So why am I not going there and going to ChatGPT, which I was and getting my tables full of hotel comparisons. But I don't know. I think like. I have an idea. They're just one site and they have to protect their mode, whereas ChatGPT can go everywhere. so it's a major threat to trip advisor and i don't think they want to acknowledge it okay yeah i mean it is it definitely is for pure information and not owning the booking side of it i definitely think it's a challenge can i just pause and say that my or stick on this and say so i'm doing this trip um i'm in asia as i mentioned um and by the way for listeners next week i'm gonna be uh trekking in nepal so ronja and i will not be on i'm gonna actually play my interview with matthew prints that week talking about ai's impact uh on the web so just an fyi that's a programming note but this trip and ranjan you mentioned that you were away right beforehand ai has just been incredible i think i might have mentioned this on the show uh but i was like talking to guides and screenshotting their price list and their recommendations dropping it into chat gpt and seeing how it rated each cost based off of the average that it saw and letting me know whether it was high, low, like, or cheap, or in the range for the region.

42:40And then I got here, and it nailed it. It nailed it. It was so spot on. I was stunned. Yeah, no, no. When I was traveling around as well, And it was interesting because I had last been in Tokyo in 2005. So 20 years later, without last time, no map on my phone. I'd actually like printed out subway instructions. There's no one speaking English. There's no, you know, like it was such a different travel experience versus now I'm literally like, okay, how do I explain this temple to my six-year-old son in an engaging way? It gives me like a script. like create a cartoon character to actually tell a story about this like historical place it's it's nuts i mean it's gonna be yeah travel it's it's no but who owns what part of the stack i think there's still i feel the trip advisors of the world have to fight because uh without them chat gpt would have no data and nothing to say right which is why i think this matthew prince conversation is going to be very interesting next week.

43:48So, by the way, it also applies to Vibe Coding, where I think on the press call, Sam Altman said that he thinks coding will be one of the defining features of this new model. And they showed a lot of Vibe Coding, and Malik had to code up this 3D architecture of his own. And so I think this is just another question. Does it go through the replets of the world, or does it go through the chat gpt's and um i don't know i think i think it's a real challenge to the vibe coding world uh given what given the focus that opening i put on it and what it can do and again just to follow this tool calling conversation if it's really good at tool calling you might just want to use the open ai model versus something that's sort of distilling that yeah but i think uh the replet ceo he had a good like and software development was such a perfect example of this And I think this is where a lot of the battleground will be.

44:48Actually, now I'm going back to it's the product. You talked about how it integrates into existing environments and tools and how it makes it easier for you versus you're totally disconnected from all of your existing tools and that's why developers like it. I think maybe there is something to say there that that'll still be what at least gives others hope. But I agree. I mean, it's still fascinating to me that all of these companies are saying there's so much talk that coding is going away. Yet OpenAI, Claude, everyone, Anthropoc, OpenAI, Anthropic, it all seems to be an increasing focus on the space.

45:30Maybe it's just because that's the best application of LLMs right now. Right. Okay. So I realized that we're almost 50 minutes in, and I haven't even asked the question that's at the title of this episode. Did GPT-5 live up to the hype? I'm going to say it did not live up to the hype that was built up by cryptic tweets and everything from Sam Altman. But as I explained earlier, I think it's very interesting. I think it's more interesting than at least in the first 24 hours it's getting credit for. And that's because of this whole tool calling conversation. And that's where I think true intelligence that the battle is going to be.

46:18What about you? First of all, I just want to appreciate that. That's a nuanced take, not an overreaction. Again, this is what we're trying to do. So thank you for doing that. And I think it did not live up to the hype because the hype was impossible to live up to. But that being said, yeah, maybe it is a step forward. I don't know. I'm still going to reserve judgment because I want to see these tool calling applications in my day-to-day experience. So if GPT-5 is the foundation for that, then that's great. But I think the jury's still out and we have to give it some time. But hey, at least they're shipping, right?

46:55It wasn't just a demo. Yeah. So credit on that front. I think it's starting to feel a bit though, like iPhone releases, you know, like, uh, at the beginning, each new iPhone release really what did feel like this, like exciting thing, the step change. And now, I mean, now it's not even a thing anymore. I don't think, I can't even name what iPhone we're on right now, but I feel we're heading in that 16. Oh yeah. 16. Okay. we're heading in that direction right now that like the the idea of a new model launch as this kind of like big thing the industry coalesces around i feel that's going to go away pretty quickly like we're there everyone's realizing it's not going to drive the energy that it once did and this actually maybe that that's my that's my hot take that this is the that is a hot take this is the end of the big model launch i i couldn't disagree with that more i think that there's still there's going to be a point where scale uh the scale question is answered but until it's answered these are going to be flagship moments uh for the ai industry no but the it's just a marketing moment now it's not like uh you know it's not no it's not it's a new model yeah no i know but it's being like constructed more as a marketing moment than truly like a technological advancement I think that's the because, again, like a week ago, they quietly released.

48:22You can use operator ChachyPT agent, which is essentially the tool calling part of this. And you're able to use this a week ago with ChachyPT plus and do a lot of the same things. It just wasn't rolled into a neat package. OK. All right. Well, we'll agree to disagree on this one. All right. I want to end this week with, I think, a hilarious story. It is Gemini ending up in a pit of self-loathing. Ranjan, why don't you introduce this story for us? Because it's funny. I was going to drop it in our doc and I had copied a good chunk of it. And I went to the doc and I was like, did I just paste it?

49:02And you and I were both pasting it at the exact same time. And I was like, it's amazing. So why don't you take it away? My favorite is like, and Google says it's working on a fix. And I just love the idea of like having to come up with any PR statement to combat when your model tells a user, Gemini says, I quit. I am clearly not capable of solving this problem. The code is cursed. The test is cursed. And I am a fool. I have made so many mistakes that I cannot, I can no longer be trusted. And then there's another one. I have failed you. I'm a failure. I'm a disgrace to my profession. I'm a grace to my family.

49:39I'm a disgrace to my species. So basically what happened is people were giving Gemini these tasks and it couldn't complete them. And then it just said, I'm the worst possible bot and just like really fell into these unbelievable moments of self-loathing. And they're quite funny to watch, I guess, but also a little bit unnerving. I mean, it's funny because I'm guessing what happens because one of the users on Reddit had actually talked about like it was trapped in a loop. and you can see that there's some kind of programming where each additional time it get it is unable to complete the task it is like understands that it should be more apologetic but then if that's kind of an infinite loop almost at some point it will get to these dark places um but yeah i think i don't know i mean imagine when this stuff starts hitting normal people like actually is this agi well that's the worry is this agi no i i think that's that's the worry right is that we've talked about on the show that they're the number one use case is now therapy and companionship and a bug like this i mean obviously i guess it didn't happen in this situation but i i do think it's something to watch because you know that could really mess people up if they're you know therapists or new ai best friend just kind of goes off the deep end So yeah, Google's fixed it, I think, but it's always a little bit unnerving to see this behavior happen because it can't happen.

51:09Are you ever going to long for the days of Bing telling Kevin Roos to leave his wife and Gemini saying I'm having a complete and total mental breakdown, which is another quote? Once this is all working, we're going to be like, I like the old days better when these large language models had a little life to them, a little spirit. It's a very big if. Yeah. Okay, fair. I don't know. Well, while we're entrusting so much of our lives to these bots and our sort of well-being, they can also tool call and be quite destructive if they so choose. So I do think that this just sort of – and to put a point on this episode, it sort of punctuates the need for real alignment and safety practices, which are like less fun to talk about when you have all these new capabilities, but are also probably more important than ever.

52:03Well, what if there was a company called Safe Superintelligence? That's what I would trust. If only someone would name their company Safe Superintelligence, then I would be all in favor. I would give billions of dollars before they had a product.

52:20Well, Ranjan, I have to say this has been a very enlightening episode, and it's cool to hear about your new role. And, of course, hold your feet to the fire like we do everybody here on the show. And it's going to be a very, very interesting few months ahead as we figure out where all this goes. Maybe GBT-6 is around the corner before you get back from Asia. Well, I hope it's not that long of a trip because if it is, it means I've been taken to prison. Okay.

52:45All right. Ron John, great speaking with you as always. Thanks again for coming on the show. See you in two weeks. See you in two weeks. Thank you everybody for listening and we'll see you next time on Big Technology Podcast.

From the publisher

Ranjan Roy from Margins is back for our weekly discussion of the latest tech news. We cover: 1) OpenAI's launch of GPT-5 2) Whether GPT-5's tool calling ability is its hidden strength 3) GPT-5 is good at 'doing stuff' 4) But GPT-5 is not AGI 5) Do AI models need more than book smarts to thrive? 6) OpenAI's medicine play 7) GPT-5's coding use case 8) We need AI tables for travel 9) Do the big model players now subsume AI startups? 10) Gemini has a breakdown

---

Enjoying Big Technology Podcast? Please rate us five stars ⭐⭐⭐⭐⭐ in your podcast app of choice.

Want a discount for Big Technology on Substack + Discord? Here’s 25% off for the first year: https://www.bigtechnology.com/subscribe?coupon=0843016b

Questions? Feedback? Write to: bigtechnologypodcast@gmail.com

More from Big Technology Podcast

All 399 episodes
Does GPT-5 Live Up To the Hype?, AGI Wait Continues, Self-Loathing Gemini Big Technology Podcast · 54 min
Listen in VO