OpenAI’s New Model, Jensen’s Bold Claim, Alexa+ Is Here

28 Feb 2025 · 59 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Big Technology Podcast Episode Summary

Episode Title

OpenAI’s New Model, Jensen’s Bold Claim, Alexa+ Is Here

Podcast Host: Alex Kantrowitz Guest: Ranjan Roy from Margins

---

Episode Overview In this episode, host Alex Kantrowitz and guest Ranjan Roy discuss the latest developments in generative AI, focusing on OpenAI's release of GPT-4.5, Anthropic's Claude Sonnet 3.7, and Amazon's new Alexa+. The episode also touches on NVIDIA's earnings and industry reactions to these advancements.

Key Topics Discussed

  1. OpenAI's GPT-4.5 Release
  2. Overview: OpenAI released its latest large language model, GPT-4.5, available first to ChatGPT Pro users.
  3. Key Features:
  4. It improves computational efficiency by over 10 times compared to GPT-4.
  5. Despite being the largest model to date, it does not exceed the performance of previous reasoning models (e.g., O1, O3 Mini).
  6. Discussion Points:
  7. The excitement surrounding GPT-4.5 feels muted compared to earlier releases.
  8. Questions arise about whether model improvements translate into tangible user benefits.
  1. Emotional Intelligence and Reasoning Enhancements
  2. Discussion: The episode explores how GPT-4.5 is designed to better handle creative tasks and offer more emotionally intelligent interactions.
  3. User Experience:
  4. Users reported that interactions with GPT-4.5 feel more natural and human-like.
  5. The model’s ability to provide concise, conversational answers was highlighted as a significant improvement.
  1. Comparison with Anthropic's Claude Sonnet 3.7
  2. Release Details: Claude Sonnet 3.7 focuses on a hybrid AI reasoning model that allows for both quick and thoughtful responses.
  3. User Experience: Ranjan shares his experience using Claude as a diet coach, noting its effectiveness compared to other models.
  1. Jensen Huang's Claims on AI Computation
  2. Earnings Report: NVIDIA reports significant revenue growth and discusses predictions of higher computational needs for next-generation AI models.
  3. Key Statement: Huang asserts that future AI models will require 100 times more computation than previous models due to new reasoning approaches.
  4. Discussion: Both hosts expressed skepticism about this claim, given advancements in cost efficiency demonstrated by other companies like Anthropic.
  1. Amazon's Alexa+ Introduction
  2. Overview: Amazon introduced Alexa+, which promises improved conversational capabilities and integration with various services.
  3. Comparison with Competitors:
  4. Amazon's strategy is viewed positively, especially in contrast to Apple's Siri and its limitations.
  5. User Engagement: Features like interacting with Prime services position Alexa+ as a compelling offering in the AI assistant market.
  1. The Future of Skype
  2. Closure Announcement: Microsoft is retiring Skype, which was once a pioneer in internet calling.
  3. Impact: The hosts reflect on Skype’s legacy and its decline in favor of Microsoft Teams, marking a nostalgic farewell to the platform.

Key Takeaways

  • Model vs. Product Debate: The ongoing conversation about whether the focus should be on developing better models (like GPT-4.5) or improving the products that utilize these models remains central to discussions in the AI community.
  • Industry Competition: There's increasing competition among AI companies, with advancements from Anthropic and Amazon providing significant challenges to OpenAI's status.
  • User Experience is Key: Enhancing user interactions through emotional intelligence and conversation fluidity is a priority for AI models as they evolve.

Conclusion The episode emphasizes the rapid pace of development within the AI sector, highlighting both excitement and skepticism about the new technologies being introduced. The hosts encourage listeners to consider the practical implications and user experiences associated with these advancements.

---

Enjoying Big Technology Podcast? Please rate us five stars ⭐⭐⭐⭐⭐ in your podcast app of choice. For weekly updates on the show, sign up for the pod newsletter on [LinkedIn](https://www.linkedin.com/newsletters/6901970121829801984/).

Feedback? Write to

[bigtechnologypodcast@gmail.com](mailto:bigtechnologypodcast@gmail.com)

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Let's break down what the release of GPT4.5 means for OpenAI and the future of generative AI. Plus, Anthropic also has a new model. NVIDIA CEO Jensen Wang makes a bold claim. And Amazon introduces a better version of Alexa. That's coming up right after this.

0:20You're used to hearing my voice on the world bringing you interviews from around the globe. And you hear me reporting environment and climate news. I'm Carolyn Buehler. And I'm Marco Werman. We're now with you hosting The World Together. More global journalism with a fresh new sound. Listen to the world on your local public radio station and wherever you find your podcasts.

0:47Welcome to Big Technology Podcast Friday edition, where we break down the news in our traditional cool-headed and nuanced format. We have so much to talk you through this week. It feels like this week, among many crazy weeks, has been one of the craziest. We have a new model from OpenAI, a new model from Anthropic, a new Alexa, NVIDIA earnings, and Skype is dead. So it was a very, very promising week for a lot of companies, but not for Skype, which will forever live in our memory. So we will say goodbye to Skype at the end of the show. But in the meantime, joining us as always on Friday is Ronjan Roy of Margins.

1:19Ronjan, great to see you. Happy New Model Week, Alex. How have all these new models changed your life as of today, February 28th? Not at all, but we will talk about whether that will matter in the long term, because, of course, I put your is it the model or is it the product question to OpenAI head of research, Mark Chen. We talked about it and now you're going to get a chance to respond. But first, let's just break down the news, because yesterday we had the release, of course, of GPT 4.5 OpenAI for the first time ever put a spokesperson on this podcast. We broke the news here with Mark. And now we're going to analyze what it means because we sort of left the fog of war and we have some perspective on whether this is disappointing for OpenAI, whether this is promising for OpenAI, and whether this means that generative AI can continue to progress or not.

2:13Now that we've seen some more reactions outside of Mark Chen saying, yes, scaling is still alive. So this is from The Verge. OpenAI announces GPT 4.5. GPT 4.5 is the largest and newest large language model from OpenAI. It's going to be available as a research preview for ChatGPT Pro users to start. And here's like a weird thing, though, that happened. There was some documentation. We're going to get right to it right away. There was some documentation that OpenAI released about this model and then removed. And it's very mysterious. They said GPT 4.5 is not a frontier model, but it is OpenAI's largest LLM.

2:52improving on GPT-4's computational efficiency by more than 10x. It does not introduce seven net new frontier capabilities compared to previous reasoning releases, and its performance is below that of O1, O3 Mini, and deep research and most preparedness evaluations. OpenAI has since removed this mention from an updated version of the document. So they did remove it. I don't think they disputed it, though. and I found what was interesting was, yes, this was a step change improvement over GPT-4. It was not over the reasoning models. So you would think maybe you could build reasoning on top of this.

3:31We're going to talk about that in a moment and it will be even better. But for the meantime, OpenAI has a new model that does not exceed the reasoning models in certain benchmarks and seem to admit that in a document. So Ranjan, you've been following along this whole way. what do you think the implications of this are? This week had me thinking. I feel with iPhone releases in recent times, a lot of us have been saying, do we really need a big event to release every new iPhone? Certainly the 16e was not exactly the iPhone launches of yesteryear. I'm starting to feel like that with all of these large language model releases, Cloud 3.7, GPT 4.5, even as you're listing out all of the kind of release notes around this, and then there's some ingredients that are not listed, or these things are removed from the actual release documentation, it's not that exciting.

4:30It's not exciting enough to have to try to launch a live stream and get everyone hyped up around it. GPT 4.5, and we're going to get into, there's some elements of emotional intelligence or emotional quotients that are around it. Perhaps creative writing is a little bit better. Perhaps there is a bit of computational efficiency introduced to it. Even Sonnet 3.7, and I was trying this, Claude Code is a pretty big release and a pretty big step change, but it's not revolutionary. So I think a lot of these companies have gotten caught in this hamster wheel of needing to do these big model launches.

5:08And there was a time where the step change was so big that it was actually exciting for all of us. But now I think 4.5 is probably the least interesting model release from OpenAI to date. Because even 01 and adding reasoning models to the overall suite was a pretty big deal. 4.5, I still cannot tell you what the big deal is. Maybe you can tell me. We are going to get some commentary from Andre Carpathy about this that he put on Twitter yesterday, which does answer that point. But even from OpenAI itself, there was some very interesting communication, shall we say, around this model. So Sam Altman came out with this tweet and he said, GPT 4.5 is ready.

5:49The good news, it's the first model that feels like talking to a thoughtful person to me. I've had several moments where I sat back in my chair and have been astonished at getting actually good advice from AI. The bad news is it's giant, it's expensive. We really wanted to launch it to pro and plus users at the same time, but we've been growing a lot and are out of GPUs and we'll add tens of thousands of GPUs next week and roll it out to the plus tier then. And there's hundreds of thousands coming soon. So I'm pretty sure you'll be able to use it once we can rack up. This isn't how we want to operate, but it's hard to perfectly predict growth surges that lead to GPU shortages.

6:28Remember, ChatGPT has gone from 100 million to 300 million users in a very short amount of time. Heads up, this isn't a reasoning model and won't crush benchmarks. It's a different kind of intelligence and there's magic to it I haven't felt before. Really excited to have people try it. Look, it's very interesting because again, we're going to go into my interview with Mark Chen very, very quickly. But Mark, I was like, you know, hey, listen, does this show that, you know, we're getting diminishing returns from scaling? And he said, absolutely not. But then you have these endorsements from Altman and it's fairly muted.

7:01So make sense of that for me, Ranjah. The world's greatest product marketer cannot market his own product. I mean, it's not as great at some things, but trust me, there's this magic, which I felt, but I'm not going to actually tell you what that magic is. I think it kind of, it actually, his statement really captures the overall feeling I have of 4.5. It just tries to put a positive slant on it. But I think that's exactly it. They have to keep pushing new models, new narratives, pushing towards GPT-5 whenever, if and when that will come. But to me, they need – and actually, this is going to get back to our product versus model debate.

7:47They need to show more product. Again, operator, deep research, those were exciting moments. 4.5 as any kind of announcement is not incredibly interesting to me. Like you had asked Mark Chen during your interview, you know, what are the new use cases or what are the use cases where this will be better at? And I was actually sitting there waiting with bated breath, ready to hear, okay, this is how this is going to help me or other people. And there was a somewhat generalized answer around how with creative writing tasks, whatever that might mean, this is better. And that was kind of it that I got out of the interview.

8:28So I think, and which lines up with the whole idea around emotional quotient, emotional intelligence, more creative writing, more thoughtful answers. And I've seen a lot of examples of out there of 4.5 answering and being a bit funny and people saying, this is the first time AI has made me laugh. But if you're just trying to get a little bit more grokky with your model, I don't know. That doesn't seem like that's going to fill in that SoftBank valuation for me. Yeah. Look, I didn't find it grokky at all. Because as we've talked about on the show - Not grokky, but more on the trying to make it funny or interesting as opposed to just giving you information.

9:11So having experimented with it, you know, as I was going to say, both of us have paid for that$200 a month upgrade because we wanted to try deep research. And I guess mine is still live. I think yours just dinged. But I'll say that I spent a good amount of time chatting with GPT 4.5 yesterday. And what they're saying is real. Like it is definitely much more pleasant to talk about it. And I spoke with Mark about this a little bit yesterday. The responses are shorter. They're more human-like. Like it doesn't feel the need to like print out, you know, a master's thesis for each answer. Like you can actually have a back and forth with it.

9:47And it was actually one of the more enjoyable conversations I've had with a bot to date. That's okay. I'll give you, that is a good point. If the big functional change is we've all gotten very used to this idea that you, you know, query a chat bot and you get this really overly, thoughtful answer that tries to both hedge itself from any kind of safety consideration and lists out 10 bullet points and, as you said, a master's thesis. So maybe there is something very important there where it actually starts to be able to answer you correctly in a concise way, in a more conversational human way. Maybe there's something there.

10:32But to me, again, why not just put that out there in the model? Why have a big event around it? Why make a big press push around it. Why not just put that in the product? Well, here's why I would say it's important to do that is because, and this is what Mark was saying, that you have linear progression of the model's capabilities. Based off of what you predict, if you put this much compute in it, you get this much output. And I think OpenAI is saying that this 4.5 is the next step on that progression. And it's met with the amount of compute that they've put in, the benchmarks that they've expected to hit.

11:10And that's why I said to him, did you find the scaling wall? And he said, GPT 4.5 is really proof that we can continue the scaling paradigm. So basically, I think that is sort of like, that is the march. But I also think it's important to kind of talk about like what it's going to feel like to all of us. And then this gets to Carpathie's comments. And it's basically here, he describes really well, the progress from the original models, because it's going to feel less as you get better. So he says, GPT-1 barely generates coherent text. GPT-2 was a confused toy. 2.5 was skipped straight into GPT-3, which is even more interesting.

11:52And GPT-3.5 crossed the threshold where it was enough to actually ship a product and sparked OpenAI's ChatGPT moment. He says, I went into testing GPT-4.5, which he's had access to. And he says everything is a little bit better and it's awesome, but not exactly in ways that are trivial to point to. Still, it is incredibly, incredibly interesting and exciting as another qualitative measure of a certain slope of capability that comes from comes for free just by pre-training a bigger model. He says we actually expect to see improvement in tasks that are not reasoning heavy. And I would say there are tasks that are more EQ as opposed as opposed to IQ related and bottlenecked by world knowledge, creativity, analogy making, general understanding, and humor.

12:39So these are the tasks that he was most interested in during his vibe checks. Basically saying that like you use this model, it's a little bit better, and that matters a lot because we've already come so far from the barely coherent part to where it is today. I think I'm going to nominate you as the new product spokesperson for OpenAI because I think you just convinced me right here. I think you just turned my entire view of 4.5 in this moment. So basically, I've been talking a lot about AI has a branding problem. The idea that people say that's written, quote unquote, written by AI. Everyone has this really narrow view of what AI text generation is.

13:20And that's because of this very dry, weird, almost inhuman way that it responds to you. And every model, whether it's Gemini or Claude or ChatGPT, everyone has this view of this is what an AI response looks like. So actually, if the real advancement here is it can move beyond that and make things more human and conversational, that actually could be very interesting overall in terms of getting people to use these products. So I think if that's the real change here, I'm surprised that they didn't hone in on that, that this is going to be what takes ChatGPT to the next 700 million people outside of all early adopters and makes people comfortable and happy with it and makes AI much more natural within all types of mediums and channels and outputs.

14:15If they positioned like that, and if that's what's really happening here, that is kind of exciting for me. I think that is how they're positioning it. They are talking about the fact that this has great AEQ, and that is where they want to seem to focus people with this release. And you look at some of these benchmarks, and so I'll just read a few of them. Simple QA accuracy. GPT 4.5 has 62.5 % compared to, let's see, 47%, the closest model, which is OpenAI 01. The hallucination rate is 37.1%. Again, lower is better. GPT-40 has a 61.8 % hallucination rate, which seems high. So those are like the standard benchmarks.

15:04But then you get into the everyday queries. And they say that for everyday queries, people prefer 4.5, 57 % of the time over GPT-40. For professional queries, they preferred 63.2 % of the time over 40. And for creative intelligence, 56.8 % over 40. So that's not nothing. No, I think if Kantrowitz and Roy were behind this marketing campaign and launch, we could have just come up with a simple make AI less AI. What about that one? Something just pushing the idea that that's what this is really about. Not getting caught up in the scaling law side of it, the compute efficiency side of it, and really saying this is the first model that makes AI less AI.

15:54It makes more people feel comfortable using this on an everyday basis. I think that I would have been, it would have been a little more exciting for me. Definitely. And so there's a very interesting debate that's going on about like, where did it get this more EQ oriented positioning? Was it pre-training? Like, is it because of its abilities or was it post-training where like they just added this personality after the model was built? We don't fully know. And actually, if I was going to have one question that I'd want to ask Mark Chen, if I could get him on the phone for like another five minutes, it would be that question.

16:30And I feel bad having left that out yesterday, but I have seen some very interesting debates about it over the past couple of days, where there's this one, Princeton academic, Arvind Narayan. He says, apparently the main thing we're getting with GPT 4.5 is an exchange for 30 times percent price, in exchange for a 30X price increase is fuzzy stuff like IQ. The ironic thing is this is an aspect of behavior, not a capability. My bet is that any difference in EQ are due to post training, not the parameter count. Okay, so that's an interesting thesis. Ethan Malik from Wharton slides into his mentions.

17:08And Ethan Malik, of course, he's a professor, he's been on the show. He's been pretty good at sort of following the pulse of AI. He's pretty positive. So he tends to take the sunny side of things. But he says disagree on this one. Stuff like theory of mind or EQ are deeply rooted in abilities, not behavior in humans. And I would bet the same for AI. But again. We don't know yet. So basically if this did come out of like just training the model, making it more able, and then it all of a sudden produces like a more human style of communication, I think that's pretty interesting. Well, yeah, I do think, and I was thinking about this as well after reading these, on one hand, it could be essentially kind of a party trick.

17:51It could be more instruction level after the actual core training where it's just speak in this voice, give concise answers, try to lean your behavior towards a certain way. I think that would actually be very sad because that would be easy. What Ethan's saying I think is the more interesting part. And I have to say, if it's OpenAI doing this, I have to imagine for this kind of product and model, they're not going to be going the party trick route. And genuinely changing the way the model thinks and produces knowledge would be a very big deal, as Ethan's saying. But again, we don't know what that means or what it looks like.

18:33Is it in the supervised fine-tuning layer? Is it in the base training layer? We don't know. I'm actually surprised. Yeah, we got to find out. You got to ask Mark Chen again, because to me, again, that is the really interesting stuff they should loudly be talking about rather than Sam Altman just saying it's kind of magic and not giving us any more. Exactly. And so there's been this other thing that's happened though, which is that people have taken a look at the evaluation scores and have noted that this is not as good as reasoning models in a lot of different fields. So I think we should talk about that because it has been used as a discussion point about whether OpenAI has lost the magic.

19:17So let me just go through some of these, you know, whatever they're going to mean to you. I'm just going to read them out. So there's GPQA, which is science, 4.5 gets a 71.4 % compared to 79.7 % for OpenAI 03 mini. So it's down by eight, seven, eight percentage points there. There's AIME24, which is math, GPT 4.5, 36.7%, compared to 03 mini, 87.3%, less than half his performance. It's amazing. It just beats on this multilingual test. And it is a little bit, no, it is a little bit better on one coding benchmark, and then a little bit worse on another coding benchmark. But basically, people have taken this, And I think this was also something I saw afterwards.

20:08I was like, oh, dear. You know, like there is a reason. These reasoning models are outperforming this on a lot of benchmarks. And I think we should say that the reasoning models use the intelligence of these, you know, standard models. And they learn how to attack things step by step, which is like, yes, the reasoning models are doing the things that they're supposed to do. And it just shows you how impressive the reasoning is. But then there's also just like, why is it lagging? People have been like, all right, that's really disappointed. Here is, let's hear from trustee Bob McGrew, former chief research officer at OpenAI, that always seems to hop in the discussion at an opportune time.

20:42He says, don't be disappointed that GPT 4.5 isn't smarter than O1. Scaling up pre-training models, pre-training improves responses across the board. Scaling up reasoning improves responses a lot if they benefit from thinking time and not much otherwise. Wait to see how the improvements stack together. I think this is really important, right? is that this 4.5 is going to be the basis of the next reasoning model that OpenAI is going to put out. And I think Mark hinted on this, is that GPT-5 will bring both of those capabilities together, where you're going to have the smarter basic foundational model, which is going to be GPT-5 or something built off of GPT-4.5, and then you're going to add the reasoning in, and then it should even further outperform the stuff we're seeing with like 01 and 03.

21:29So what do you think about that? Yeah, trusty Bob McGrew making sense of it, I think. That makes sense that building this more apt, able, emotionally intelligent foundation model and then incorporating that with the Riesing model and ideally that getting us to GPT-5 seems like something ambitious enough to actually push forward on. I guess I still have such a difficult time again, though. When we're looking at GPQA benchmark, AIME24, even when we're looking at what you had shown earlier, there's like on everyday queries, the GPT 4.5 beat 4.0 by 60%. What does that actually mean? What does that look like?

22:17What kind of real life problems? Because I'm so fascinated by what is an everyday query in one of these tests that if you have an AI researcher creating a benchmark, what is their everyday query versus your or mine everyday query? Like, I think that that's the part that's still worries me about OpenAI, that so much focuses on that research house part of it and the much like the very, very research oriented approach to all of this going back to product versus model. But it feels like we're still locked in that rat race here. OK, well, that just takes us to our model versus product question again, because I did bring this up to Mark and I said, all right, you're the head of research at OpenAI.

23:00You're a model guy. So just like, I am trying to figure out how to argue this to Ranjan. Maybe you can help me figure it out. And he did. He gave an explanation, basically saying that as the models get smarter, these products, like for instance, Deep Research, get smarter. We talked last week about how if they're hallucinating, they become useless. So the less hallucinations you would imagine, the better. Maybe unless you're Benedict Evans, who wants zero hallucinations. So I'm kind of curious to put that to you and get your thoughts on what it means. I listened to it and it still felt like a relatively generalized statement for something that shouldn't be a generalized statement.

23:42Even going back to what are the real use cases? Is it creative writing? Is that really what you're pitching me with 4.5, that it's going to be better? Is it everyday users will have a better experience with a chatbot and feel more comfortable? Is it AI therapists are going to get a lot better because now it can actually talk to you in a more emotionally connective way? To me, that's the part that the hallucination rate side of it, I think obviously matters. But if the idea is it's like to me, the 99 % versus 98 % versus 97 % for most AI use cases in the world, I think will probably be okay. to me again it's more it still doesn't answer that question like deep research can get better and better but does that mean a financial analyst will actually trust everything that they is put into a deep research report in a week in a month in a year what does that actually look like yeah i think we still don't know i mean we can definitely say for sure that like improving the model from GPT one to where we are today has mattered.

24:55But I think that the question is, yes, what are these incremental improvements going to really lead to? And like, yeah, I mean, Mark was like, it's all about getting to the frontier of knowledge in AI. The smarter these things are, the more they can do, just like a smarter human can do more. I think it's great. I love that they are pushing, you know, the cutting edge on this and that every AI lab is trying to get push, push the cutting edge on it. But I, you know, I cannot staunchly sit in my position for much longer unless I see some tangible outcomes from this. But anyway, I'll still be on team model for the time being.

25:36All right. It makes for better Fridays, knowing we still have product versus model. Yeah, I'm not, I don't really see myself going away from that position anytime soon. I want to see the better models. Thanks for shipping the better models. I'm waiting for GPT-5 to show up and magically solve every use case perfectly. And I will eat my hat, whatever one does on that day. Well, the last thing I'll say about Mark is I did say like, so aren't you setting expectations too high? And he said, I don't think so. Okay. All right, Mark. GPT-5, baby. I don't know what trusty Bob McGrew would say about that, but let's see.

26:18All right. So now that I've become the sort of de facto product spokesperson for OpenAI, let's go to Gary Marcus. Because I feel like we should talk about, we should at least give some time to those who've said actually that GPT 4.5 launch shows that OpenAI is toast and sort of discuss their points. And one of those people are Gary Marcus, longtime critic, former, well, He's been on the show as well. I'm sure we'll have him back soon. He messaged me on LinkedIn after he saw my Mark Chen interview and said, allow me to give a rebuttal. I said, all right, send me something. We'll read it on the show.

26:54I haven't got anything back, but I will read a LinkedIn post from him and we can discuss it. So he says, OpenAI is in serious trouble. They still have the brand name, a lot of data, and tons of mostly unpaid users. But GPT 4.5 is usually expensive. Even so, it offers no decisive advantage over competitors and zero moat. Scaling hasn't gotten them to AGI. The GPT 5 project was a failure. There is already starting to be an is-that-all-they-have reaction, including from some people who've said they have to adjust out their prediction of when we hit AGI. He said DeepSeek led to a price war that cuts potential profits.

27:29There is still no killer app. OpenAI is still losing money on every prompt. A bunch of investment turns to debt if they can't make the transition to the nonprofit fast enough. And Elon has perhaps upped the cost. Many, many top people have left. Some have started serious competitors with similar IP because OpenAI's burn rate is so high. They have limited runway. Microsoft no longer fully has their back. altman's credibility has diminished sora went nowhere uh whatever uh lead they had two years ago has been squandered and if masa changes his mind they will have a serious uh cash problem and elon is right that they don't have all the money for stargate man uh what do you what do you think in responding to that list what takes ed zitron about 5 000 words to write i think gary Marcus did in about, in one tweet, actually that reminded me of Sora, that it exists, which I played.

Read the full transcript

28:24Have you used it recently or? I have not. The text to video or image to video model, that one definitely went nowhere, could have been a good product demonstration. I think overall they are making this bad. It's what we keep talking about, but it has to be, GPT-5 has to be, oh my God, this solves everything. Like this is where there's no hallucinations. It's reasoning. It's a huge foundation model. It's relatively low cost somehow. I think it really, the way they're positioning their entire business is that it's going to be the kind of silver bullet to everything. Otherwise, I really don't see again at the zero moat part of it you're seeing more and more and that which is to me maybe that is why they push so hard on these constant model releases because they have to stay relevant because the moment they're just an API in the background then you're the most commoditized thing imaginable and then then that will kill you anyway so that's a pretty compelling case right there.

29:31Yeah. And I think that one point I think that I should make here, and I did speak to one more point about the market interview. I spoke to him about starting and stopping. And he said, that's a normal part of training any model. But if you're starting and stopping on a model that's this expensive to make, then your costs go way up. So I think Gary is right that the errors or whatever, the changes, the tweaks that you have to make become very expensive tweaks when you're starting to work on projects this size. Yeah, the cost of the model training, and we're going to get into how Anthropic, supposedly the new Claude, was much less expensive.

30:08DeepSeek, we know whatever, whether it was$6 million or$60 million or whatever it was, was significantly less expensive. I think overall, you have one side of the industry showing us that it actually can be cheaper and cheaper and cheaper. But then those with the best interest, remember OpenAI's competitive advantage could be talent to an extent, even though a lot of talents left, they have a pretty deep bench that's pretty impressive. Or it could be resources, cash, and access to compute. So they almost have to make that their game because if that's not their game, they're not going to win. If that's not the game, they're not going to win.

30:48Yep. So we talked about the competition you teased Anthropic. We have so much more to talk about, including the new Anthropic model, what Jensen Wong has talked about, how expensive reasoning is and NVIDIA earnings. And of course, the new Alexa. We're going to do that right after this. Did you know your credit card points and miles can lose value to inflation? Credit card companies often reduce the redemption value of your points and miles. Now, imagine a credit card with rewards that can grow in value. With the Gemini credit card, you can earn Bitcoin or one of over 50 other cryptos instantly with no annual fee.

31:21Every swipe at the store or gas pump earns you instant rewards deposited straight to your account. Plus, sign up now for a$200 Bitcoin bonus to kickstart your rewards. Visit Gemini.com slash card today. Check out the link in the description for more information on rates. Again, if you're looking to invest in Bitcoin but don't know where to start, the Gemini credit card makes it easy. The Gemini credit card is issued by WebBank. In order to qualify for the$200 crypto intro bonus, you must spend$3 ,000 in your first 90 days. Some exclusions apply to instant rewards in which rewards are deposited when the transaction posts.

31:59This content is not investment advice and trading crypto involves risk. The Gemini credit card cannot be used to make gambling-related purchases.

32:10You're used to hearing my voice on the world bringing you interviews from around the globe. And you hear me reporting environment and climate news. I'm Carolyn Beeler. And I'm Marco Werman. We're now with you hosting The World Together. More global journalism with a fresh new sound. Listen to The World on your local public radio station and wherever you find your podcasts.

32:38We're back here on Big Technology Podcast Friday edition, talking about all the latest AI and tech news, including the fact that Anthropic has a new model. Jensen Wang has a stance on how much compute reasoning uses, and the new Alexa. And by the way, Skype is dead. So let's see if we can get to that all in the second half. The first is that GPT 4.5 wasn't the only model here. We have Anthropics Cloud 3.7 Sonnet. It's here. This is from TechCrunch. Anthropics is releasing a new AI frontier model called Cloud 3.7 Sonnet, which the company designed to think about questions for as long as users want it to.

33:15So like we've been talking about, it's a hybrid AI reasoning model, a single model that can give both real-time answers and more considered thought-out answers to questions. And you just choose, do you want the quick response or do you the thinking response and the model represents Anthropik's broader effort to simplify the user experience around its AI products. We're longtime claw heads, I would say on this show. I've gotten a chance to use it. You've gotten a chance to use it. I believe the thinking toggle that we talked about is pretty good. It's almost as good as deep seeks. What is your response that we have another model from Anthropik and the fact that we went not from 3.5 to 4, but from 3.5 to an incrementally better 3.7 what about 3.6 i was waiting for 3.6 that was going to be the big one but we just skipped straight ahead to 3.7 baby i think i've been using as a clodhead i've been using 3.7 regularly on again from the model side it doesn't the the thinking toggle mode which i'll still categorize a bit as product.

34:27Maybe that one lives between product and model is good. Claude Code is definitely a very new offering from them. And I think it's going to be very interesting because still coding to me is the most monetizable, direct to actually productive use case for generative AI as of today. So I think the way they approach this is kind of how I want these model launches to be approached. There's a blog post, there's some tweets, you know, there's, there might be an explainer video here and there, and that's it. And, and we keep getting improvements. And as we wait for 3.9, maybe not 4.0, because that's AGI probably.

35:094.0 is AGI. So yeah. So I will say just having, we're going to get into how they trained the model because it's interesting. But I will say I did an experiment this morning where I've been using, I mean, I think I've mentioned this on the show. I've been using Claude every day as a diet coach where I basically like write down the meals I've had, weigh in, will give me how I did a letter grade based off of the prompt that I gave it about the way that I want to be eating. And it will like count up the calories and grade the foods. It's very good. And it has lots of memory. And so I like, I copied the history, which goes back like probably a month and a half at this point in the latest chat and dropped it into OpenAI's GPT 4.5, Claude 3.7 reasoning, and DeepSeek.

35:53And unfortunately, I'm here to report that DeepSeek did the best job of all of them. You didn't try Grok and have it yell at you and make fun of you? No, I'm good on that. Thank you, though. I think, see, that's like, to me, I want that as its own standalone benchmark. The Alex, what did I eat benchmark that is the leading benchmark for all frontier models going forward. I mean, that's the real life stuff that's actually interesting to me. I do that very regularly. I'll have three tabs open, try the same question across three and see what I get a good answer on. Those are the use cases and the kind of ways that I think everyone, all of our listeners, like approach these models in that way, try different models and try the same question just see what happens and see what you like better i think that's the real way to try to decide what's really happening in terms of progress here versus the more theoretical stuff can i just say one of the big takeaways for me this week is just that reasoning is just freaking unbelievable like it's a true breakthrough and when you use those models you just get better stuff and that to me has been sort of like i think discounted i think in some of the conversation here, but not in our show.

37:10I think we've already always talked highly of reasoning, but in the broader, like the walls hit open AI's toast, like reasoning is both useful and better. I'll agree it's better, but it's still better for many things, but not all things. Again, simple queries, a back and forth, analyze this text, something that's like where all the information is right there in front of you and doesn't require a great deal of complexity, you don't need reasoning for that. And that's more expensive or that's more complicated or it's more time consuming. So I think I agree. I'm still wowed, but there's also the UI element of that, that DeepSeek, again, listing out the questions of the chain of thought as they're coming up was the kind of using the term party trick again, but it just UI feature that makes it so much more real.

38:03And now everyone's doing it. It's amazing how quickly, sometimes it's almost annoying now on ChatGPT where it starts walking me through what it's doing. And now it's thinking when I actually don't want it to, when I'm like, I'm good. I'm good. Just give me an answer. I'll take a little, I'll switch to another tab and come back. It is like your very talkative friend who's like, let me tell you exactly how I got to this. And you're like, nah, it's good. We're just going to go with your answer. I'm going to start on this tab and then I'm going to go here. It's almost, I want like the log afterwards if something's wrong to go back.

38:38But we've talked about this before. The problem that remains is if something is broken in that reasoning process, you can't simply fix it. It's not like I go back and I'm like, okay, on step three of eight, I would rather you have done this than this. That does not exist yet. So at that point, the reasoning is nice, the show of reasoning, but it's not actually, you can't utilize it in any meaningful way. That's fair. So Ranjan, I want you to talk a little bit about this cost efficiency that Anthropix seems to have found in training 3.7, because I think that's pretty significant when we think about how these businesses will operate and whether they need to spend as much money as they are training their latest models.

39:27So on the cost side, 3.7 sonnet, it apparently costs just a few tens of millions of dollars to train. We already talked about DeepSeek. I think, again, it goes back to showing like, what are the real costs involved? There's gathering up some large amount of data. If it's a reasoning model, there's a supervised fine tuning side of it. There's a reinforcement learning side of it, that could involve bringing lots of humans. And again, that literally is like, what is the correct way to get to this answer? Is the answer correct? What rank these outcomes and actually going through hundreds, thousands, tens of thousands of times and training the model that way.

40:10Obviously that's time intensive and it's expensive, but I think it's important to recognize that even Anthropic who has kind of been in the whole big models, expensive models game so far, the fact that they are moving towards this, it almost means, I guess, OpenAI is probably the only player left that's still kind of trying to sell. You need big expensive models to win. So then what do you think about this comment from Jensen, where he talks about, now we're going to go to reasoning and inference, and that's going to be more expensive. So Nvidia earnings came out this week. So they had revenue jump 78 % from a year earlier to 39.33 billion in the quarter, they're projecting 43 billion in the next quarter.

40:57They had they delivered 11 billion of their Blackwell chips. So life is good for Nvidia, but everyone's getting the sense as to like, how is your business going to look if we get more efficient if we go toward these reasoning models. And this is a very interesting statement from Jensen Wang, where he says, AI has to do 100 times more computation now than when ChatGPT was released, basically talking about how the reasoning approaches are more expensive. Next generation AI will need 100 times more compute than older models as a result of new reasoning approaches that think about how to best answer questions step by step.

41:34The amount of computation necessary to do that reasoning process is a hundred times more than what we used to do. So it is interesting to me because, I mean, you look at what DeepSeek did and they found a way to not only do reasoning, but do it more efficiently. And Jensen is saying this thing that seems to disagree with this a little bit. Well, I'm curious what you think, Rajon. I mean, never to speak ill of Jensen Huang. I think he's saying what he needs to say. I mean, if the thesis that things are going to get much cheaper and require less compute holds, we could have the Javon's paradox, which I haven't heard in a little while, but we all heard about that one week.

42:21Again, the idea that the more ubiquitous AI would get because it's cheaper would actually require more aggregate compute. But it's, I mean, it still feels like NVIDIA has to tell that story. And I'm, again, the company blew out numbers again. And even though it's getting caught up in the larger stock market route as of today, but this is still an insane company in terms of its ability to produce and deliver, it still hurts their longer term story, at least with the expectations that have been set by the market. it. Yeah. I mean, it's just one of those things where I'm like, I see his logic and I see where he's going, but I don't really see how, I mean, yes.

43:02I mean, they've talked about how inference is 40 % of their revenue, but I just don't really see how it's going to cost a hundred times more to do reasoning. Maybe I'm missing something. No, I think it's very difficult to try to calculate out? Because even, I guess, the more complex the use cases get, and maybe we'll start unlocking use cases that we haven't even imagined, or AI is going to be applied to areas where we haven't even started to, and those will be the ones that really soak up all that compute. But I agree with you that the idea that it's going to require a hundred, it's going to a hundred times more compute, especially as the trend is everything's getting cheaper, doesn't make sense to me either.

43:51Let me ask you this one thing that I saw from earnings. And I'm curious if you think that it's right. I mean, the fact that they shipped 11 billion in Blackwell chips, the expectation was like three and a half billion. So clearly there's a huge amount of demand for the Blackwell chips, which are the latest generation of NVIDIA chips. All the hyperscalers are saying we'll take as much as we can get, including Andy Jassy at the Alexa event this week. Does that show that there's already enough tangible process. Sorry, does that show that there's already enough tangible progress within AI that merits this further investment of chips?

44:25Or do you think we're just still in the finding out phase? You never want to be in the find out phase. I think we all know what happens after. But I think from the hyperscaler side, it's still like no one backed down from actually Microsoft a bit seem to hedge. And I believe there's some reporting that they're canceling some data center leases. But overall, the hyperscalers are playing the same game that we're going to, it's an arms race for a compute and we're going to continue down this road. And we're going to get into the Alexa event. Maybe it does start to seem like the more complex Alexa gets.

45:05If every single person who has an Alexa is actually actively engaged all day, with Alexa plus, then you start to see that, okay, it's going to require a lot of compute. So if those, like, if it really lands in the way that it's being promised to, maybe that does make sense. But I think as of today, it's just, everyone is to, all the hyperscalers are taking the exact same bet. Right. So we're still in the like scale up infrastructure and maybe this will work not, and this is working enough that we're going to keep investing. Exactly. So I think that, I mean, I am still, I'm still bullish on NVIDIA, but I think if you take this kind of like dubious proclamation about reasoning being cheap, being much more expensive, combined with the fact that like, yes, they're still ordering, but you know, there's a big if at the end of the tunnel.

45:55I do wonder a little bit, like if there's like a potential nasty surprise for NVIDIA coming in a couple of years. No one ever won that prediction in the last few years, at least. But I'm not disagreeing with you, but it's one of those things that it's almost, I'm almost like fearful of saying out loud. Yep. I'm sure I'll eat the words there. And maybe Mark Zuckerberg will be the one that continues to keep NVIDIA running, taking all that ad money and pushing it right into this chip and server company, right? Because now Facebook is going to potentially spin off a Meta AI app in an effort to compete with OpenAI's ChatGPT.

46:35This is according to CNBC. Meta AI will soon become one of the social media company's standalone apps. Joining Facebook, Instagram, and WhatsApp, the company intends to debut a Meta AI standalone app during the second quarter. Of course, they're going to have all the app install power that you have on Facebook to get people to use it via their ad slots and news slots they're going to put in. And Mark Zuckerberg is really intent on basically taking over OpenAI's lead with ChatGPT. He sees it. We've talked about it in the past. He sees this as a big consumer app. He sees that it's growing fast and he doesn't want somebody else to do it.

47:10Same way that he sort of cut off Snapchat and did it with some success against TikTok and Reels. Very funny response from Sam Altman when he sees this. He says, okay, fine, maybe we'll do a social app. He says, it's funny if Facebook tries to come at us and we just uno reverse them. It would be so funny. I mean, I think he's, you get that response from Sam Altman when he's not completely sure of his footing. And I don't really feel that he's sure of his footing against this one because you don't want to go up against Facebook when it comes to a consumer app. It doesn't usually end well. So what do you think, Rajan?

47:46I think that I'd actually, it's been so long since I played Uno, I had to look up what Uno reverse was first. I was like so surprised. I was like, what are you, Uno reverse? Who speaks like that? But anyway. Sam Altman speaks like that. But the same guy who's coming up with everyday queries for our AI benchmarks. But I thought this was a very interesting one because I have been using meta AI more for image generation just because it's very easily accessible. And it's still in a weird place because it lives in the search bar for Instagram and WhatsApp or Facebook. and living in the search bar. I've also accidentally used it when I'm searching for something on Instagram and somehow it pushes me towards meta AI and gives me a weird chatbot response.

48:36So I think spinning it out is a very interesting idea. And then being meta, they would be incredible at quietly guiding people towards that app from all of their other apps. But it's still, is it needed? Is there other ways to integrate it more as a tab in existing Facebook Blue and Instagram itself? That part, I would think it would just be another tab on the regular app versus having people download something. But I see this one as another threads. They'll get some big numbers, but I don't think it's going to be anything too impactful. Yeah, I don't think it's going to work. I would like to see an open AI social network and not for nothing, but open faces out there for the taking.

49:28The AI first social network where all of your posts are commented on extensively, where you have a million friends who are all fawning and love you very much. I think they could go down that road. Open face. Open face. open face i'd be day one user social networking needs like an entire remake and maybe sam altman is the one to bring that to us i mean if anybody can do it maybe it is open ai they're the best product in ai so yeah you don't even need a better model for that i'll give you that ron john that's the product that we've all been waiting for so we got about 10 minutes left and i've saved this for last and I don't want to spend too much time on it because I am going to have a podcast next week covering it.

50:20But Alexa, the new Alexa app is out this or it is not out has been introduced the new Alexa revamp. It's called Alexa plus it is conversational. It is able to accomplish things in the real world. It seems to have an awareness of what happens between your Amazon services. So you can ask it to play a song and then say, all right, can you take me to a point in this movie. That song is on and it will do that based off of Amazon Prime Music and video. It will search your ring cameras for you. It will potentially order you an Uber. You can use it to control the sound in your apartment if you have echoes with conversational tones like, can we have this song play in that room or can I want to hear it over there?

51:05it was a very impressive, very impressive demo, I felt. And it was live, unlike Apple Intelligence, where Apple Intelligence was a promise. And it seems like a lot of this Alexa stuff is going to work. So I do want to do this preface that we're going to have Panos Panay, along with Daniel Rauch, the head of Alexa. So it's going to be a fun conversation that's coming up on Wednesday, you're gonna hear a lot more about that. But Ranjan, I'm very curious, like, what your reaction was in our chat, I think in the discord, you dropped this, uh, German tweet where he talked about how like this was chat GPT voice on steroids and, uh, Apple has to be embarrassed at this point for how bad, uh, Apple intelligence is.

51:46But what was your takeaway looking at the Amazon news? And, uh, and if you can, if you want to, I'm just going to say, if you want to, you can say what it might mean for Siri? Oh man, I spent so much time rewiring my entire house for HomePods and I watched that event and I want to just go back. I want to switch all to Alexa and I'm sure we'll definitely talk about it more after your next episode, but like it looked good. It looked exactly like what it should be. It looked like putting chat GPT voice mode on a device or Gemini voice on a device or just what basic voice interaction should do right now.

52:32And I'm avoiding hitting my table right now. So there's feedback on the mic because you can tell I'm worked up right this time. I'm just telling me, sorry, listeners. Oh my God. Like that's all it should be doing right now. And it did what it's supposed to do in voice. We know generative AI voice is that good right now. I think the only thing that I think could be a little problematic for Alexa Plus is like Amazon does not have a great reputation in terms of privacy or just overall, like it can still be a little creepy. The reason I got rid of my Lexus was it would always ask these follow up questions, which you could not turn off.

53:13You'd be like, what's the weather? Oh, here's the weather. And can I interest you in these three other things? And you couldn't turn it off so i think i mean the way they portrayed it becomes your like really trusted companion that you're sitting there sharing yourself with that's a big ask in terms of trust so i think in terms of the technology i'm pretty confident they're there in terms of like getting people actually comfortable with interacting with your voice device in that way we'll see but But man, it's going to cost me a lot of money. So yeah, we did talk about it yesterday. I'll just give a quick preview of what this is going to look like.

53:51I mean, we talked about it yesterday. And they are aware that the talking back to you and being proactive can be pretty interruptive and annoying. And I think they're paying attention to that as they roll this out. So that'll be at the end of the conversation for those that listen. But yeah, I thought it was really interesting. I think Amazon has a shot here. I wrote about this in Big Technology that basically all big tech companies want to build a universal, contextually aware assistant that helps you get things done. And Amazon has a pretty good shot to be the one that pulls it off, especially because they have a working demo of this and it seems like it's going to go live next month.

54:31And I don't know. I mean, they don't have an operating system, which is on one hand, they don't have a mobile operating system. On one hand, that's a curse because that default matters a lot. We know Google pays Apple$20 billion a year to be the default search engine on the iPhone. However, that does let them use other productivity services and not privilege their default. And I think these companies privileging their default productivity services have been sort of the downfall of the modern AI assistant. Like if I'm using an iPhone and I can't use Google Calendar or Gmail in there because Apple is so dedicated to whatever, Apple Mail, that ruins Siri for me.

55:13But Amazon doesn't have that problem. And so I was speaking with actually the head of Prime who was mentioning that, yeah, I use my Google Calendar in my Echo devices and it works just as well. So that could be a blessing for them. Yeah, I 100 % agree. Though, speaking of Prime, the way they rolled out the pricing, I think was the most like savage and amazing Amazon move ever. Again, I forget what the monthly subscription is. Is it like - $19.99. Or was it$14.99? Let me see. It's something that most people would not pay as of today. But as an Amazon Prime member, you get it for free. So it just kind of like assigns this additional benefit to being a Prime member, which if you're a Prime member, you shop on Amazon X percent more.

55:59So they're going to kind of assign this incredible value to it on day one. And you're going to feel like, oh, well, now if I was questioning, should I renew my Prime membership? Well, it's a gimme. I mean, I'm getting Alexa Plus for free. Ronjan, listen to this. Alexa Plus costs$19.99 a month. Prime costs$14.99 a month. Oh, wait. I thought it was free for Prime. No, if you have Prime, it's free. So basically you could pay an extra$5 to just get Alexa Plus or five less dollars to get all of Prime and Alexa Plus. Oh, okay. That is savage. That is savage. I mean, if Lena Kahn was still around, I don't know what she would say, but my goodness.

56:42But she's not. But she's not. Okay, before we go, we need to talk about Skype. Microsoft is killing Skype. This is hot off the presses and makes me very sad. As from TechCrunch, after kickstarting the market for making calls over the internet 23 years ago, Skype is closing down. Microsoft, which acquired the messaging and calling app 14 years ago, said it will be retiring it from active duty on May 5th to double down on Teams, Skype users of 10 weeks to decide what they want to do with their account. It's not clear how many people will be impacted. The most recent numbers that Microsoft shared were in 2023, where it said it It had 36 million users, a long way from Skype's peak of 300 million users.

57:29We do look at tech with a critical, sometimes hopeful eye. I will say that Skype is one of the products that I've loved the most on the internet. I just have good feelings about it helping me make international calls and calls to friends. And you could play different games on it back in the day. And that little squeak that it makes when you get a message will forever remain in my heart. Rest in peace, Skype. We bury you the week after the Humane pin. goes the way of the Neanderthals. And I'm much sadder about losing you than the wearable AI device. I had my first remote job interview on Skype.

58:06I agree. International calls. It opened up the world. Sold for$8.5 billion to Microsoft in 2011. And sorry you had to get caught up into a corporate battle with Microsoft Teams that clearly won and ended up basically, I think the last few times I ended up on Skype, I had all these messages that were clearly phishing and scam things that were just barraging my Skype account and did not open it after. Goodbye, Skype. Goodbye, Skype. And goodbye to all of you. But just hopefully for a couple days because I'll be back on Wednesday with those two Amazon executives and Ranjan and I will be back on Friday.

58:47Ranjan, thanks so much for coming on the show. See you next week. alright everybody thank you for listening and we'll see you next time on Big Technology Podcast

From the publisher

Ranjan Roy from Margins is back for our weekly discussion of the latest tech news. We cover 1) OpenAI's release of GPT 4.5 2) Is GPT 4.5 a major advance or what? 3) What better EQ gets you in an AI model 4) What reasoning advances can be built on top of GPT 4.5 5) Is AI product or model more important? (cont.) 6) Gary Marcus says OpenAI is in trouble 7) Anthropic releases Claude Sonnet 3.7 8) NVIDIA earnings 9) Jensen says reasoning costs 100x typical LLMs 10) Meta wants to build a standalone AI app 11) Will Alexa+ work? 12) RIP Skype
---
Enjoying Big Technology Podcast? Please rate us five stars ⭐⭐⭐⭐⭐ in your podcast app of choice.
For weekly updates on the show, sign up for the pod newsletter on LinkedIn: https://www.linkedin.com/newsletters/6901970121829801984/
Want a discount for Big Technology on Substack? Here’s 40% off for the first year: https://tinyurl.com/bigtechnology
Questions? Feedback? Write to: bigtechnologypodcast@gmail.com

More from Big Technology Podcast

All 399 episodes
OpenAI’s New Model, Jensen’s Bold Claim, Alexa+ Is HereBig Technology Podcast · 59 min
Listen in VO