In short
Big Technology Podcast - Episode Summary
Episode Title
Too Many AI Companies, Amazon's Alexa Upgrade Awaits, RIP Humane Pin
Host
Alex Kantrowitz
Guest
Ranjan Roy from Margins
Date
[Insert Date]
---
Episode Overview In this episode of the Big Technology Podcast, host Alex Kantrowitz and guest Ranjan Roy discuss significant developments in the tech world, particularly focusing on the saturated AI market, Amazon's anticipated Alexa upgrade, and the demise of the Humane Pin. The conversation offers a nuanced exploration of the current state of AI companies, the implications of recent advancements, and the broader cultural impact of technology.
---
Key Topics Discussed
- AI Industry Landscape
- Satya Nadella's Criticism:
Nadella criticizes the practice of "benchmark hacking" in AI, suggesting that claiming AGI milestones without practical impact is nonsensical. He emphasizes the importance of real-world applications over meeting arbitrary benchmarks.
- Emergence of New Startups:
Discussion on the influx of AI startups, including ex-OpenAI CTO Mira Murati's new venture, Thinking Machines Labs. The proliferation raises questions about differentiation and sustainability in a crowded market.
- Foundation Models and Commoditization:
The commoditization of foundational models is noted, with many companies offering similar capabilities, leading to concerns about a potential bubble in the AI sector.
- Evaluating AI Models
- Benchmarking Models:
The hosts delve into how AI models are evaluated, emphasizing the need for practical assessments rather than relying solely on benchmarks that may not reflect real-world effectiveness.
- Chatbot Arena:
A mention of Chatbot Arena as a platform for comparing AI model outputs, with Grok3 currently leading the rankings.
- Amazon's Alexa Upgrade
- New Features for Alexa:
Amazon is expected to announce an upgrade that allows Alexa to handle multiple prompts and act on behalf of users, increasing its utility in everyday tasks.
- Concerns and Expectations:
Kantrowitz expresses hope for the improvements while noting the historical inconsistencies of Alexa's performance.
- The Fall of Humane Pin
- End of the Humane Pin:
The episode reflects on the closure of the Humane Pin, which failed to resonate in the market despite raising significant capital. The hosts humorously lament its failure and discuss the marketing missteps that led to its downfall.
- Cultural and Market Implications:
The conversation touches on whether every attempt at innovation should be applauded or if some failures deserve criticism due to poor execution.
---
Key Takeaways
- Market Saturation:
The AI market is experiencing oversaturation, leading to questions about the viability and uniqueness of new companies.
- Importance of Practicality:
There is a pressing need for AI models to demonstrate practical utility rather than merely competing on benchmarks.
- Cultural Reflections on Technology:
The discussions highlight broader cultural attitudes towards technology and innovation, particularly how society balances excitement for new tools with skepticism about their long-term impacts.
---
Closing Thoughts The episode encapsulates a critical moment in the intersection of technology and society, urging listeners to consider not just the innovations themselves, but their implications for everyday life and the future of the tech landscape. The hosts encourage a thoughtful approach to evaluating new technologies and their potential to both enhance and complicate human interactions.
---
Join the Discussion Listeners are invited to engage further through the Big Technology Discord community to share insights and continue conversations about the evolving tech landscape.
---
Note
For those interested in the full episode, it is available on various podcast platforms.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00There are probably way too many AI companies, but why not? Let's add another. Humans criticize Open AI's deep research. Amazon has an Alexa upgrade upgrade on the way. and the Humane Pit is dead. Off to HP to check your toner. That's coming up right after this.
0:20You're used to hearing my voice on The World bringing you interviews from around the globe. And you hear me reporting environment and climate news. I'm Carolyn Buehler. And I'm Marco Werman. We're now with you hosting The World Together. More global journalism with a fresh new sound. Listen to The World on your local public radio station and wherever you find your podcasts.
0:47Welcome to Big Technology Podcast Friday edition where we break down the news in our traditional cool-headed and nuanced format. We have a wild show for you today because there's a whole lot of news in the world of artificial intelligence and the tech world in general. We got a new AI company. It is from the ex-open AI CTO, Miramirati. We also have new models from Grok. We have big news from Microsoft, which we're going to touch at right at the top. And we'll look at the death of the Humane Pin. Oh, by the way, there's also an Amazon Alexa upgrade coming. Joining us as always to talk about it on Fridays is Ranjan Roy of Margins.
1:20Ranjan, great to see you. Welcome to the show. RIP Humane Pin. What could have been? I know we're both feeling a state of loss. We're both feeling a state of hurt and harm. We're going to try to do today's show. Work through this. Yep. With respect, but also holding back some of those feelings till the very end so we can analyze what's happening with the appropriate nuance, the appropriate tone. We'll analyze and we'll emote at the end for the humane pin. I want to start this week with a quote from Satya Nadella, who's been obviously listening to this show because he's channeling Ranjan in a recent interview with Dwarkesh Patel.
2:04He's on the Dworkish podcast talking about Microsoft's new announcements, one in quantum, another world model. And he says this, us self-claiming some AGI milestone. That's just nonsensical benchmark hacking to me. The real benchmark is the world growing at 10%. Very interesting statement for him to make, quite provocative, especially because we keep hearing from the company that Microsoft has funded, OpenAI, that they keep getting closer and closer to these AGI benchmarks. Yet, Satya is coming out there and I think rightly saying, hey, listen, it's not about getting those benchmarks. It's about can you actually have practical impact on the world?
2:45Just one statement in a long podcast that I think is worth listening to, but pretty revealing in terms of where Microsoft stands today, where the discussion is going today, and also maybe how Satya feels about OpenAI. Am I reading too much into it, Ranjan? I don't think you're reading too much into it. I think we could break this down. There's two main levels you can look at this. One is actually the legal and contractual side, because remember, Microsoft and OpenAI's entire deal actually hinges on when AGI is achieved. And that's a key part of it. And we've been joking for weeks or even months now that all Sam has to do is say, AGI, and we're there.
3:23So I think him recognizing this and openly kind of calling out this self-claiming of AGI milestones is ridiculous, is important in whatever is going to happen in the future of OpenAI and Microsoft's relationship. But the other is nonsensical benchmark hacking is my favorite line. I feel like, I mean, I want to put that on my own LinkedIn right now because Satya is just after his we're good for the$80 billion and now him coming up with phrases like this. This guy, he's got some words. He's got some lines that I would not have expected from him. Yeah, Satya is a low-key baller. I mean, he has some sharp elbows.
4:05Yeah, I think put him in the cage match. put in, Elon never challenged him, but I think Satya actually is currently now my favorite in the, uh, in the octagon that never happened. And it's also just so interesting to me that he's talking about it, the nonsensical benchmark hacking in the context of, uh, open AI, which is very proudly proclaiming how they're doing against the benchmarks. All right. So speaking of open AI, they have lost a lot of people recently, shall we say? 2024 was a year of exodus for them. Where are these people gone? Well, we know Ilya Sutskever is doing super safe super intelligence or safe super intelligence.
4:42Like, I keep forgetting exactly what it's called, but they're raising billions and billions of dollars without building anything yet. And now we see the emergence of another former OpenAI executive, former OpenAI CTO Mira Mirati. She was the CEO for the weekend that Sam Altman was fired. She is out with a new company called Thinking Machines Labs. but it really is true that we have so many companies that are working on, I think, the same thing and you're getting to the point where you're like, what are we all doing here? So let me just read from TechCrunch. It's called Thinking Machines Lab. The startup intends to build tooling to make AI work for people's unique needs and goals and to create AI systems that are more widely understood, customizable, and generally capable than those currently available.
5:28Roddy is heading the lab as the CEO. She's also brought over a cast of characters from her former company. The co-founder of OpenAI, John Shulman, is the chief scientist at this new company. And Barrett Zoff, the former OpenAI chief research officer, is the CTO. So they've got the former royalty from OpenAI over there, outside of Ilya, who's doing his own thing. Let me just read you one more bit from the TechCrunch story because I do think I got to the point this week where I was reading words and I was just like, what does this all mean? Thinking Machines Lab plans to focus on building multimodal systems that work with people collaboratively that can adapt to the full spectrum of human expertise and enable a broader spectrum of applications.
6:17We talk a lot about PR on the show and about marketing and positioning because we do think that oftentimes that does reveal the inner soul of a company or at least the inner direction of a company. And I can't for the life of me figure out what this word salad means, Ranjan. And maybe you have some clarification here because I'm scratching my head. What's funny is I was just about to ask you as you were reading it, what does that actually mean? What does Thinking Machines Lab do? So I'm a little disappointed that you already preempted me by acknowledging it's a word salad, because I think you also had left out this part from their blog post.
6:54The scientific community's understanding of frontier AI systems lags behind rapidly advancing capabilities. Knowledge of how these systems are trained is concentrated within the top research labs, limiting both the public discourse on AI and people's abilities to use AI effectively. Again, more word salad. I don't want to say an AI wrote this, but it certainly seems like there's no really concrete, tangible thing that this company will do. It's a tough one, because you have a lot of very smart people. We have very regularly said OpenAI is effectively a research house that had a business that kind of accidentally walked into a business.
7:37Sam Altman himself has declared or described his company as this kind of institution. So to me, it's almost odd that you look at how Ilias Itskever and Safe Superintelligence, which again has raised or is raising some ungodly amount of money to achieve some nebulous goal where they're essentially saying, I think they said that they're not even planning on having a product anytime soon. Here you're having Mira Morati, who I'm sure is going to be able to raise a bunch of money, launching a company with a name that's kind of hard to say, like Thinking Machines Lab. I feel it should be labs. I lost it pretty good right at the beginning.
8:18By the way, it's also a copy of a name of another company that was established back in the day. Oh, really? Yeah. I mean, yeah, Thinking Machines is a good one. Call it labs, though, Mira, please. But overall, the whole alumni ecosystem from OpenAI, you can tell it's a problem because they can raise money so easily without any plan, product, or anything. simply on promise that I think this whole, you know, like the PayPal mafia ecosystem of companies did quite well. I don't think the open AI alumni ecosystem is going to do as well. Maybe they're going to produce some of the greatest research out there, but I don't think they're going to build the actual businesses of tomorrow.
9:08Yeah, look, I think we're going to also have to talk about how there are so many AI companies and what are they going to actually offer that's differentiated? Like, do we need another one working on? It seems like that's going to work on foundational models. I don't know. We still don't really know. By the way, I've invited Mira Maradi on the show. I think I got her email right. I tried about like 17 ,000 different permutations until I figured it out. So Mira, if you're listening or if Mira's representation is listening, please come on. But before we go on to the fact that there are so many of these companies and how's this going to shake out.
9:40I just want to ask one simple question because you did mention Elias Itzkever. He's raising this money. He has this safe super intelligence lab. It's not going to release a product. I mean, Ranjan, to me, I just want, I'm just wondering, how are you going to have safe super intelligence if you have to build a product that delivers returns that VCs anticipate? Like you have taking that VC money, the contract is growth. And if you're going to actually do the job, you're going to have to grow the way that they expect. And I tried to put this to Reid Hoffman when he was here and he seemed to brush it off, but it just doesn't seem to me like this is compatible.
10:20And it's sort of been baked into the entire AI mythology. We need more money, so therefore we grow big, but also we're concerned about the impact of this technology. But then again, like you're raising unprecedented rounds. How is that possibly compatible with safety and slow rolling things if you're actually afraid of them. Yeah, just confirming here, the safe superintelligence reporting was they're looking to raise over$1 billion in capital at a valuation of$30 billion. Now, how you even come up with some kind of math to get to a valuation on something where they're not even announcing or even pitching a product, I don't understand.
11:03But I think if I'm to try to see the other side of this, let's look at OpenAI is generating revenue. They're losing a lot of money, but they're still, they're growing quite fast. Claude came out of this model. Anthropic, I mean, Anthropic has been playing this same game. Mistral, I don't know how they're doing in terms of revenue, but they clearly played the same game. So I guess if we take Anthropic and OpenAI, the game presented this really specific way of working has had some reasonable success stories. I hesitate calling them success stories because they haven't proven themselves as actual sustainable businesses.
11:46But they're generating revenue pretty quickly in terms of VC growth expectations. So maybe the assumption here is we're just going to get more of that exact pathway that these two companies got or three or the early entrance to this. Just kind of basically, it's clear in my mind, like there is no safe path to developing this. I don't know. I'm just kind of. Oh, the safety, I think, is a whole different conversation. Yeah. It's just like, I don't know. I don't see how it's feasible. And I understand there's like the promise. But I don't know if you're taking a billion dollars of venture capital and you're telling me that like you're I'm sure you're in the in the meeting with the VC talking to them about how they're going to deliver the returns.
12:29You're not really working for society. You're working to deliver that money back. I would love to have seen that pitch from Ilya and Save Super Intelligence. like it almost feels like an episode of HBO Silicon Valley where you can imagine him sitting across the table. There's some VCs on the other side and him looking them in the eye and saying, we are not going to build a product. We are not. And then be like, Oh my God, this guy, this guy knows something. What do you need? What do you need? Yeah. It's, it's totally, it's totally bizarre. So there's, so I think this sort of brings us home to this piece that Casey Newton wrote.
13:08there are probably too many AI companies now. It's a great piece and there's a really hilarious subhead. Everyone has a model. Almost no one has a business. I'm just going to read a little bit from it. He says, I've been talking with tech executives about the likelihood of a bubble in artificial intelligence. Everyone I've spoken to described the experience of seeing another AI company come along with a slightly better or cheaper model than they've been using and quickly swapping it in to replace the previous one. There's a reason foundation models have become so commoditized. Most of the original research was published openly for anyone to make free use of.
13:40Catching up to the state of the art is just is often just a matter of acquiring the necessary hardware and a smattering of talent. All the big AI labs are losing money as models improve and prices fall. It seems increasingly certain that a wave of consolidation and even outright failures will follow. And that is kind of, you know, when I when I think about what Ilya is doing, when I think about what Mira is doing, when I think about what, even the fact that we have all these established ones, the open AIs, the Anthropics, and Grok now, and DeepSeek. I think Casey has a point. Well, I think he's missing one big part of this.
14:17Currently, you kind of have two categories. You have the model builders, and then you have the model users, the productized side. I don't know what we want to call it. But like on the more product-oriented side, there's still endless products coming out like granola ai have you been using it for uh note taking in meetings it's really good tell me about it yeah it's otter and others it's like you know are all competing in this space but it basically both transcribes your meeting it does it without i don't know if anyone's ever used otter but there's this like kind of really creepy the otter assistant joins the meeting and shows up on zoom or google meet as like a separate guest and everyone gets freaked out versus Granola uses your local sound on your computer, takes notes.
15:04However they've trained the model to actually summarize the notes and give you action items is good. So here's a product that is probably worth paying like$10 a month or whatever their pricing is, and it could work. Maybe it'll be a sustainable business or maybe OpenAI will somehow or any other large player will be able to kind of just abstract away that specific service and kill their business. But there's that side of it. Then on the actual foundation model side, as you said, we have the OpenAI's and the Anthropics and the Mistral's and now Grok from XAI. All of these, their products all look kind of similar.
15:47There's a chat interface. Obviously, OpenAI's operator was a completely new interface and new looking thing. Gemini and OpenAI all have good voice modes. You know, it still all looks the same. So that's the side where I do believe, Casey's right, that that part of the entire stack will get completely commoditized. And the idea that they're the ones continuing to raise the most money and have the richest valuations does not make sense to me. Let's go to Benedict Evans. He's talking about this commoditization issue that Casey pointed out. OpenAI and all other foundation model labs have no moat or defensibility except access to capital.
16:30They don't have product market fit outside of coding and marketing, and they don't really have products either, just text boxes and APIs for other people to build products. Deep research, which we talked about last week, is one attempt amongst many, both to create a product with some stickiness and to instantiate a use case. But on one hand, perplexity claims to launch the same thing a few days earlier. And on the other, the best way to manage air raids today seems to be abstract the LM away as an API call inside software that can manage it, which of course makes the foundation models themselves even more of a commodity.
17:06I think this is such a good point from Benedict Devins, basically saying like, what is defensible if the best way to use these things is to sort of implement them in your software and put the proper controls in so that you cannot, so you don't have the same errors that they do out, you know, off the shelf. And I think he makes this other good point, which is that, and it's also something that Casey pointed to, look how many deep researches we have at this point. I mean, there's four of them. There's four deep researches. OpenAI has one. Google has one. Google was first. Grok has one. And now if you ask me like what are the modes for these companies and where are these you know trillion dollar valuations coming from i don't know i think some of these skeptics have a point no i welcome benedict to team it's the product not the model we still haven't gotten our t-shirts made but if and when we do i'll make sure to send you one because yeah he he laid it out perfectly that But having built on this stuff, it's so easy often to just switch whatever model you're looking at.
18:17Like you build something and then you just change an API call and that's it. And as you said, deep research, what's almost terrifying and amusing is when I read the words deep research in this quote, I actually didn't know what he was talking about because I've actually started using Perplexity's deep research this week, which they launched off of DeepSeaCar1 within days. So yeah, I think it's very clear that those kind of interfaces, those kind of use cases are going to be completely commoditized. I still give Google an advantage here because the distribution side of it is going to become even more important.
19:02Gemini is still getting there in terms of the integration in Gmail and other areas is still not great. But I think as those get commoditized, the distribution becomes even more important because if you can inject these type of features into places people already are, it becomes a lot more useful than having to get more people to your platform. Though ChatGPT actually just came out today and said now they have 400 million. So speaking again of this commoditization of everything and has DeepSeek been commoditized at this point. So this is actually coming from the Big Technology Discord. And look, folks, I'm not going to, I won't spend too much time hitting you all over the head with the Discord pitch, but it has been pretty fun in there.
19:49We have 54 people in there and it's for Big Technology paid subscribers. And if you want to join, you can just go to bigtechnology.com, find the story that says, let's talk DeepSeek, AI, et cetera, on Big Technology's new Discord server. sign up for the paid tier and then join us. I think it's been awesome, Ranjan. The signal to noise has been insane. Like it is some of the highest value conversations I'm having on AI already. Oh, no, I completely agree. Signal to noise ratio, I think, is how my other chats certainly do not equal the same quality in signal to noise. But yeah, I've been learning a good deal.
20:27It's been cool. So we have some real builders in there and I just started a memes channel. um so don't just that we i apologize for what i'm going to put in there but it's going to be all chats go so yeah sign up for the big technology uh premiere uh the premium uh subscription it's just eight dollars a month or 80 a year and you can uh join us in the discord um and and if not no worries we'll just talk with you on the podcast here so this is from the discord though someone wrote it looks like google just deep seeked itself deep seek of course dropped the price of a lot of stuff. And here we go.
21:01This is Gemini 2.0 flash thinking. An input token is seven and a half cents per million on flash thinking. And it's 55 cents for DeepSeek R1. Output token, 30 cents per million. And DeepSeek R1,$2.19. It is, I mean, I, you know, whether the quality is the same or not it is just amazing to me that like talk what happens when it's a commodity right the margin compresses i gave a talk about this at web summit two years ago the margin will compress when everything is a commodity margins are flying out of of the game right now for these foundational model companies well you wanted javon's paradox everybody you're getting it because if when we're all two weeks or three weeks ago now talking about javon's paradox and the idea that when things are commoditized, they will be used more.
21:54And that's the like really, I don't want to say desperate, but it's the argument that, okay, if this goes down from what used to be like three or four bucks down to seven cents, but we're going to now distribute this at a much larger scale because people are going to build a lot more with it. So we'll make more money. I don't know. I don't know. I think it's the product, not the model is just getting even more real for this. And Gemini 2.0 Flash Thinking, not the greatest name, which we've discussed many times. They get worse. It's amazing. They get worse. But you know what? The models are getting better.
22:30I've started using Gemini a lot more. Gemini Voice is really, really good. The latency on it, the voices they use on it, the conversational ability of it. So I think that, again, this going back from a cost standpoint. And also remember, everyone who is already sitting in Google Cloud and Microsoft Azure, if these models and the actual building capabilities within those get really good, they're always going to beat an open AI. Like if you already have a contract and an entire customer success team with Google, why would you then waste your time with open AI if it's more expensive and not as good?
23:12I mean, you wouldn't. Yeah. That was a leading question. And an argument for Microsoft also. And there was this hilarious moment in the Dwarkesh interview with Satya where he goes to Satya. He says, all right, you tweeted about Jevin's paradox after DeepSeek. What is so expensive about artificial intelligence today that I would use more of it if it was cheaper? Because for my vantage point, it's already pretty cheap. It was a great question. Of course, you and I just paid 200 bucks for unlimited Chachapiti. So this is like a stupid thing for me to be saying in that context. But I think the big problem is, again, like what are people going to do with it?
23:51Not they want to use it so much that it has to be cheaper for them to use it. Now, maybe on the enterprise side, it makes sense to, you know, this Jevons paradox things make sense, especially in the eyes of Satya Nadella, who's selling to enterprises and selling that compute. But it was a really good question. And Satya was basically like, listen, And he's like, the price needs to come down and the model needs to be better also. I would disagree on that because it's still such, even when OpenAI, let's say it's not 400 million, but I think whatever the other numbers, 200, 300 million of users of ChatGPT, it's still tiny relative to the overall population.
24:30And the number of use cases for the average person are still tiny. So I do think that most people, the idea of spending$30 a month on ChatGPT Pro or$20 or whatever it is, they're not ready to do that. And also knowing that most of these companies are losing gobs of money, so they're essentially subsidizing our, meaning you and me, our use. So I do think that there is a world that more use cases that the average person adopts, it'll be important, the actual cost to the individual. I accept that. Ranjan for Microsoft CEO. That was even a better answer than Satya gave. I don't know. After his nonsensical, what was it?
25:16Benchmark hacking. Nonsensical benchmark hacking, Satya gets to stay. He'll stay. He'll be the king. There's another, Lord Almighty, another AI model to talk about this week. We'll go quick through this. Elon Musk's XAI has released its latest flagship model, Grok 3. This is according to TechCrunch. Grok is XAI's answer to models like OpenAI's GPT-4.0 and Google's Gemini. It can analyze images and respond to questions and powers a number of features on X. XAI has used an enormous data center in Memphis, which we've talked about here on the show, containing around 200 ,000 GPUs to train Grok 3.
25:51XAI claims Grok3 beats GPT-4-0 on some benchmarks. An early version of Grok3 also scored competitively in chatbot arena. I used it. It spells out its chain of thought in some really interesting ways. It is supposed to be kind of edgy and real, but I always find it to be... I want it to be good, but I find it to be cringe, and I feel like it's like the Steve Buscemi thing, like how do you do, fellow kids, when it responds to me. But I'll just say this about Grok. I want to know how well this performs. I think that Elon is such a polarizing figure that the people that like Elon say it's the, you know, the next generation of model and better than everything else.
26:33And the people that hate Elon say that it's filled with nonsensical benchmark hacking. I mean, it's really, really tough to get a solid read on how it's performing, you know, maybe outside of chatbot arena. What do you think about this, Ranjan? Well, I think the overall space of benchmarking, for me personally, I don't want to call it nonsensical, and Satya is using his words, but like, to me, the way like the benchmarking is always done at such a theoretical level, using whatever test has been developed or using, and we saw this with like, what was it, 4.0 was 90 % or 80 % of the way to AGI basis.
Read the full transcript
27:16on a test that was defined in 2021. I think to me, like you see this stuff and you feel it when you're actually using these products. And it's so hard. And I don't know, like, I feel there needs to just be like a normie benchmarking where you just ask, you don't try to like fool it with a math problem, or you don't try to like, have it create some like PhD level thing, but you just ask it some like pretty straightforward questions and see if it gets it wrong or right. I mean, I've even done that where if you're trying to test some type of RAG tool, like retrieval augmented generation, put in a bunch of documents and ask it like five questions that you already know the answer to, see if it gets it right or not.
27:57And a lot of the times, like OpenAI connecting to Google, ChatGPT connecting to Google Drive, even with Claude, you ask questions and it just doesn't get stuff right. So like numbers, really specific facts. I don't know what benchmark would actually show that, but... I think it exists, though. I mean, isn't that chatbot arena where basically they put two outputs of chatbots side by side and they let you vote on what the better one is? That's the closest we have. I agree. That's the closest. But who is using chatbot arena? Come on. People who care about this stuff. Exactly. But the real value is going to be accrued to people who don't currently care and who just want to use it and not spend their time in chatbot arena.
28:39No, I'm not saying Chatbot Arena is a valuable program in and of itself, but I think we can rely on it as a pretty good evaluation. And, you know, as we're talking, I'm taking a look at Chatbot Arena. And guess what's number one on Chatbot Arena? It is Grok 3. So there's an answer. Yeah, but you don't think that can get gamed pretty well by... I guess it could. I guess if Elon's fans go into Chatbot Arena and they say, let's find the edgiest, realist answers and vote them up. Yeah, ask some questions where you want something slightly more not politically correct. And then you'll guess pretty easily which one is Grok3.
29:21And again, if they're the type of user that would spend time on Chatbot Arena, I think in terms of persona would be an Elon fan. So I think like this stuff can get gamed, basically. I want to know when my mom is using ChatGPT versus Claude or whatever it is and whatever questions she wants to ask, who's going to give the best answer? Let's go back to Casey Newton. He says, if you accept that Grok is a state-of-the-art model, not a single person working in AI believes it will stay there for long. Leading AI labs push out new models every few days, and any innovations are almost all quickly copied and absorbed by their rivals.
30:02Speaking of people who work inside these labs, you get the sense that none of that really matters. To the true believers, AI is the final technology, the one that will invent all the others, and almost all the rewards lay simply in getting there first. They seem not to be building traditional moats for their businesses out of a sense that they don't really need them. That once super intelligence arrives, the world will shift from scarcity to abundance and the need for money will disappear. For now, that vision remains strictly in the realm of science fiction, but a lot of bills are going to come due between now and then.
30:41I think that is a really great analysis and sort of really ties together a lot of what we've been talking about. Where's the ROI on the Ilya thing? Can Mira Murati's company exist? Where's all the products coming from these labs? I think Casey puts it really well and really succinctly. The people that believe in this, and I've heard this before, either believe it goes to infinity or zero. And so therefore they're investing it and therefore they're betting on it. and basically like it doesn't, as long as it survives long enough to the point where it gets to be good enough that it can invent other things, then they're good.
31:18But this is, and again, I'm speaking about this with Reid Hoffman a couple of weeks ago, that's a big bet. That is a very big and dangerous bet to make. It's a big and dangerous bet. I'm guessing a lot of the investment community either made paper returns on OpenAI already or has FOMO for not being in the early rounds of open AI. So I think, yeah, this feeling or thesis and me sitting here in New York, I don't come across that many people. I feel like if you're in Silicon Valley, maybe you'll hear this more, but I think Casey put it really well. There is this almost religious belief that the game here is to just build this either AGI or super intelligence or whatever you want to call it, and that will magically solve all the business challenges.
32:08Everyone will pay you for the product. Everyone will pay you untold amounts of money for the product. And I guess I don't look at it like that. It's kind of like a science fiction investment thesis. I mean, I guess you have to be believing in science fiction. In some sense, you're investing in technology, but this is pretty bold. Let's go back to that pitch scenario of Ilya and the VCs because I wish Silicon Valley on HBO was still around because right now... I could never watch that show. It's way too close to home for me. Just couldn't. I loved it. I lived it when I was living in the Bay Area.
32:43So I was good. That's true. I saw some crazy stuff. There's a scene in Kevin Ruse's book about some young kid telling him that he's building an automation company and calling it a boomer remover. And yes, and I was there when he was telling both of us this. I'm good with the silver I'm good with Silicon Valley man I mean I I'm happy I spent time there I'll probably live there again at some point uh but I don't need to watch that show I've seen it that's fair that's too much and and what about um research removers so you know let's go back to Benedict Evans just for a minute so we've wrapped our too many AI companies uh segment probably too many but if they're right about this science fiction vision then jokes on us uh but as as As I was reading through, I actually went to, I read through the full Benedict Evans post.
33:35And it was actually quite interesting because Benedict Evans is a researcher, an analyst, and he put deep research through the motions, basically trying to see if it was capable. And what he found was quite interesting. He found that basically it does decent research, but it's often quoting from surface level sources, often misses obvious sources that are more definitive. and it's just incomplete. Let me read from him. Are you telling me that today's model gets this table, one that he had to produce, 85 % right, and the next version will get it 85.5 % or 91 % correct? That doesn't help me. If there are mistakes in the table, it doesn't matter how many there are.
34:20I can't trust it. If on the other hand, you think that these models will go to being 100 % right, that would change everything. But that would also be a binary change in the nature of these systems, not a percentage change, and we don't know if that's even possible. We don't know if the error rate will go away, and so we don't know whether we should be building products that presume the model will sometimes be wrong, or whether in a year or two we will be building products that presume we can rely on the model by itself. That's quite different to the limitations of other important technologies, from PCs to the web to smartphones, where we knew in principle what could change and what couldn't excellent analysis here basically he's saying like these models keep improving so they get like 85 right 90 right 95 he's like why am i going to use a research report a research tool that i know may never get to 100 right and if it does get to 100 right then it's a totally different tech what do you think about this i like this i think this is a really smart take on it, especially the part around, like, should we be building products with the assumption that this is 85 % right or 91 % right or whatever it is?
35:31Should we be educating people today on how to use a research product that is, call it, 85 % right? There's tremendous value in that. I do believe. I've been using it more and more myself. And again, I'm not going to pay chat gpr open ai 200 bucks now because perplexities is pretty good it's good enough that i'm okay not spending 200 bucks is it a starting point is it is that 85 good how should i look at the data being presented how should i look at the sources like me spending time on how to use it at this error rate is actually valuable for me and then it makes me a better user of it But these companies are not going to do that, though.
36:18They're not going to build a product for an 85 % because they have to promise that it's going to get to 100%. So now they're going to keep pushing. They're not going to improve the product. And now this is what I worry about. And I've been saying AI has a brand problem over and over again because they're going to keep building for that 100 % accuracy. Maybe they get there in six months, a year, two years. but in the meantime more and more people are going to be like oh well this is useless because they're going to see wrong things rather than just even it's like if you use google search or wikipedia you learned to use it you learned that if i use google search the best result if i click the first one might not be the best one that's okay i'm going to work my way through the list this is a process so i think this is the smartest point on this is that companies like the way we actually release these products into the wild, there's a big disconnect here right now.
37:15Yep, definitely. And speaking of AI's branding problem, you got to think about Humane, but we'll talk about that in this. Well, it's not the second half because we're in the second half already, but after the break. But before we get to the break, I want to talk about one more piece of the consequences of deep research, which I think the handing tasks off to AI already, there's already some early research that it's atrophying people's brains uh and this is from 404 media it's about a microsoft study microsoft study finds ai makes human cognition atrophied and unprepared a new paper from researchers at microsoft and carnegie mellon university finds that as humans increasingly rely on generative ai in their work they use less critical thinking which can result in the deterioration of cognitive faculties that ought to be preserved So it's a key irony of automation that by mechanizing routine tasks and leaving exception handling to the human user, you deprive the user of the routine opportunities to practice their judgment and strengthen their cognitive musculature, leave them atrophied and unprepared when the expectation is dualitis.
38:22Interesting that it's coming out of Microsoft research, but this makes sense to me, right? Like people say that when you use GPS as opposed to a map, that part of your brain kind of goes away. Lord help me, I couldn't get around with a map today. And so maybe we're going to experience the same thing on a wider scale when it comes to these products like Deep Research or even just the ChatGPTs. What do you think about this, Ranjan? I think it's both correct but not worrying. I think, as you said, GPS. I mean, I honestly have outsourced a part of my brain where I remember where I keep things with AirTags.
38:55And now literally before I would look around for my phone or AirPods or keys and now I go straight to my phone and find the item and beep it. Like I've completely outsourced that part of my brain. So I think there's always good and bad in these kind of things. But I do think, and you have a Paul Graham piece a bit later, I think the process, and it's going back to what I was saying a second ago, the process of researching in itself is valuable. Understanding how to look through what you're presented, if it's 85 % right, is valuable. So I do think that certain parts, but the actual act of having 20 tabs open and having to click on and copy paste into a document, that skill will probably go away.
39:41Maybe that's not the worst thing in the world. So I think this one's okay. Okay, let's hear from Paul Graham being quoted in The Economist, which uses the most economist word ever to start this paragraph. and that is ineluctably? Ineluctably? Ineluctably? Why do people write like this? No one speaks like this. Because you're the economist. No, no, the economist editors, they speak like that. They do. I worked at the Financial Times. I can guarantee you the economist editors speak like that. I'm going to just skip this word. Paul Graham, a Silicon Valley investor, has noted that AI models, by offering to do people's writing for them, risk making them stupid it.
40:25Writing is thinking, he has said. In fact, there's kind of a thinking that can only be done by writing. The same is true for research. For many jobs, researching is thinking. Noticing contradictions and gaps is in the conventional wisdom. The risk of outsourcing all your research to a super genius assistant is that you reduce the number of opportunities to have your best ideas. I completely agree with this about writing in particular. I kind of see your yes and no perspective, but I'm also kind of like probably a bit more concerned than you are because this stuff all seems pretty real to me. I'm not denying that I have zero concern on this, but again, I think that over time people will figure out how to use it.
41:07There's going to be some large number or some number of people that do use this in a lazy way and do lose the ability to think through writing. But I think overall, I actually had the most random deep research also related to last week we were talking about model picking and uh do you do you know what the opposite of a cake donut is there's two main kinds of donuts a is it a potato donut no no no no but is it is a cake donut it is there's cake and you know like the kind of fluffier ones that are uh okay no anyway this came up a bread donut the answer i'm not a huge donut guy but this is It just came up because someone was like, is this a - I'm a little ashamed of myself now, but sorry, yeah, let's hear it.
41:54You can become an expert thanks to deep research. So I was asking, and because I use not Google for most of my searches nowadays, I asked, you know, what is the other type of donut other than a cake donut? I asked perplexity. I accidentally had deep research still clicked. And rather than giving me a simple answer, I received an entire term paper, essentially a PhD paper, The Opposite of Cake Donuts, a comprehensive exploration of yeast raised donuts and their contrasting characteristics. And it goes on and it provides like 6 ,000 words. It's a yeast donut around this. And this was where I was also like, how much random stuff is going to get generated by either people asking dumb questions, asking simple questions and picking the wrong model.
42:43But these research tools are, they're quite something. Okay, I changed my mind. I'm into this. This will expand the human capacity for cognition. I mean, I might, you might become a donut expert now. I've learned more about the opposite defined by contrasting leavening methods, textures, preparation techniques, and culinary applications is unequivocally the yeast raised donut. That's pretty good. I think that these deep research tools are going to be a godsend to stoners everywhere who instead of browsing Wikipedia will start to read long, half true research, deep research entries about lots of weird stuff.
43:27All right. Well, maybe deep research will give stoners something else to do outside of talking with their Amazon Echoes all day long about the meaning of life. And I think they're going to have some competition because the Echo and Alexa are about to get better. We'll talk about that right after the break. Did you know your credit card points and miles can lose value to inflation? Credit card companies often reduce the redemption value of your points and miles. Now, imagine a credit card with rewards that can grow in value. With the Gemini credit card, you can earn Bitcoin or one of over 50 other cryptos instantly with no annual fee.
44:02Every swipe at the store or gas pump earns you instant rewards deposited straight to your account. Plus, sign up now for a$200 Bitcoin bonus to kickstart your rewards. Visit Gemini.com slash card today. Check out the link in the description for more information on rates. Again, if you're looking to invest in Bitcoin but don't know where to start, the Gemini credit card makes it easy. The Gemini credit card is issued by WebBank. In order to To qualify for the$200 crypto intro bonus, you must spend$3 ,000 in your first 90 days. Some exclusions apply to instant rewards in which rewards are deposited when the transaction posts.
44:39This content is not investment advice and trading crypto involves risk. The Gemini credit card cannot be used to make gambling related purchases.
44:51You're used to hearing my voice on the world bringing you interviews from around the globe. And you hear me reporting environment and climate news. I'm Carolyn Beeler. And I'm Marco Werman. We're now with you hosting The World Together. More global journalism with a fresh new sound. Listen to The World on your local public radio station and wherever you find your podcasts.
45:18And we're back here on Big Technology Podcast Friday edition. A couple of news items to hit before we get out of here for the weekend. And next week, Amazon is going to have a big revamp of Alexa that it will be announcing on Wednesday. This is from Reuters. The AI service will be able to respond to multiple prompts in a sequence and, as company executives have said, even act as an agent on behalf of users by taking actions for them without their direct involvement. This contrasts with the current iteration, which generally handles only a single request at a time. So it looks like at long last, we're going to see an Alexa update.
46:00I am planning to be in attendance, and I'm hoping that we'll be podcasting in or around the event. So stay tuned for that. But I'm personally looking forward to this. I have three Echoes in my apartment here in New York, and I really want them to work well, to work at all. to even just play some music when I ask them to do it, where sometimes they do and sometimes they don't. Am I getting over my skis being excited about the Amazon Alexa overall, Ranjan? And what are you looking out for here? Well, I don't want to rain on your parade as a Apple Home user with multiple HomePods everywhere who just gets more and more disappointed by Siri in every new release.
46:44But I do, I want to root for you. I want to believe that Amazon is going to figure out. I do think, and I'm almost confused how good voice has become. Gemini on the Gemini app, ChatGPT advanced voice mode. They're so good that there's no reason these voice assistants should not be good. I don't understand. Like, I'm genuinely baffled why even the latency with Gemini and stuff like that, like it's really good so i've been very confused as to what the hold up with all these companies are and honestly amazon brought us voice assistance alexa was one of the most revolutionary devices i think of the 2010s so i think if anyone can do it i think they have a shot well it's coming and it looks like it's going to potentially be a subscription product this is all speculation but the post says amazon previously planned to launch the improved alexa with the free trial period after which customers would have to pay a subscription fee.
47:49It is going to be delayed, the post says, so they're going to announce it on the 26th, but we might not see it until the end of March. And the opportunity is massive for Amazon here. Also from the post, Alexa is free on the more than 500 million Echo devices or Alexa enabled devices around the world that can play music, dim lights, and read the headlines. So if they get this right, massive opportunity. But if they don't, just be more disappointment. But the cool thing about having a piece of hardware that is software enabled is you can update it at any time and take something that's mildly disappointing and make it shockingly beautiful and wonderful.
48:31And so maybe that is what Amazon will do. The thing that would worry me here is, And I'm curious if you have use cases beyond, I don't know if you have it turn off and on lights, but ask the weather, ask sports scores, maybe ask some like recipe information. Like, do you have any more advanced use cases currently? I used to do lights. I don't do lights anymore. I think it stopped working as well as I wanted it to and sort of gave up on the whole smart home concept. I do music a lot. I do timers. but every now and again when my wife and I have an argument about something I will summon Alexa and say what's the answer here and it does get it right every now and again and I just think that that is going to be like one of the cool use cases if you have it you know be somewhat conversational start conversations facilitate things and be able to answer accurately and sort of actually get the context of your question when you ask something.
49:32Do you know what I will never do? Do you know what I will never do? While in an argument with my wife, turn to my smart, my voice assistant and ask who is right. Because does that, does that really go well? It's a huge miss. It's a huge miss. Honestly, honestly, it really does diffuse some of the tension because you're just like, well, at least we're not as dumb as this echo here. And then, you know. All right. All right. But if it gets smart. Peace and love again. What if it gets smart and starts, Alex, you're wrong. You're wrong. You're like, is that okay? Look, then I'm owned. I have to admit it.
50:09That's true. Two against one. Yeah. I just think, I mean, this is not marriage advice or relationship advice, but I think everybody out there, if you're trying to find ways to bring common ground with those that you love, I think that you might want to just try bringing a voice AI into the conversation and seeing what happens. Maybe it will be great. Maybe it will be terrible. But at least you tried. And that's what matters. If any listeners do go down this path, write in and let us know how it went. You can find us at ronjohnroy at gmail.com. I don't even know if that's your email address. But yeah, please don't do this in any high stakes moment.
50:58No. Okay. No. So that will not be our fault. Speaking of high stakes. Just keep it in the low stakes argument. You know, what's the capital of Greece? When you say Athens, she says Sparta. And then you call in your voice assistant to tell you the capital of Macedonia. It's just how it goes. Okay, that is fair. I'll give you, if it's purely information or factual, maybe. But if it's like... Can you believe that? Full blown. You never pick up the kids. We're always at your parents. Alexa, who's right? Who's right? I think that Alex is wrong. Well, I've been tracking both of your movement regularly for the last two years, and I can tell you exactly who's right.
51:39Yeah. Honestly, it's true innovation that we need. It's going to herald the new American century and, once again, take the AI discussion away from model and towards product. See, just trying to help you, Ranjan. It's the product. All right. A few minutes on the Humane pin before we leave. The Humane pin is dead. If you all remember, it was this pin that you would wear on your shirt, basically, and you could touch it and could give you AI stuff. And it would project stuff on your hand when you needed some information. HP will acquire the assets from Humane, the maker of the wearable AI pin introduced in late 2023 for$116 million.
52:19The deal will include the majority of Humane's employees in addition to its software platform and intellectual property. however it will not include umane's ai pin which will be wound down i saw a great meme about this i should drop it in the discord meme section about someone uh hitting the umane pin and projecting on their hand that their printer is low on toner toner um but um it is really a by the way they raised 230 million dollars so it really is a ignominious uh end is that the right way to pronounce it for one of the ai era's worst conceived and worst marketed products rest in peace humane pin we will not miss you we will remember you with the likes of quibi and quickster and other terrible failures that began with q but you began with an h humane pin hp hp now stands for humane pin that's the rules i don't make them up that's the rules i think it's not the worst outcome again And as M.G.
53:21Siegler put out, they're selling for$116 million. So for something that feels like this just dramatic a failure, they somehow got something out of it, at least to initial investors. But I was thinking about, remember, we came up with the concept of pin casts instead of podcasts specific to the pin? That's what I'm a little disappointed at, where you're like, it's some kind of interactive podcast that you're listening to, and you're talking to your pin and it's projecting some information. I'm a little disappointed. I did not bet it all on making pin casts. That's right. We should have done that, but we didn't.
54:01And we could have really, I think, saved Humane with either pin casts or just imagine you're in a dispute with a loved one. It all comes back to solving arguments. Things are going rough, like real rough. And you're like, sorry, honey, let's bring in our pin. Let's bring in the pin. You tap your pin twice and then it settles the matter for you. And it's right above my heart as well. You ask who is right and it projects the answer on your hand in green. Let's read M.G. Siegler just before we go. He goes, a regular person might read that headline that the company sold for$130 million and think, wow, a startup sold for nine figures.
54:44Impressive. Of course, it's not impressive in this use case. It's a fire sale for a company that has been under duress for months after their product, the AI pin, failed to catch fire in the market. Actually, that's not technically true. There was a literal risk of fire when charging the device, which led to a recall. I completely forgot the UMaine pin lights on fire side of this recall. Never a good thing. Never good. Of this story, but apparently that's what happened. I do think, I do wonder these kind of things. If they weren't, if those launch videos weren't so bad, did it stand a chance? Like sometimes you wonder like the, was it really the technology?
55:25Also, it was released at a point where voice interaction with generative AI was, it was again, it was only two years ago, year and a half ago. The voice mode and voice interaction have like exponentially jumped in quality. So it's a timing issue of one. but the marketing and those launch videos will live in history as some of the worst i've ever seen and did could this have stood a chance if their launch videos weren't just memed into oblivion yeah that's what mg siegler is basically saying here he says the pr strategy was obviously a disaster from the get-go this is easy to say in hindsight but many people were saying it in real time from their grandiose but nebulous introduction video to their first product tease on stage at TED, it was seemingly less of a sound strategy than a vignette of cliches.
56:17And by God, we've heard a lot of vignettes of cliches lately. This is me, not MG. But he says it also includes a$699 price point. That was obviously never going to work for this product. And on top of that, a monthly subscription fee. Everyone was like, oh, that might be cool. Maybe we weren't. But just the dynamics of this business were never going to work. And that kind of leaves me to our final question of the week, which is there is something to be said of like, you kind of got to hand it to these people for actually going ahead and building. But the other side of that is like, I don't know, do you have to applaud everybody that tries or can you just say that sucked and I'm disappointed that happened and I'm kind of in the second category on this one?
56:59Well, I think I have a lot more sympathy or empathy for people who weren't able to raise $240 million before even launching the product. And then the production value and just knowing how expensive those launch videos must have been to make and how much they must have spent on the agencies. So no, I don't have sympathy for them. I think like you could just come on, you can make a good engaging video pretty easily nowadays. You should have made one. And maybe we would all be wearing our pins as we argue with loved ones and listen to our pin casts. And the world would actually be okay right now.
57:41Everything would have been okay if Humane didn't screw it up. They should have called in Cantroitz and Roy. We would have told them, got to show the use case, got to show the product, and got to show the humanity of your pin. And instead, they burned through$240 million. Could have saved democracy. could have saved it all democracy yeah instead there's no more pin it's now it's now with the printers that is the the ultimate graveyard of innovation i think that we've ever you can one can ever imagine hewlett packers printers division there goes uh any chance that hp will ever sponsor this podcast but it was worth it so thank you thank you for that ron john but if you're listening out there hp um good luck with humane maybe you'll do something cool and we're waiting for it bring Bring back the pin.
58:29Bring back the pin. In the meantime, we're wishing you all love and success and happiness and tranquility at home. And remember, if anything goes wrong, just ask Alexa. Ron John, great to see you as always. Thanks for coming on the show. All right. Hopefully my marriage survives the weekend and I don't ask Alexa to resolve anything. Okay, everybody. Thank you for listening. Chris Hayes is coming on next week. We actually have a very lively and fun conversation about how social media and publishing differs and all this other stuff that he's been talking about on his tour about the Sirens College is his number one best-selling book so hope you stay tuned for that thank you listening thank you to Ranjan and we'll see you next time on Big Technology Podcast
From the publisher
Ranjan Roy from Margins is back for our weekly discussion of the latest tech news. We cover 1) Satya Nadella's criticism of AI benchmark hacking 2) Ex-OpenAI CTO Mira Murati's new Thinking Machines Lab startup 3) There are too many AI startups 4) Why foundation models have commoditized 5) Did Google 'DeepSeek' itself? 6) Grok3 arrives 7) How do you evaluate whether models are good? 8) Grok3 at the top of Chatbot arena 9) Benedict Evans on Deep Research 10) Does using AI tools make our brains atrophy? 11) Amazon's incoming Alexa upgrade 12) Actually, voice AI helps during marital disputes 13) RIP Humane Pin
Join the Big Technology Discord here: https://www.bigtechnology.com/p/lets-talk-deepseek-ai-etc-on-big


