In short
a16z Podcast Episode Summary: Dylan Patel: GPT-5, NVIDIA, Intel, Meta, Apple
Episode Overview In this episode, host Erik Torenberg is joined by Dylan Patel (Founder & CEO, SemiAnalysis), Erin Price-Wright (General Partner, a16z), and Guido Appenzeller (Partner, a16z) to discuss the rapidly evolving landscape of AI hardware, focusing on the competition in the chip market, particularly in relation to NVIDIA's dominance.
Key Discussion Points
- State of AI Chips and Competition
- NVIDIA's Dominance: NVIDIA remains the leader in AI hardware, making it difficult for competitors to catch up.
- Challenges for Competitors: Simply copying NVIDIA's approach is insufficient; new players need innovative strategies that leapfrog existing technologies, aiming for a significant competitive advantage (5x better).
- Custom Silicon and Market Dynamics
- Custom Silicon from Big Tech: Companies like Google, Amazon, and Meta are investing heavily in custom silicon, which could potentially reshape the market dynamics.
- AI Model Launch Economics: There's a significant shift towards cost efficiency in AI model launches, influenced by the high costs associated with running AI models.
- Infrastructure Bottlenecks
- Power and Cooling: Current bottlenecks in AI infrastructure include power supply issues and cooling requirements, critical for data center operations.
- Global Supply Chain: The challenges in the global supply chain affect the availability and cost of necessary components.
- The Rise of AI Silicon Startups
- Funding Challenges: Many AI silicon startups are attracting significant capital despite not having launched products, highlighting investor confidence in AI technology's future.
- Competition with Established Players: New startups face challenges in competing with established entities like NVIDIA and need to find a way to differentiate themselves effectively.
- Geopolitical Insights
- Export Controls and Geopolitical Tensions: Discussions around China’s ambitions in AI and how U.S. export controls are impacting competition and innovation in the AI sector.
- Big Tech's Future Strategies
- Advice for Tech Leaders: The episode discusses strategic advice for leaders like Jensen Huang (NVIDIA), Sundar Pichai (Google), Mark Zuckerberg (Meta), and Elon Musk (X), focusing on areas of investment and innovation.
- Monetizing Free Users: The conversation touches on the importance of finding sustainable monetization strategies for products that attract large user bases without direct revenue models.
- GPT-5 Insights
- Performance Evaluation: Dylan Patel shares his analysis of GPT-5, discussing its performance relative to previous models and how it optimizes computational efficiency.
- Market Reactions: The impact of GPT-5's features and pricing strategy on user experience, particularly for free users versus paying subscribers.
Key Takeaways
- AI Infrastructure is Critical: As the demand for AI hardware grows, being able to effectively manage power supply and cooling solutions is becoming increasingly crucial.
- The Competitive Landscape is Shifting: Companies must innovate beyond existing frameworks to compete with NVIDIA effectively.
- Monetization Strategies Must Evolve: The challenge of monetizing free users in AI applications remains a significant issue for tech companies.
- Geopolitical Considerations are Vital: Understanding the global implications of AI technology and chip production is essential for navigating the competitive landscape.
Conclusion This episode of the a16z Podcast provides valuable insights into the current state of AI hardware and the competitive landscape, emphasizing innovation, strategic investment, and the importance of infrastructure in sustaining growth in the AI sector. The discussion highlights both the challenges and opportunities facing tech companies as they navigate this rapidly evolving field.
For more details and to listen to the episode, visit [a16z.com](https://a16z.com).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00In videos going to have better networking than you, they're going to have better HBM, they're going to have better process node, they're going to come to market faster, they're going to be able to ramp faster, they're going to have better negotiations with whether it's TSMC or SK high nicks in the memory and silicon side or all the rack people or like copper cables, everything they're going to have better cost efficiency. So you can't just like do the same thing as Nvidia. You have to really leap forward in some other way. You have to be like 5x better. Today we're talking AI, hardware, chips in the infrastructure powering the next wave of models with three people at the center of it all.
0:31Dylan Patel, founder and CEO of Semiann Allisus, one of the sharpest voices on chips, data centers, and the economics driving AI's explosive growth. Aaron Price -Rite, General Partner at A16Z, investing in the technologies and infrastructure shaping the future. Widow Appenzeller, Partner at A16Z, with decades on the front lines of AI, Cloud, and networking. From GPT -5's launch to Nvidia's Dolmins, Custom Silicon, and the Global Race for Compute, We're covering what's happening behind the scenes. Let's get it to it Dylan welcome to the podcast. Thank you for having me. We've been trying to get you for a while You're busy man, but it worked out.
1:09We don't want to introduce why we're so excited to have Dylan and podcast and what we're excited to discuss I think Dylan you've done except no drop in and covering what's happening in the AI hardware space AI semi space and now more and more data center space as well and just Just looking at it, currently the most valuable company on the planet is an AI semi company, right? I think biggest IPO so far in AI was an AI cloud company. This is currently where it's happening. In any gold rush in the early days, it's the peaks and troubles that make money. I think this is the stage that we're in. So it's super excited to have you here today.
1:41Awesome. Thank you. Happy to talk about my favorite topics. Amazing. Maybe let's start with GB5. We just had some of the research of Christina and Isabella on here last week. You said it was disappointing. One, you share your reactions or what capabilities you are hoping to see or overall. I think it depends on what tier of user you are. Right. If you're just using GPT -5 and before you were $20 or $200 a month, subscriber, you no longer have access to 4 .5, which I, in my opinion, is still a better pre -trained model for certain things. Or you no longer have access to O3, which would think for 30 seconds on average, maybe, right?
2:17Whereas GPT -5, even when you're using thinking, only thinks for like five to 10 seconds on average, right, which is an interesting sort of phenomenon, right? But basically like GP -D5 is not spending more compute per se. The model did get a little bit better on a vanilla basis, right? 4 .0 to 5 is actually quite a bit better. But when you think about, you know, what is this curve of intelligence, right? It's like the more compute you spend, the better the model gets. And that's whether it's a bigger model, which GP -D5 isn't, right? You can see it's not a bigger model. It's roughly the same size, you know, or you think more, right?
2:50But again, And this is something that OpenAI's first thinking models, first few generations of 0103 would think for a long time and waste a lot of tokens, if you will. And when you look at, for example, inthropics thinking models, even when you put them in thinking mode, they think a lot less, right? To get to the same results are better results, right? As OpenAI was. And so OpenAI, I think, like, optimized a lot of, like, well, if I ask, like, I think the silliest one I had asked was like, I asked 03 once, is pork red meat or white meat? And it thought for like 48 seconds, it's like, what are you doing?
3:20Like, this should just like tell me the answer. And so like, the nice thing is that GPT -5 will think a lot less, even if you select thinking manually, but more importantly, they have the sort of auto functionality, the router, which lets them decide whether or not, hey, do I route to the regular model, do I route to maybe many, if you're at a rate limits, or do I route to thinking, right? And how much do I think? But in general, the thinking model will think less. So there's less compute going into a power user's average query than before. But is it even more interesting? Open ARC cannot control how much computer wants to allocate to you, right?
3:55If we're in a high load situation, maybe tune the router a little bit so it's less, right? Maybe we've got... I have no idea what they're doing behind the curtain, but there's this meme out there at the moment that basically all they did, which is a meme, right? It's not true, but all they did is take all three plus a couple of smaller models, put a router in front and offer that at a lower -blanded price, essentially, right? I think there's a little bit of that. I cost suddenly matters and they figured out a way how they can steer that. I think, yeah, I mean, and they talked about how they've been able to as dramatic they increase their infrastructure capacity.
4:25Because I myself was just regularly using O3 or 4 .5, right? And now I'm forced to use auto, which sometimes gives me the O3 equivalent thinking model, but sometimes gives me the regular base also, which sucks. But like, I think for the free user, it's actually quite interesting, right? The free user was not getting thinking models pretty much ever or not using them or in many cases they just opened the website and asked their query and now sometimes their query gets routed there so sometimes they get away better model but now sometimes the opening I can gracefully degrade that if they need to right and I think the router points to the future of opening I from a business right like you can look at Sort of the model companies right and Thropic is fully focused on B2B right API code etc right or a cloud code whatever it is, right?
5:09Open AI, yes, they have that business, codex and API business, but really they're majority of the revenue is consumer, right? And it's consumer subscriptions, but they have no way to upsell, you know, to make money off of all the free users, right? In any other application consumer app, the free user still pays via ads. But this is not compatible with AI, right? Like, it's a helpful assistant. You can't just make the users with result worse by injecting ads. Vanner ads don't really work in AI either. So it's like, how do you now monetize them? And I think with the router, they're getting really close to figuring out how to monetize that user, right?
5:45With the new CEO of applications, if you saw her product that she launched at Shopify, I think it was Shopify was an agent for shopping, right? And now this like immediately clicks, like, oh, if the user asks a low value query, hey, why is the sky blue? Just route them to many, right? The model can answer perfectly fine. And that is a chunk of queries, right? But if they ask, what's the best DUI lawyer near me? Right? All of a sudden, this is like, you're in jail, you have one shot, you're like, screw it, let me ask Ted GPD what the best DUI lawyer is. And now all of a sudden, the model's not capable of it today, but soon enough, it'll be able to contact all the lawyers in the area and figure out what their results are and maybe search their like court filings and whatever, right?
6:24Book the best lawyer for you. Or an airplane ticket. Maybe you go, Shane, a cot, as part of that. Yeah, of course they're going to take a cut, right? But this is a much better way of monetizing the free user is like, you know, it's like at sea 10 % of their traffic now comes from chat and OpenMix nothing off of that, but they really really will soon, right? And partially that's because Amazon blocks chat, but there's a way to make money from shopping decisions whether it's booking flights or looking for items and those you now say free user. I don't care. I'm gonna send you to my best model I'm gonna send you to agents.
6:53I'm gonna spend and ungodly amounts of compute on you because I can make money off of this. But if it's a query that's like, help me with my homework, I'll send you like a decent model, right? I don't need to spend money on you. And so this is how I think like, opening I can finally make money off of the free user. And I think that's the biggest like, think about the router, right? This is so interesting. I think this is the first time that we've seen that there's a launch of a new model where to some degree, cost is the headline item, right? I mean, just so far, I was always like, who is the smartest model?
7:22Who is the highest ML or US core? Now, we have suddenly people who use models for coding for eight hours a day and surprise that if you take a large context window and the best model creates thousands of dollars of cost a month. So cost matters. And so, in some degree, so far, you're on the parade of frontier between cost and performance is the new benchmark for model competitive and no longer cost alone. Is that what we're seeing here? I mean, I think definitely, right? Like, opening eyes said they doubled their rate limits for big amounts of users. They've dramatically increased the number of tokens they're serving from this launch, which effectively says this is an economic release.
7:54Probably also means that tokens are not cheaper. Yeah, for sure, for sure. I think the funniest thing is this whole cost thing you mentioned is like, we've seen this in the code space, right? Kersher had to pull away the unlimited, clog code. Initially they have this super expensive plan and it had like unlimited rates and then they were only like a weekly rate limit. Now they have like our based rate limits and I saw the craziest like thread on Twitter where this guy said he changed his sleep schedule, right? modeled after like how sailors in the bay. If you're sailing, you can't sleep, right?
8:24Like it's solo sailing. They'll take like power maps when they get to the right spots so that they can like still be safe. In the morning when it's not very windy. Well, but like they can't sleep uninterrupted, right? And so because Anthropic had to put rate limits that are like not just weak base, but like a number of hours based. And like he like basically sleeps multiple times a day, but small chunks just so you can maximize the usage. And there's also a leaderboard on Reddit where people are like competing to see how many tokens they're using through their subscription. And there's like a dude spending like $30 ,000 a month.
8:56So I'm gonna find some developer in India that I can do a pair of programming with. So I can get the day cycle, he can get the night cycle and we both can maximize the together the quota for the account, is that the future then? I mean, but it's clear like people are taking advantage of the negative gross margin, like sort of subscriptions that are offered. I think Anthropic probably makes a positive gross margin off of my subscription, I don't know how to code enough. But there's plenty of people that are definitely losing money. And so as you said, it's an economic. It puts more and more to, I think, just usage -based pricing, right?
9:22I think if you have an underlying commodity that you're reselling to some degree that has that is that larger part of your cost of goods, right? You need to go to use a space pricing. How much do you think the like customer capture and stickiness for these code products is? I'm curious what you think on that, right? Once you use an ID, once you integrate one of the CLI products in, like, how sticky is it? Or is it a billion dollar question? There's a very conservative estimate. Look, Andrew Parthy has this great slide where he basically says, if you're building an agentic system today, right? But fundamentally what it is to solve this loop, right?
9:54Where half of the loop is the model thinking, right? And I'm trying to do some. The other half is then the user verifying what did the agent do? Is it the right thing providing feedback and trying to steer it in the right direction? Because we can't run forever. I eventually need to steer it back. One half of that is the model provider, right? They're trying to build the best models. The other half is really about, I think, designing the best possible UI to enable the user to get feedback. And I think there's value on that. So I think there's a certain amount of stickiness in there. Right? So what are all the different tools like in terms of like say take hold editing, right?
10:20How can a most easily visualize what the cold changes are? How can it most easily visualize, you know, what they impact which files? You know, how can I for small changes get very quick feedback versus for complex ones, you don't get complex feedbacks? And that's not some tools that actually draw diagrams for you for what they do. Right? So I think this will be the battle is I think there's stickiness in that. Right. How much exactly? So in that sense, like people should be doing subscriptions to get people locked in, right? Instead of moving to usage -based pricing. Well, I think it's the customers that don't want to do usage -based pricing because it's so hard to guarantee, it's so hard for it to get away from them.
10:52You actually want guarantees and you're willing to commit to pretty high spend in order to not have usage -based pricing. I think it's the model companies that want usage -based pricing. I think with consumers, it's frankly very hard to not have usage -based pricing. It's because the variability is so massive. it's us coding versus somebody who does their full -time job, right? You just have a factor of 20 or so difference in usage. That costs a lot of money, right? I think it's just one. I think for enterprises, we could see that like more flat fee pricing was an average at all. You have a developer that's using it all day.
11:23You kind of know in general sense how many hours a day they're programming and what that sort of looks like. The five quotas are harder. Before we leave opening, I want to ask a broad question, which is if someone was sitting here and saying, hey, Dylan, I'll listen to anything you tell me to do. any advice you have as long as it makes OpenAm more valuable. Who would you tell them? Oh, it's a immediately launched a method for you to input your credit card into chat GPT and agree that for anything it like agentically does for you, it'll take X cut and then launch that product because where it does shopping, right?
11:55Because like, everyone knows that like Anthropic and OpenAI and all the other labs are buying oral environments of Amazon and of Shopify and of Etsy and of all the different ways to shop on the internet, oh, of airline websites, right? Now, just like, hey, integrate my calendar, I want to fly to there on Thursday, make sure I don't miss a meeting cool book, right? Do that integration, like, super well, know my preferences on whether I like I other window, all this stuff, right? And just take a take, right? I think this will make them so much money the moment they launch it. And I think they're working on it already.
12:26But I'd like to hear how he thinks about it, because he shifted his tone massively on like ads over the last six months, right? He used to be like no way. And now he's like, maybe, you know, there's a way to do it without harming the user. And I think this is how you monetize the free user, right? So I think that's probably what I tell him slash ask him about like a whole line of questions around this. Well, he's coming on the podcast in a few weeks. So I want to shift to Nvidia. I mean, it is having a monster year. They're up almost 70%. What are the possible paths from here? I do see it playing out.
12:58Depends like how pill do you are on like the continued growth. But I think you guys have a good vantage point. We have a good vantage point of how fast revenue is growing for a lot of these companies, especially the code companies, but even many other applications. I think we can clearly see the demand side is accelerating, right? And then if you look at the training side, I think the race is on. That is upping hugely. Google's upping hugely. If you just look at, again, just open -air an Anthropic and the compute that they have and are getting this year from Google and Amazon for Anthropic and from Microsoft, soft core weave oracle for open AI, 30 % of the chips are going to them, just those two companies.
13:36But that's actually like, okay, well, like 70 % of the stuff, like who's making off of it? Well, one third of it is like ads, right? Whether it be bike dance or meta or many of the other people who are doing ads. So then it's still like, okay, well, where are the rest of these one third of the chips coming from? Well, they're like mostly uneconomic providers who I don't think it's like an obvious bet that they're going to keep raising bigger and bigger around. So what happens there? I think with the, you know, we talked about like coding, right? Like earlier, actually the Qen Coder 3 model is actually super cheap if you're running it on PrAM or if you're running in the cloud with all these inference libraries.
14:10And so like there's stuff like that as well. So I think the question is like, how much does it keep growing? Because clearly, I think the first third is definitely skyrocketing, right? Of open AI and theropic lab spend. The second third of like ads is going to grow. It's not going to grow like crazy. But I think there's definitely an inflection point that could be hit with Jenny. I know Met has been experimenting with it a lot, but I could totally be convinced that there's gonna be a huge inflection and take right there where you start showing me personalized ads. Like every person that's an ad is like, looks like me and I'll be like, okay, yes.
14:39Except like slightly better, so I like feel better. Right? And I'm like, oh, yeah. I've no idea how this is gonna scale, right? But if you ask the question, how much could it scale? Right? How much value will be creating here? Is it, can we create enough value to actually keep growing for a long time? If you just take AI software development, right? Yeah. We know we can easily get about 15 % more productivity out of it. I don't think that's right. I think it's way higher. No, no, but with the straight, like I talked a lot of enterprises, like a classical enterprise straight up GitHub Co -Pilot deployment that gives you about 15%.
15:08We can do much more than that. But bro, like you know how bad GitHub Co -Pilot is, like how did they how good they look at the revenue error? This is so funny. It's so funny. If you look at the revenue error chart, it's like cloud code, it's three months has surpassed that cursor, you know, easily surpass them. And then like even like companies like Replator like at wind surf slash cognition are like going to pass them like it's like, you're pretty good. That's going on. So look, let's assume we can get this 100%. Yeah. So as we can double the productivity of a developer, right? Well, 30 million dollars worldwide give or take, right?
15:39Let's say 100k value ad per developer. It might be the little high worldwide US is low, but one one is high. Just $3 trillion. Yeah. Yeah. Right. So so we're probably building technology here, which adds $3 trillion of GDP value in theory. We could put that into GPUs because that's the main cost. Just from a coding model. Just from a coding model. It's already every other use case. So at least in theory, the value generation is here to keep growing. Right? Now how that translates to the industry is a much more public case. I think we've already seen AI's value creation. So sort of there's like the whole like the famous like oh, 300 billion problem, our 200 billion problem now, 600 billion problem.
16:13I'm sure so quiz going to put out like a one of the So you know how I'm right soon enough, but like, like, there is some like reality in that, of course, but you know, it ignores that like infrastructure spend today is accounted for five years of revenue, not like one and the revenue looks like this, not like flatline. But I think, um, I think the main thing is that AI is already generating more value than the spend. It's that the value captures broken. Right. Like I legitimately believe open AI has not even capturing 10 % of the value they've created in the world already. just by usage of chat, right?
16:47And I think the same applies to, you know, anthropic and cursor and whoever else you're looking at, I think the value capture is really broken. Even internally, I think, what we've been able to do with four devs in terms of automation, our spend on Gemini API is absurdly low, and yet we go through every single permit and regulatory filing around every single data center with AI. And we take satellite photos of every data center, and we were able to label our data set and then recognize what generators people are using, what cooling towers and the construction progress and substations, all this stuff is automated and it's only possible because of Gen AI, but when we do it with very few developers and then the value capture that I'm able to generate by selling this data, by consulting with it is so high, but the company is making it, it's like they get nothing out of it, right?
17:39I think there was a value capture challenge here that far out exceeds the sort of creation, right? And as you get models like GPT -5 or open source models, like continuing to drive it down, it's like the value capture is just harder and harder and harder for these companies because they're making 50 % gross margin on inference if they're, you know, or less in many cases. And so many words you're saying, we're getting commoditized and therefore you can't capture the value and thus you should tempt your expectations of how much, how much you can spend in GPUs. Well, no, I think I think you can, I think there's still ways to like, inflect hugely on value capture, right?
18:17But I mentioned that ads are a huge value capture. But that needs to happen before, before we see a massive increase. No, I think the other thing is like, there's a lot of capital that's not been spent, right? Like the hyper scaler still can grow CapEx 20, 30 % next year, right? From what they're doing this year. In addition, companies like CoreWeave and Oracle, because they're tapping capital markets can raise is way more than 20 to 30 % cat -backs. And then you go down the list further and it's like, oh, the largest infrastructure funds in the world, like Brookfield and Blackstone. Well, actually, they're turning all of their eyes to investing even more into infrastructure, AI and FRA.
18:54And then you're like the sovereign wealth funds of the world, like the G42s or the Norway one or GIC and Singapore, like these people have barely started touching AI. And so I think there's a whole lot more cat -backs that can come without it being necessarily economically motivated day one. I'm also saying economically motivated catbecks can only grow so much, but there's so much other where it's not clear from, if you have a spreadsheet and you're basically on real business that you should actually spend this much. But people will, because they believe, I believe, I think you believe, in for people believe that this will be, you'll get profit out of it, but there's no like 100 % certain like, you know, way to argue it.
19:41Yeah. How threatened is Nvidia by my custom silicon? I think that's the biggest thing, right? Is when we look at orders from Google and from Amazon, right? Especially their and meta, their custom silicon is not not Microsoft, their custom silicon kind of sucks. But the other three, they're really upping their orders massively over the last year. You know, Amazon is making millions of of train -yum, Google's making millions of TPUs. TPUs clearly are like 100 % utilized, right? Yeah. That's right. I'm training them's not there, but I think Amazon will figure out how to do that and then Thropic will.
20:22So I think that's the biggest threat to Nvidia is that people figure out how to use custom silicon more broadly. And this sort of becomes the sort of like, if AI is concentrated, then custom silicon will do better. And that's not even talking about like, open AI is silicon team and stuff, right? Like if AI is really concentrated, then they'll do better, custom silicon. But if it gets dispersed broadly because there's all these open source models from China, and there's all these open source software libraries from Nvidia and China. And it makes the deployment costs like rock bottom than potentially.
20:58Him and I, if Google's TPU is able to compete with Nvidia, it can theory could do it on the open market. And video is worth more than Google these days. Shouldn't Google start selling their chips to everyone? I mean, theory, they should be able to achieve a higher market cap. I absolutely think so. I think Google is even discussing it internally. I think it would require a big reorgive culture and a big reorgive like how Google Cloud works and how the TPU team works and how the jack software team and XLA software teams work. I totally think they could. It would just take them like shaking themselves pretty hard to be able to do it.
21:38Yeah, but I totally think Google should sell TPUs externally, not just renting, but like physically. It's kind of funny. If a side hobby in theory has a higher company value of potential as you make it. Especially if you think about the degradation of search. So I think business, I mean, yeah, I think, but I think like if you were to ask like surgate, right, like, hey, do you think selling chips and in -friend racks is more valuable or cloud or or Gemini, um, he'd be like, no, no, no, no, like Gemini is going to be worth way, way, way more. It's just not yet today, right? Um, and so I think like, like today, you say Nvidia is the most again, it's like a whole concentration thing, right?
22:19The role is super concentrated in terms of customers, then Nvidia will not be the most valuable company in the world, right? But if it gets dispersed more and more, which arguably we're starting to see with a lot of these open source models getting better and better and better and with ease of deploying them getting better, then you would see, I think you could argue in video, will remain the most valuable company in the world for a long period of time. Historically, no pun intended software has eaten the world in most markets, right? I mean, like if you look at early networking days, Cisco was the most valuable company on the planet, right?
Read the full transcript
22:55For a while. It's no longer, right? The guys that build services on top, like Google or Amazon or Meta eventually eclipsed. Which is why Nvidia is making all the software libraries, right? And they're trying to commoditize inference, right? You guys don't, I think, even have an inference API provider investment, do you? Well, with all kinds of model providers. Model -Coviders. I'm talking about a pure API provider investment. I think, right, is that correct? I think I talked to one of the team members, maybe Rajko or someone about why you guys didn't invest in a fireworks. And the argument was, well, we think just serving models alone without making them will sort of be commoditized.
23:38Yeah. We have some in the CBTF system, with like a file. Without a decay. Yeah, it's a little bit different dynamics there, I think. They tend to make much more component models than the OM folks. I think they're different. Yeah, absolutely. But you guys don't have one of these like, you know, base 10 or any of these like sort of like API investments because you think this is from someone on the infrared team that you guys think it'll get commoditized because of software and video making because VLM and S .D. Laying, which is like open source software coming out of Berkeley and now sort of has their own environments now and supported by many like this being commoditized means that like API the I providers aren't necessarily worth a ton, right?
24:18Is sort of your argument, maybe. I think that's relevant to this whole thing, which is, you know, why, right? Like, why would you do this? Shifting gears. What about the Silicon startups? What's your take on those? I mean, there's a ton of capital flowing into that, where I've seen, I have not a number, it's probably billions being invested in and ship startups. Yeah, for sure, for sure. I mean, like, whether you're looking at like, you know, companies like, I think it's pretty impressive that a few companies like Etch and Revos and a number of other companies, you know, Maddox and others have gotten the amount of funding they've had without even launching a chip, right?
24:58You know, in the past, like, yes, Silicon companies would make money or raise money, but they would at least launch a chip before they get a, you know, a big round, but like Etch and Revos have raised, you know, a lot of money without ever launching a chip publicly, which I think is, I mean, it speaks to, Well, yes, Silicon is super capital intensive if you're building a chip, especially an accelerator, which has so many moving pieces. And there's like 10 different AI accelerator companies out there, right? That are new -ish in the last few years. It gives a lot more. That are like, yeah, that's fair.
25:32And then there's the old guard which continues to raise money, right? Like Grockens, Rebrests, and Sominova, and Tens Torren, and so on and so forth, right? like, or Graph Quirking bought out by Softbank and Softbank dumping money into this effort as well, right? There's a lot of capital being invested to disperse, to spell sort of Nvidia's top dollar or top position. But it becomes challenging, right? It's like, how do you beat Nvidia, right? Like, the hyper scalers, I think, are like kind of lucky in that. They can do mostly the same thing as Nvidia, right? They've got a couple of customers, which is themselves.
26:07Sorry, that's a huge estimate. And it's, they can just win on supply chain, right? I'm using cheaper providers. It's a margin compression exercise, essentially. Yeah, yeah. And maybe for certain workloads, like metaphor recommendation systems, they'll have a better, they can specialize more. But for the most part, it's like, no, we're targeting the same workloads. We can just simplify supply chain or in house a lot of it and compress margin. And it'll be fine. But in the case of these other companies, it's like, well, they don't have a captive customer. So now you have to contend with, well, I'm using the same ecosystem.
26:39them. And either I can use some custom silicon provider who's going to take a margin anyways on top and that's going to compress my like what I can sell for or I can try and in house everything. But then it's like, this is really hard, right? Like I'm going to do all the software design. I'm going to do all the silicon design. I'm going to build all this different IP. I'm going to manage the supply chain on chips on racks on everything, right? Ends up being a huge effort in terms of team size, all in the end, like, hey, I make a 75 % gross margin as Nvidia, AMD sells their GPUs for 50 % gross margin, and they have a hard time out engineering Nvidia, and they're great at engineering, right?
27:19But yet, they still take more silicon area, more memory to achieve the same performance, and they have to sell for less, so their margin gets compressed. That makes sense. I think, historically, if you look at it, But typically, if a new entrance in markets didn't win by marginally improving on something existing, they happen sometimes. But more likely, they jumped up some kind of disruptive technology leap. Right? Whereas we have a different approach, we have different technology. Is that possible here? I mean, to some degree, maybe there's over -fantons simplifying a little bit, but I think part of the reason why the transformer model won was because it runs so incredibly great on GPUs.
27:54Like a recurrent neural network is similarly performing it looks like, but it runs terribly on now on a GPU. So did we sort of pick the model for an architecture? And now it's hard to come up with an architecture that, you know, really... Well, it's hardware software's design, right? Like, there's all this hype about neuro -morphic computing, right? Like, theoretically, it's amazing and super efficient. It's like, okay, great. Like, there's no ecosystem of hardware. There's no ecosystem of software. It would take like, you know, tens of thousands of people who are the best AI today, focusing on that to even prove out if it's worthwhile or not, right?
28:27on a hardware side, on a software side, on a model side. And so like you look at like Grox or Rebrace, Samanova, they all sort of over indexed to the models that were leading at the time when they designed their chips. And so they made certain trade -offs, right? They put a lot more memory on chip. And Nvidia was like, well, we're not gonna do that. That's faster at least, right? Well, more like if you could bring the amount of memory of SRAM on Nvidia's chips, it's much, much lower. Yes, correct. Then when SRAM instead of DRAM, but then they usually have less DRAM. So there's a trade -off there as well.
28:56There's less DRM, there's more SRM, and because there's more SRM on the chip, you have to have less compute on the chip, and so they ended up losing, right? Because the model size is got too big and all this, right? And so you have this like super weird dynamic where they bet on something that was actually better, right? Like I have no doubt that cerebris would run certain types of models better than Nvidia or GROC, or hey, Dojo, right? Dojo runs certain, you know, in Tesla's Dojo, would run certain types of models way better than Nvidia's chips. because they're optimized to that. But then it's like, oh, well, actually, even envisioned has to use vision transformers now.
29:31So it's like, okay, cool. Gives model sizes grew and all these things. So it ends up being a, you know, catch 22 and that like you optimize for something. And so now like today, you have this new age of AI accelerator companies are like, okay, we're going to optimize for transformers. But the time they started designing, they're like, okay, transformers are dense models that are this big. What's the best, you know, the hidden dimension is 8K and your back sizes are this big and your sequence ones are this big. So let's just make a super large systolic array. So you can create the maximum efficiency and that turns out, oh look at deep seek or go look at what the labs are doing.
30:04Actually, there are shapes are much smaller. Actually, you need to do a bunch of small matrix multiplies, not massive, massive, massive, you know, singular matrix multiplies per layer. And then it ends up, you know, oh, well, that chip you're designing for that is actually not super effective for that. And so the software is evolving constantly because of what, because of what works best on Nvidia. and you see that with, you know, whether it be what Deep Seeks doing or other Bobbos doing, or what the labs are doing internally. And you even see this like for Google, right? Like their open source Gemma models make different decisions because the shapes of a TPU are different than a GPU.
30:41And those, the GPU and the TPU are actually not that far apart, right? Like you would say yes, they're very different, but like Blackwell and TPUs are very, very, they're converging on similar designs, actually. Whereas, to be in video, you can't just have this supply chain win. You don't have this captive customer. So now you need to do something that will give you 5X advantage in hardware efficiency for a certain type of workload. And then pray the workload doesn't shift. Because in videos also optimizing their architecture generation, they've added a lot of stuff to make their chips way better for the existing models.
31:16but it's like they're taking large steps every year, every two years towards something, whereas you have to go way over there and left field and hope that models stay over there, right? Because you have to win by 5x because Nvidia is gonna have supply chain efficiency over you. They're gonna have time to market over you in terms of like a new process, note, or new memory, or whatever technology, right? Even AMD, right? They got to two nanometer before Nvidia. They had higher density HBM. They use 3D stacking. All these things on supply chain that should be better than Nvidia, and yet they still lose.
31:51They're still the software angle. I mean Nvidia is fantastic. Yeah, and then there's software as well, right? But it's like Nvidia's going to have better networking than you. They're going to have better HBM. They're going to have better process node. They're going to come to market faster. They're going to be able to ramp faster. They're going to have better negotiations with whether it's TSMC or SK high nicks in the memory and silicon side or all the racked people or like copper cables. Everything they're going to have better cost efficiency. So you have to be like 5x better. But to be fair, if somebody had a viable competitor, which would even be marginally cost competitive, if my guess is many of the big consumers of GPUs would immediately shift some revenue there, just to have a number tool, right?
32:23Just to just do that. That's AMD today, right? And Microsoft stopped. Yeah. I mean, like, there is still pre -limited traction, though. Sure, but yeah. Medic continues to buy from them. And Microsoft did buy a bunch, and then they stopped, because it's like, well, yes, they're, you know, AMD's giving you all these advantages, but ends up still not being better on a performance per lot basis and they have a way bigger software team. They're somewhat competitive on like all these dynamics that I mentioned, right? So you can't just like do the same thing as Nvidia. You really and do it better, right?
32:52Or try and execute better like AMD. Like you have to really leap forward in some other way. But that's the design cycle takes so long that models will shift, right? Because they're like, oh, what's the next inertia in TPU and GPU look like? Okay, let's optimize for that. And the research path is, you know, like great, like yes, neuromorphic computing could be the most optimal thing for us to do. But no one's working on that because you have to advance in the tech tree, you've chose it, right? If you restart the tech tree, you're going to be like, well, the sucks. And so like if it branches this way and you're over here, you're screwed.
33:25Because you have to be five x times as a mode. Because the supply chain stuff means that five x actually turns into a two and a half x. And then Nvidia can compress their margin a little bit if you're actually competitive. and then that two and a half X becomes like a 50 % better. And then yeah, so it's like it ends up being way too difficult to, and the software stuff, right? Everything like takes your 5X and makes it like, oh, you're actually only 50 % better. And defense supply chain for sure. Yeah, defense supply chain. And then like, they get that, right? Like so it's like, and let Nick himself said, we had to do this for Rural Thin Minerals.
33:57That's like interesting. China, there's like provinces in China that have like rules that say the H20 is not efficient enough to be deployed, which is like super bizarre because it's clearly the best AI chip China has. Paul is still a little bit behind. Well, what's interesting is that efficiency is just not, is so much less of an issue in China than here because they just have the power infrastructure to be able to support. So even if they're running less powerful chips, you know, you would imagine that it doesn't really matter because China has just such an infinite supply, infinite supply of power that they'd sort of be okay with it.
34:32So it's interesting. Which is a big challenge in America, right? Like, there have been companies that were like, they wouldn't, they'd like, you know, Johnson keeps saying he couldn't give away H20 in America for free. But I've literally like heard companies like say like now say like, yeah, no, I mean, I wouldn't because like, I only have this much power, how am I going to, you know, in data centers ready to go over the next year, if I bought an H20, I'd literally have less compute capacity and then I'd lose, right? Even if it was free, like it doesn't make sense. Whereas China doesn't care, they can build these things.
35:05They have the muscle. I'm curious how this all shakes out, you know, China's posturing really hard. They've been like put out something and I was like, we're investigating to see if there's back doors in the H20, it's like there's no back door in the H20 like chill. So, you know, it's like, you know, GPUs are usually like firewalled from the public internet anyways. Like you step through stuff before you get to the GPU clusters. So like a backdoor wouldn't even matter. I don't know, I think it'll be interesting to see because China can definitely deploy way, way, way more power to AI the moment they decide to.
35:46But there's these like, there's like competing interests, right? Like, because they want Huawei to be better than Nvidia. Yeah, and then this is how Nvidia argued to the administration. They're like, if we don't do this, I think it's like a very like powerful argument that like, like, for example, within Triton, which is a common ML library anyway, like, like bite dance has open source some stuff that plugs into this that is like super awesome. And there's like all these other libraries. It's not just models that China open sources. It's like software for Nvidia that Chinese companies open source.
36:18In a sense, like by Nvidia selling GPUs, Nvidia's argument again, like was like they were able to stop Huawei from building up a software ecosystem and the Western ecosystem is better. But then in the flip side, it's like, again, if you believe the models deliver more economic value to society than the hardware, which I actually think they do. It's just there's a value capture problem today. Then you're giving China way more by giving them H20s and soon a version of Blackwell that's cut down like Trump said, right? Versus, you know, selling in the chips, right? The economic value derived from selling in the chips is not as large as, you know, being able to somehow sell them AI services.
36:56So is China gatekeeping power for AI? I don't think so. I think, again, like there's a lot of like, What we see is that like even with H20 being sold to China into China and future versions of the chip, H20E and other chips, we still see like Chinese companies like Alibaba, renting GPUs outside of China because the GPUs they can get outside of China are just so much better on a dollar spend per performance basis, renting them or even like going through sort of like a Singaporean company that is effectively a Chinese company and building data centers and putting chips in them. So it's like, I don't think China's limiting the power per se.
37:40It's that it's, you know, you can only, if you can spend like your Chinese companies are growing their capex way more than US companies on a percentage basis next year. The dollar, absolute dollar number is, you know, obviously the US company are spending more still on AI. The percentage basis Chinese companies are growing more next year. And you, you still have the problem of like, well, dollars spend to AI output in tokens or in whatever is going to be lower because these chips are worse. So power is not the gating factor. It's always capital, right? At least today, right? Now, China can spend a lot more capital if they wanted to.
38:16They're subsidizing the semiconductor industry to the tune of like $152 ,200 billion a year through SOEs, through CapEx that's not generating revenue, et cetera. So it's not like they couldn't do this today, AI ecosystem, right? Given, you know, Metas CapEx is like 60 billion, right? and Google's capex is like 80 billion, right? Like they could totally spend way more than that on a single effort they just haven't decided to. And I just think for the US, our buildouts are constrained by power, right? Like Google has a ton of TPUs sitting waiting for data centers to be powered and ready. As does Meta with GPUs, right?
38:50We posted about how Meta's now building these like effectively tense. Is this just something we also coupled to the unwillingness to sell them to a broad ecosystem? system. I mean, if they want to be confined in their own data centers and they, you know, didn't ram data center build out for their own high for their own, so high basically, it's because it's quickly enough, right? And yes, that constrains them, right? If they were on the open market, we'll be still be constrained. Yeah, yeah, for sure, because like, companies like CoreWeve, you know, why is CoreWeve valuable is really because they build infrastructure really fast, right?
39:22And their software is nice, I think, but like a lot of their customers are bare model, right? Just just replace the GPUs whenever they're broken and networked. They think they grew more aggressively and I think I think they'll go anywhere. Jens and Lyson as well to be there. Yeah, they'll go. Yeah, yeah, that's very important as well. But they'll like go because it's it it it unconstitutes the ecosystem which is better for Nvidia. Right? Have you worked at it? I know exactly what you're going to as my. Yeah. So I think what's really important is that like core we have doesn't care, right? They're like, Oh, crypto data center.
39:50I will convert it to a AI data center, right? They bought a company for like $10 billion that's doing crypto mining, which is worth like $2 billion like a couple years ago. And it's not because they're Bitcoin mining businesses growing. It's because they have powered data centers, right? Like anywhere and everywhere, people are trying to build power data centers. And company is like like CoreWeave and Oracle are moving to the actually today. Google just didn't bought 8 % of a crypto mining company called Teril Wolf, right? Right. Yeah, because they're getting into crypto mining. No, because they need the data centers, right?
40:24They need the power, right? And it's like, all the hyper scalers have like said, screw off to my sustainability pledges because they need power as fast as possible, right? They're doing things that are not, that they take a little bit longer to move the ship, but like, even if you didn't do it in your own self -built data centers, there's still a lot of challenges in the open market. There's a deficit, right? And that's, that's constraining American chip buildouts heavily. Yes, others could maybe do it a little bit faster, like Corvieve or others, right? Oracle's got an open mind as well. But it's still constraining US buildouts heavily.
41:07Even though the capital has been spent, right? The chips are 60 to 80 % of the cost of the cluster, depending on what chips you're getting. So it's like, they've already bought the chips. They just can't put them anywhere because the data centers aren't ready. Supplies to Google, Pyson, Microsoft, PysDemeta, PysTo a lot of folks. I mean, it's really hard to build power in the US. Yeah. Power, greener connections, transmission, substations, all of this stuff. Electrical contractors, electricians, in Texas, if you're willing to be a travel electrician, it's like, it's like oil pay, right? It used to be that if you're physically adept, You could go make, you know, a hundred grand in West Texas, but like who the fuck wants to do that?
41:47Now it's like, well, you could go like 200 miles away from Dallas and what's still a reasonable town and build a data center and work on the wiring within the data center and all this other stuff, the transmission stuff, and your pay is up like 2x now versus what it was just a few years ago. This labor problems are challenged too. And it's, yeah, I think in China, they don't have any of these problems, but they just haven't spent the capital yet. but capital is an issue as well. Because of the scale of what's being spent, right? Like, in videos revenue this year is going to be like over $200 billion, and next year it expects over $300 billion, plus Google is going to spend like $50 billion on TPU data centers, right?
42:27And it's like, and Amazon's going to spend tons and tons on training data centers. It's like, the scale of dollars is quickly growing to nation state level stuff. And what's more important is being able to decide to spend the dollars in what's cost effective. And so to some extent, China still constrained by that, but they can smuggle chips in, they can build data centers outside of China, they can rent data centers outside of China, and have the most cost effective, blackwell chips or whatever, right? Bike dance is either the biggest or the second, biggest customer of Google Cloud for a reason, right?
43:03And they're getting, you know, over, you know, they're getting many, many blackwell from them, right? And the same with Oracle and the same with Microsoft often, all these other companies are renting tons of chips to China anyways, because it's more cost effective to do that than build it yourself. So it's not like China has this mentality where we only have to, well, the government does, but the infrastructure companies don't, like Alibaba, Tencent, Baikdance, etc. So what's the endgame for data centers? I mean, like we need more power, we need more cooling, well, the MPE, every, all data centers would be next to a nuclear reactor, lots of solar, you know, next to a deep level, like deep sea water that we use for cooling or something like that or what's the thing?
43:42I think that cooling is, like the physical cooling of a data center are like, you know, there's this little narrative about like, oh, yeah, I use this so much power and it's like not really, you know, farming else alpha uses like 100x the water of AI data centers, even by the end of the decade, it'll be the same. And it's like, alpha is like worth very little. So it's like, I don't know. There's like, it's like cooling is like not that. You know, people have like experimented with like, you know, undersea data centers to reduce the cooling cost. Yeah, that doesn't make sense. It's like 5, 10 % savings, but then like, if you were to get the water out of the ocean then, then with the data center, the ocean.
44:16And so if you want to service it like you're screwed, right? So like the same with power, it's like, we talk a lot about like, the power is not actually that expensive. It's just hard to build, right? How to get to the right place? I'm going to get to the right space and convert it down to the voltages and all the stuff that chips need. So it's less than magnitude of power and more where it is and how it moves. Well, the magnitude too, right? Like it's going to be. It does a total world wild energy consumption. AIJ does it is still the full. It's like, it's like nothing. Yeah, it's a fraction of a percent.
44:44It's not that. Yeah, yeah. Even by the end of the decade, you know, the US will be like 10 percent of our power will be AI data centers, which is still like electricity. Of our electricity. I know it's a bit of energy that's even a smaller fraction, right? Oh, yeah, yeah, because you think about. My shift into electric vehicles, I was like, you can probably make a bigger swing than and then, you know, with all the data centers, we can build with the data. But outside, like, it's like in Europe, like that number's not moving up that fast and like all these other countries, I think we need to build a lot more power, but it's not like some crazy, crazy amount.
45:16It's just like doing it properly is the hard thing. And again, like the cost of power, like, you go look at like these deals, people are signing, they're still signing, like even though the price is skyrocketed from like, if you sense a kilowatt hour for these massive, massive purchases to like 10. It's still, you know, when you think about the full TCL, the cluster, you know, the GPU cost of networking, all of this stuff far out, trips the power. Yeah. And same with cooling. But what percentage is power from like a V8, do a four year amortized GPU data center? What percentage will be power 80 % of the cost of a GPU data center if you're building black well is capital, right?
45:53It's the GPU purchases. It's the networking. It's the it's the physical data center conversion, conversion, power conversion equipment, all of this stuff is like 80 % of the cost. And then 20 % is going to be your lack and your power and your cooling and your cooling towers and your backup power and your generators and all this stuff. It's like nothing, which is why it doesn't matter if you spend, you know, 10 % or 50 % more on that. Because at the end of the day, the expensive thing, right? Like this is why what Elon did would seem silly, right? they spent a lot more money on generators outside the data center and these mobile chillers to cool the water down for their liquid cooling instead of like the more cost -effective option because it got the data center up three months faster.
46:37And so like that three months of additional training time is worth way, way, way more on a TCO basis, right? The performance you got out of the chips and the time to market and all this is way, way faster. And therefore, it was the right decision. Even though this part of the data center ballooned at cost, everything else is still there and you're still paying for the chips. And if they were sitting idle, it's not worth it. Right. Just by like bypassing the grid, bypassing anything to do with interconnect, anything to do with public utilities. Exactly. Exactly. What's your take on Intel? Where's Intel going?
47:07I think the world, well, the US needs Intel. I think the world needs Intel. I think the world needs Intel because like Samsung is doing worse than Intel on leading edge process development in my opinion based on even on various customers in the industry having done touch chips at like Intel versus Samsung. They think industry generally agrees that Intel is further along the sort of the two nanometer class process technology than Samsung is, but both are way behind TSMC. And TSMC is a monopoly in some extent. The number one question always people ask is like why is TSMC not making more money? Why are they only raising prices next year, you know, 3 to 10 percent depending on what it is.
47:48It's like, just some season monopoly. Like they could raise a lot more, but they're they're good Taiwanese people rather than like dirty American capitalists. So TSMC was owned or was managed by Americans. I think most ownership is actually American in terms of the stock, it's on the New York Stock Exchange and all this like, you know, they would have raised prices a lot more. And so like there is this like difficult difficult thing to be done that like, hey, there's one island that controls all leading edge semiconductors and not just all leading edge, like the majority of trailing edge production as well.
48:21Yeah. Something needs to be done until it's behind, but not like, not like absurdly so, right? Like if something were to happen to Taiwan, Intel would have the most advanced technology in the world, right? It's just, it's not economic. Can you keep Intel as one company if you want them to be competitive? I think the process of splitting it would take so much executive time and so much executive effort that you would have been bankrupt by then, right? And that's the big challenge. Like I think Intel should be separate, right? But to properly split the company and for all the management time that's needed, there's like absurd.
48:54And instead like what you need is like, you need Lip -Boutan, who's the CEO of Intel. You know, there's a lot of drama going around about him because he's one of the greatest semiconductor investors ever, right? He's invested in so many different companies first. You know, he was on the board of like SMIC, which is China's TSMC effectively, which is like a big like drama or like the, some of the biggest tool companies in Chinese, the first investor in them, because you know, there's a multi -polar world there and he's making good investments. But like, you know, now like people are getting mad about that but it's like, no, he recognizes he, the company's, like he understands the supply chain.
49:29He needs to not spend his time on splitting the company because then he never actually fixes the company, right? Intel's problem is that like, it takes them five to six years to go from design to shipping the product. in some cases more. And when they tape out a chip, right? Like, you send the design to the fab, the fab brings back the chip. They go through 14 revisions in some cases, where it was like the rest of the industry goes through like one to three, right? Revisions, if they're good. Of like send the design in, get the chip back, test it, send the design in, right? For a public launch.
50:03And they'll launch a chip in three years. So, but if you look at it until today, right? They still don't have a competitive entry on the AI side. And they what? Can you, so what does it mean for their offering? I mean, they're doing great on CPUs. They don't have a good AI, a Azure product. Is it long term sustainable positioning? Right? I mean, there's a standard on chip company here. I mean, IBM still makes more money every launch off of mainframes. So it's not like X86 is dead. It's like you don't get the growth rates, but like you could totally run this as a very profitable enterprise. And I think the same with PCs, right?
50:37There's some turmoil, there's some arm entry, there's some AMD competition. Well, I think it's a very, it can be a very profitable business if it had like one third the people or half the people working on it. And so like Lipputan to fix Intel needs to go into both the design company and lay off a shitload of people, but like keep all the good people and make sure that they're designing fast and they're launching from design conception to launches two to three years, not five to six. And that's on the design side and make that profitable. And then on the FABs, he has to do the same thing. There's all these people, like, one of the heads of FAB automation at Intel.
51:15I explicitly told Lip Bootan, because we have a couple of X Intel people who are actually good in the company that worked on the FAB side. And we're like, they were like, who's the worst people in France? It's like, oh, this guy sucks. I explicitly told Lip Bootan, he had never talked to the guy because it was like, four layers down. The company has like absurd amounts of hierarchy. It's like four layers down. He goes to talk to the guy and he's out, right? it's like, he figures out who's bad, right? And who's good? And he has to go in and he's still like, hey, the vast majority of the team at Intel is the one who led the world in production and process technology for 20 years.
51:48But there's a lot of built up crap. So he has to go figure this out. He can't waste his time on all this structuring to split. I think it would be better if the company split. I just don't think he can spend the time to do that. And if the design side of the company is, you know, you're not really going to get into AI, you're not really going to, you have to make some money there. But the fabs, I think, could truly become a competitor. But they're going to go bankrupt by the time anything can happen. So they have to figure out how to get capital. So he has to figure out how to get capital. He has to figure out how to clean up all the crap.
52:21Make the yields go up, right? Make the product ship way faster. Like all of these things are basically - I think the goals are completely correct. I mean, I think the big challenge is just for a vacuum back in my time there, right? I think the big challenges that right now, if you look at Intel, right, they have essentially software, software, the chip design making it, and then there's something to the core manufacturing part, right? And they have three different, very different cultures. And it's very hard to get everything on the one on the brother, right? And so I think that is the big challenge.
52:48I think you should do that run the company separately, right? But like, you can't physically separate them entity -wise because it's going to take so long to sever all these things, because he doesn't have time, right? like Intel is literally going to go bankrupt if they don't have a big cash infusion, or they lay off like half the company, right? Which some could argue you need to lay off like 30 % of the company anyways, but there's a lot of bad things that happen if that happens, right? And they need to spend a lot more on building the next generation fat, even if they fix the fat. And they don't have money for that, right?
53:19So there's like, there's like a lot more more important problems than like physically separating the company, even though I think long -term, Yes, the fab has to be separate from the chips design software, right? Like our chip design part of the company just like that's going to make each company much more accountable Be able to service their customers better etc It's just that's going to take too long and they're going to go bankrupt by then Awesome I but I think I think I hope I hope I pray Someone does something right like you get a big capital infusion I don't know the big hyper scalers are like muscled into like oh, okay Wait if TSMC eventually grows their margin to 75 % because of the monopoly plus they intake all this stuff, like co -package optics and power delivery and all this, like all of a sudden the cost is gonna spike.
54:01So we should actually just throw $5 billion at Intel each, right? Screw it. And that could actually give Intel enough of a lifeline to potentially get to something and maybe be competitive. That's the hope. Can we finish by finishing this game that we started when we gave Sam and Altman advice if Jensen was here, what advice would you have from? If Jensen was here, I think he has a massive, massive balance sheet, right? Jensen does. He's cash -free cash flow is ridiculous. The tax cut, the new Trump tax bill institutes something really incredible, which is that you can depreciate all of the GPU cluster cost in your one, which we put out like a note about how the tax implications to meta are 10 billion dollars a year.
54:53And across each of the major hyperscalers, it's massive. It's like, well, Nvidia is going to spend tons and tons of cash, or they're going to spend tens of billions of dollars of taxes. Why don't you get into the infrastructure game somehow? Now this is obviously going to be crazy because now they're buying GPUs and putting them in data centers and doing stuff and they're competing with their own customers, but they're already doing that anyways because their customers are trying to make chips. But they should accelerate the data center ecosystem with investments, right? Because really, we think we can have very high degree of accuracy on what they're going to do next year in terms of revenue because it's just the number of data center watts that are being built, right?
55:35This is harder thing to shift up and down, right? Now there's a little bit of share difference between how much is TPU versus TPU, but you You have to accelerate the infrastructure and you need to spend all of this capital that you're building, right? Like, okay, do you want to go the route of like doing buybacks and dividends? Like, great. Like, you're a loser if you do that, right? Like you can make more money by reinvesting and building a bigger company that's not just chips into the ecosystem or servers into the ecosystem, but actually like controlling the infrastructure end to end somehow.
56:05So I think there's something he could do there with this massive war chest. And there was a reason like Nvidia has done some buybacks and they've done some dividends and increasing, but the cash on their balance sheet keeps growing and they're going to have north of a hundred billion dollars of cash on their balance sheet by the end of this year, I think. So it's like, what are you going to do with that? I think there's something moving into the infrastructure layer much more that they could do if you really wants to be the king of the world, right? Which I think he does. As Sergei and Sender?
56:39I think they should open up the Commodo on TPs, right? Like start selling them, open up the software, open source a lot more of the XLA software because there's open XLA and there's XLA but the vast majority of the closed source, really, really open up the Commodo on that and be a lot more aggressive, right? They're still pretty not aggressive on data centers. They're pretty not aggressive on a lot of elements to the company. The TPU team's next -gen designs are pretty not aggressive, partially because a lot of the TPU team has left to go to OpenAI, the best people that I knew. It was actually really annoying.
57:15I knew like four people or five people and they all went to OpenAI and it's like, fuck, like, now I don't get as much. I met some other people, right? But it's like, you know, I think they could be a lot more aggressive in many ways across the company. They don't have to be, right? But they could. Because AI, you know, like this Chad GBT take gray, the shift of search queries, the monetizable ones, especially from two purchasing agents is going to really screw Google long -term if they don't get their act together. I think they've gotten their act together on DeepMind. There's still some inefficiencies, but Sergey works on, works with DeepMind a lot in their driving hard.
57:50There's still a little bit behind, but like, I think like physical infrastructure, TPU, and how much money they could make, and how much they could take the wind out of everyone and else's sales if they start selling TPUs externally and reorg around like building data centers much faster so that they do have the most compute in the world because they did. But now there's certain companies that are gonna surpass them potentially over the next few years if they don't really get their act together. So I think that's what I would say for them. Yeah, and also like learn how to ship product. Yeah, I think you should.
58:23Zach. I think, I think, Zuck, you know, it remains to be seen what goes on with super intelligence, but like they're trying to move super fast with the data centers, you know, like screw it, we'll build tents instead of like physical data centers because we only need these for five years anyways. You know, the super intelligence moves, you could, you could say whatever you want, but like, you know, trying to buy like thinking for like 30 billion or SSI for 30 billion didn't work out. So then they spent, you know, not even that much on higher, not 30 billion on hiring all these people. So I think that he recognizes the urgency with the models, with the infrastructure.
58:59So I really think he needs to like, you know, if you read his website post about like AI, like I think, you know, he sees the vision, right? There's the wearables, there's integrating AI into that, there's being your AI assistant to do all this purchasing and stuff. I think he sees the vision, but I think he also needs to focus on like actually like releasing that faster, but also like the products that they do outside of their core IP, every time they launch something is kind of mid, right? You know, metal reality labs is doing well, but I think they should like go more explicit, like have a chat, GPD competitor, have a cloud, like cloud code competitor, like just start releasing way more products because they're really just focused on their individual gardens rather than like branching outside of it.
59:46Do you think apples should have that same sense of urgency or 10 cook was here? What would you tell them? The funny thing is like some of their pissed AI people are now like at super intelligence. They're building an AI accelerator. They're going to they're they're have AI models, but they're just like way slower. They did mention on the last earnings call they're going to allocate more capital to this, but it's like, guys, Apple, like you guys are going to lose the boat if you do not spend like $50 $100 billion on infrastructure. You don't think the kind of we'll cut it. I think like more and more you'll see people like, you know, great Apple has this world garden, but like they can only do so much to protect it, right?
1:00:23IDFA, like they shut down ads to or data sharing to meta, but meta made better models. And now they have way more data and way more power over the user than they ever did before. Kind of it was good that meta kicked the crutch off of them or Apple did. But the same applies to like AI, like yes, they have access to the text and they have access to this. But like, I think other people are going to be able to integrate user data. And agents will be able to integrate all this user data. And they'll start to lose control of what the user experience is. As more and more gets disinheredated by AI being interface rather than touch, rather than touch pad and keyboard.
1:01:00And I don't think they've truly realized what happens when the interface to computing is AI. Like they market it, but like that's going to shift computing really heavily. they have great hardware and their hardware teams are working on awesome stuff and form factors, but like I just don't know if they like get what is actually going to happen to the world in the next five years truly well enough and they're not building fast enough for it. What about Microsoft to that end? Microsoft has the same problem I think they were super aggressive in 23 and 24 and then they pulled back heavily right now like opening eyes slipping through their their graphs, there's that whole thing there.
1:01:40They cut back on data center investments heavily. They were gonna be the largest infrastructure company in the world by like a factor of 2x, which would have been, you know, you could argue maybe that was like too much and maybe it wouldn't have been economical, but like they're losing grasp on open AI, their internal model efforts are failing, spectacularly like they're on LLM arena right now and they're pretty decent there, but it's like that's just like a sycophantic model. Like it's a code name, but like whatever. Like, M .A .I. is like failing, Azure is like losing a lot of share to Oracle and Core Weave and Google and so on and so forth, right?
1:02:16Their internal chip, chip effort is by far the worst of any hyperscaler. Like, they're just like, mis -executed. Like GitHub, how does GitHub not the highest ARR software code model? I mean, they only had the best IDE, the best source code repository, the best enterprise I say it's for us, the best model company as a relationship and they're by the first market. It's like they're never going for them. And like there's just nothing, right? It's like, get up, get up, get up, co -pilot is failing. Microsoft co -pilot is like still crap, right? Like, it's unusable. It's like, what is, what is, you know, what is going on?
1:02:51Like you did a shake to crap out of the company. Like, I think they win a lot because they have the best business to business relationship with so many enterprises. That's why it's for us on the planet. Yeah, but like they end up like not having actual product to sell them, which is like really scary So they need to really work on product. Yeah, such such as done great on sales and stuff But like yeah, if Elon was here, what advice would you have him? A lot of people at XAI are mad about the porn models like porn porn stuff It's fine like you're gonna make a ton of money off of this This is how you accelerate the revenue of that company But like he's losing a lot of talent and and acting a lot of good projects but Elon has is a magnet to amazing talent and building stuff so I won't bet against him but it seems like since he left the administration and focused on stuff again.
1:03:35But I think I don't know, I think he's focused on a lot of things and I think like Robo Taxi, starting to look good actually again. Like I haven't ridden one yet but I have some friends who've ridden one. It's like looks pretty decent. He could like not make these snap decisions which often are the reason why he's amazing. But like some of these snap decisions are hurting him. I'm not sure if I can give Elon that much great advice because I think maybe it's just like focus on like the products again, right? More. But he is working on that stuff a lot. Yeah. I think that might be a good place to wrap.
1:04:05Awesome. It's a great discussion. Don't thank so much for joining us. Thank you for having me. Thanks for listening to the A16z podcast. If you enjoyed the episode, let us know by leaving a review at rate this podcast dot com slash A16z. We've got more great conversations coming your way. See you next time. As a reminder, the content here is for informational purposes only. Should not be taken as legal business, tax, or investment advice, or be used to evaluate any investment or security, and is not directed at any investors or potential investors in any A16Z fund. Please note that A16Z and its affiliates may also maintain investments in the company's discussed in this podcast.
1:04:42For more details, including a link to our investments, please see a16z .com forward slash disclosures.
From the publisher
The AI hardware race is heating up, and NVIDIA is still far ahead. What will it take to close the gap?
In this episode, Dylan Patel (Founder & CEO, SemiAnalysis) joins Erin Price-Wright (General Partner, a16z), Guido Appenzeller (Partner, a16z), and host Erik Torenberg to break down the state of AI chips, data centers, and infrastructure strategy.
We discuss:
- Why simply copying NVIDIA won’t work, and what it takes to beat them
- How custom silicon from Google, Amazon, and Meta could reshape the market
- The economics of AI model launches and the shift toward cost efficiency
- Infrastructure bottlenecks: power, cooling, and the global supply chain
- The rise of AI silicon startups and the challenges they face
- Export controls, China’s AI ambitions, and geopolitics in the chip race
- Big tech’s next moves: advice for leaders like Jensen Huang, Sundar Pichai, Mark Zuckerberg, and Elon Musk
Resources:
Find Dylan on X: https://x.com/dylan522p
Find Erin on X: https://x.com/espricewright
Find Guido on X: https://x.com/appenz
Learn more about SemiAnalysis: https://semianalysis.com/dylan-patel/
Stay Updated:
Let us know what you think: https://ratethispodcast.com/a16z
Find a16z on Twitter: https://twitter.com/a16z
Find a16z on LinkedIn: https://www.linkedin.com/company/a16z
Subscribe on your favorite podcast app: https://a16z.simplecast.com/
Follow our host: https://x.com/eriktorenberg
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Stay Updated:
Find a16z on X
Find a16z on LinkedIn
Listen to the a16z Podcast on Spotify
Listen to the a16z Podcast on Apple Podcasts
Follow our host: https://twitter.com/eriktorenberg
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

