In short
The “compute race” behind AI—chips, data centers, power/networking constraints, custom silicon, export controls, and AI economics—plus why NVIDIA remains hard to displace.
Guests
Dylan Patel (co-founder of SemiAnalysis; covers AI hardware/semis and data center trends). Other hosts/speakers include Aaron Price-Wright and Guido Eppenzeller (and the narrator).
Key claims
AI competition is as much infrastructure and economics as models. NVIDIA’s moat is supply chain + software ecosystem + performance-per-watt and fast iteration, making competitors need ~5x hardware leaps. Value capture is “broken”: labs often don’t capture much of the value they create, so GPU spend may be tempered even as demand grows. Custom silicon (Google TPU, Amazon Trainium, Meta) is a major threat if it expands beyond hyperscalers. Power and data-center buildout (grid interconnects, labor, cooling) are major bottlenecks in the US.
Notable examples
OpenAI “router” monetizing free users by routing low-value queries to cheaper models and high-value ones to “thinking”/agents. Google/Meta/ByteDance buying large volumes of chips; CoreWeave/Oracle building power-centric infrastructure. China’s H20 deployment constraints are framed as capital/power placement, not just gating.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VODylan Patel Joins the Discussion
1:09 to 1:50
Dylan Patel is welcomed to the podcast as an expert in AI hardware.
“We've been trying to get you for a while.”
Reactions to GPT-5 Capabilities
1:50 to 3:56
Dylan shares insights on his reactions to GPT-5 and its comparative performance.
“We just had some of the research for Christina and Isabella on here last week.”
Monetization Strategies for Free Users
3:56 to 7:40
Discussion on how OpenAI's router feature may help monetize free users effectively.
“OpenAI can now control how much a computer wants to allocate to you, right?”
Pricing Models and User Stickiness
7:40 to 11:28
Exploration of the impact of pricing models on user stickiness and engagement.
“And so to some degree, so where you're on the parade of frontier between cost and performance is the new benchmark for model competitive no longer cost alone.”
Advice for OpenAI's Future
11:28 to 12:51
Dylan offers strategic advice to OpenAI on enhancing their product offerings.
“you kind of know in a general sense how many hours a day they're programming and what that sort of looks like.”
NVIDIA's Growth and Future Paths
12:51 to 13:24
Discussion on NVIDIA's impressive growth and potential future trajectories.
“Like a whole line of questions around this.”
Demand Trends in AI Hardware
13:24 to 14:01
Analyzing the demand trends for AI hardware and the players involved.
“But that's actually like, okay, well, like 70 % of the stuff, like who's making off?”
Exploring AI Ads and Value Generation
14:01 to 15:41
Discussion on the potential growth of AI and its economic implications.
“I think with the, you know, we talked about like coding, right?”
AI's Value Capture Challenges
15:42 to 17:44
Analyzing the difficulties AI companies face in capturing value from their innovations.
“Right, let's say 100K value add per developer.”
Capital Expenditure in AI Infrastructure
17:45 to 19:38
Debating the future of capital expenditure in AI and infrastructure development.
“In so many words, you're saying, we're getting commoditized and therefore you can't capture the value and thus you should temper your expectations of how much you can spend on GPUs.”
Show all 26 chapters
Threats to NVIDIA and Custom Silicon
19:39 to 22:50
Examining the threat that custom silicon poses to NVIDIA's dominance in the market.
“you'll get profit out of it, but there's no like 100 % certain like, you know, way to argue it.”
Silicon Startups and Industry Competition
22:51 to 28:00
Discussion on the influx of capital in chip startups and competition with established players.
“Historically, no pun intended software has eaten the world in most markets, right?”
The Challenges of Neuromorphic Computing
28:00 to 29:48
Explore the complexities and trade-offs in neuromorphic computing designs versus traditional GPUs.
“So did we sort of pick the model for an architecture?”
NVIDIA's Competitive Edge
29:48 to 32:18
Understand why NVIDIA continues to dominate the GPU market despite competition.
“They're like, okay, we're going to optimize for transformers.”
China's Power and AI Infrastructure
32:18 to 34:24
Examine how China's approach to power infrastructure influences its AI capabilities.
“There's still pretty limited traction, though, right?”
The Global AI Chip Landscape
34:24 to 37:08
Discuss the implications of AI chip distribution and China's software ecosystem.
“because they just have the power infrastructure to be able to support.”
Capital and Infrastructure in AI Development
37:08 to 42:00
Analyze the challenges of capital and infrastructure for AI data centers in the U.S. compared to China.
“I think, again, there's a lot of like...”
Challenges in Data Center Growth
42:00 to 43:30
Explore the challenges and scaling issues in the data center industry.
“and build a data center and work on the wiring within the data center and all this other stuff, the transmission stuff, and your pay is up like 2X now versus what it was just a few years ago.”
Cooling and Power in Data Centers
43:30 to 46:20
Discuss the significance of cooling and power management in data centers.
“So what's the end game for data centers?”
Intel's Current Standing and Future
46:20 to 48:38
Analyze Intel's position in the semiconductor industry and future prospects.
“Because at the end of the day, the expensive thing, right?”
Intel's Operational Challenges
48:38 to 53:36
Delve into Intel's operational inefficiencies and necessary changes for improvement.
“Can you keep Intel as one company if you want them to be competitive?”
NVIDIA's Strategic Recommendations
53:36 to 56:00
Evaluate strategic recommendations for NVIDIA to enhance competitiveness.
“Just like that's going to make each company much more accountable, be able to service their customers better, et cetera.”
NVIDIA's Strategic Moves and Cash Reserves
56:00 to 58:24
Explore NVIDIA's financial strategy and potential in the infrastructure space.
“So I think there's something he could do there with this massive war chest.”
Zuckerberg and Meta's AI Focus
58:24 to 1:00:50
Discuss the urgency and direction of Meta's AI initiatives and product launches.
“Yeah, and also learn how to ship product.”
Apple and Microsoft's Competitive Challenges
1:00:50 to 1:02:59
Analyze the competitive challenges facing Apple and Microsoft in the AI landscape.
“And I don't think they've truly realized what happens when the interface to computing is AI.”
Advice for Elon Musk and Future Directions
1:02:59 to 1:04:09
Considerations and advice for Elon Musk regarding talent retention and product focus.
“Yeah, but they end up not having the actual product to sell them, which is really scary.”
Transcript
Automatic transcript. May contain errors.0:21Erik Torenberg:The AI race isn't just about models. It's also about the infrastructure underneath them. chips, data centers, power, networking, and the economics that determine who can keep scaling. In this conversation, Semi Analysis co-founder Dylan Patel joins Aaron Price-Wright, Guido Eppenzeller, and me to discuss the state of AI hardware, why NVIDIA remains so difficult to compete with, and how companies like Google, Amazon, Meta, and OpenAI are approaching the next generation of AI infrastructure. We also explore custom silicon, AI economics, robotics, export controls, and what founders and investors should be paying attention to as the compute race
1:08accelerates.
1:09Erik Torenberg:Dylan, welcome to the podcast. Thank you for having me. We've been trying to get you for a while. You're a busy man, but it worked out. Guido, why don't you introduce why we're so excited to have Dylan on the podcast and what we're excited to discuss? I think, Dylan, you've done an exceptional job in covering what's happening in the AI hardware space, AI semi-space, and now more and more data center space as well. And just looking at it, currently the most valuable company on the planet is an AI semi-company, right? The, I think, biggest IPO so far in AI was an AI cloud company. This is currently where it's happening, right?
1:40And any gold rush in the early days is the peaks and troubles that make money. And I think this is the stage that we're in. So super excited to have you here today. Awesome.
1:47Dylan Patel:Thank you. Happy to talk about my favorite topics.
1:50Erik Torenberg:Amazing. Well, maybe let's start with GPT-5. We just had some of the research for Christina and Isabella on here last week. You said it was disappointing. Why don't you share your reactions or what capabilities you were hoping to see or overall? I think it depends on what tier of user you are, right? If you're just using GPT-5 and before you were$20 or$200 a month subscriber, you no longer have access to 4.5, which in my opinion is still a better pre-trained model for certain things, or you no longer have access to O3, which would think for 30 seconds on average, maybe, right? Whereas GPT-5, even when you're using thinking, only thinks for like five
2:26Dylan Patel:to 10 seconds on average, right? Which is an interesting sort of phenomenon, right? But basically, like GPT-5 is not spending more compute per se. The model did get a little bit better on a vanilla basis, right? 4-0 to 5 is actually quite a bit better. But when you think about, what is this curve of intelligence, right? It's like the more compute you spend, the better the model gets. And that's whether it's a bigger model, which GPT-5 isn't, right? You can see it's not a bigger model. It's roughly the same size. Or you think more, right? But again, this is something that OpenAI's first thinking models, the first few generations of 01, 03, would think for a long time and waste a lot of tokens, if you will.
3:05And when you look at, for example, Anthropics thinking models, even when you put them in thinking mode, they think a lot less. to get to the same results or better results as OpenAI was. And so OpenAI, I think, optimized a lot of, well, if I ask, I think the silliest one I had asked was, I asked 03 once, is pork red meat or white meat? And it thought for 48 seconds, it was like, what are you doing? They should just tell me the answer. And so the nice thing is that GPT-5 will think a lot less, even if you select thinking manually, but more importantly, they have the auto-functionality, the router, which lets them decide whether or not, Like, hey, do I route to the regular model?
3:42Do I route to maybe many if you're at a rate limits? Or do I route to thinking, right? And how much do I think? But in general, the thinking model will think less. So there's less compute going into a power user's average query than before. But isn't it even more interesting? OpenAI can now control how much a computer wants to allocate to you, right? If we're in a high load situation, maybe tune the router a little bit so it's less, right? I have no idea what they're doing behind the curtain. But there's this meme out there at the moment that basically all they did, which is a meme, right? It's not true.
4:12But all they did is take O3 plus a couple of smaller models, put a router in front, and offer that at a lower branded price, essentially, right? I think there's a little bit of that, right? Cost suddenly matters, and they figured out a way how they can steer that. I think, yeah. I mean, and they talked about how they've been able to dramatically increase their infrastructure capacity. Because I myself was just regularly using O3 or 4.5, right? And now I'm forced to use auto, which sometimes gives me the 03 equivalent thinking model, but sometimes gives me just a regular base, which sucks. But I think for the free user, it's actually quite interesting.
4:46The free user was not getting thinking models pretty much ever, or not using them. Or in many cases, they just opened the website and asked their query, and now sometimes their query gets routed there. So sometimes they get a way better model. But now sometimes the opening icon gracefully degrade them if they need to. I think the router points to the future of OpenAI from a business. You can look at the model companies. Anthropic is fully focused on B2B, API, code, etc. Or cloud code, whatever it is. OpenAI, yes, they have that business, codex and API business, but really the majority of the revenue is consumer.
5:22And it's consumer subscriptions. But they have no way to upsell, to make money off of all the free users. In any other application consumer app, the free user still pays via ads. But this is not compatible with AI, right? Like it's a helpful assistant. You can't just make the user's result worse by injecting ads. Banner ads don't really work in AI either. So it's like, how do you now monetize them? And I think with the router, they're getting really close to figuring out how to monetize that user, right? With the new CEO of applications, if you saw her product that she launched at Shopify, I think it was Shopify, was an agent for shopping, right?
5:57And now this like immediately clicks like, oh, if the user asks a low value query, hey, why is the sky blue? Just route them to many, right? The model can answer perfectly fine. And that is a chunk of queries, right? But if they ask, what's the best DUI lawyer near me, right? All of a sudden, this is like, you know, you're in jail, you have one shot, you're like, screw it. Let me ask Chad GPT what the best DUI lawyer is. And now all of a sudden, the model's not capable of it today, but soon enough, it'll be able to contact all the lawyers in the area and figure out what their results are and maybe search their like court filings and whatever, right?
6:30Book the best lawyer for you or an airplane ticket. Negotiate a cut as part of that. Yeah, of course they're going to take a cut, right? But there's a much better way of monetizing the free user. It's like Etsy, 10 % of their traffic now comes from chat. And OpenAI makes nothing off of that. But they really, really will soon, right? And partially that's because Amazon blocks chat. But there's a way to make money from shopping decisions, whether it's booking flights or looking for items. And those you now say, free user, I don't care. I'm going to send you to my best model. I'm going to send you to agents.
6:59I'm going to spend ungodly amounts of compute on you because I can make money off of this. But if it's a query that's like, help me with my homework, I'll send you a decent model, right? I don't need to spend money on you. And so this is how I think opening, I can finally make money off of the free user. And I think that's the biggest thing about the router, right? This is super interesting. I think this is the first time that we've seen that there's a launch of a new model where to some degree cost is the headline item, right? I mean, so far I was always like, who has the smartest model? Who has the highest MLU score?
7:29Now we have suddenly people who use models for coding for eight hours a day and surprised that if you take a large context window and the best model creates thousands of dollars of cost a month. So cost matters. And so to some degree, so where you're on the parade of frontier between cost and performance is the new benchmark for model competitive no longer cost alone. Is that what we're seeing here? I mean, I think definitely, right? Like OpenAI said they doubled their rate limits for big amounts of users. They've dramatically increased the number of tokens they're serving from this launch, which effectively says this is an economic release.
8:00Probably also means the tokens are not cheaper, right?
8:02Dylan Patel:Yeah, yeah, for sure, for sure. I think the funniest thing is this whole cost thing you mentioned is like, we've seen this in the code space, right? Cursor had to pull away the unlimited cloud code. Initially, they have this super expensive plan and it had like unlimited rates and then they were only like a weekly rate limit. Now they have like hour-based rate limits. And I saw the craziest thread on Twitter where this guy said he changed his sleep schedule, right? modeled after like how sailors in the bay. If you're sailing, you can't sleep, right? Like solo sailing, they'll take like power naps when they get to the right spots so that they can like still be safe.
8:35Erik Torenberg:In the morning when it's not very windy.
8:37Dylan Patel:Well, but like they can't sleep uninterrupted, right? And so because Anthropic had to put rate limits that are like not just week-based, but like number of hours based. And like he like basically sleeps multiple times a day, but small chunks just so he can maximize the usage. And there's also a leaderboard on Reddit where people are like competing to see how many tokens they're using through their subscription. And there's like a dude spending like$30 ,000 a month. So I'm going to find some developer in India that I can do pair programming with so I can get the day cycle, he can get the night cycle, and we both can maximize together the quota for the account.
9:09Is that the future then? I mean, but it's clear like people are taking advantage of the negative gross margin, like sort of subscriptions that are offered. I think Anthropic probably makes a positive gross margin off of my subscription. I don't code enough. But there's plenty of people that are definitely losing money. And so as you said, it's an economic... It'll push more and more to, I think, just usage-based pricing, right? If you have an underlying commodity that you're reselling to some degree that is that large a part of your cost of goods, you need to go to usage-based pricing.
9:36Dylan Patel:How much do you think the customer capture and stickiness for these code products is? I'm curious what you think on that, right? Once you use an IDE, once you integrate one of the CLI products in, how sticky is it? That is a billion-dollar question. That's a very conservative estimate. Look, Andrew Carthy has this great slide where he basically says, if you're building an agentic system today, right, but fundamentally what it is is of this loop, right, where half of the loop is the model thinking, right, and then trying to do some. The other half is then the user verifying what did the agent do?
10:06Is it the right thing? Providing feedback and trying to steer it in the right direction because we can't run forever. Eventually, you need to steer it back. One half of that is the model provider, right? They're trying to build the best model. The other half is really about, I think, designing the best possible UI to enable the user to give feedback. And I think there's value in that. So I think there's a certain amount of stickiness in there. So what are all the different tools in terms of, say, take code editing, right? How can I most easily visualize what the code changes are? How can I most easily visualize what they impact?
10:31Which files? How can I, for small changes, get very quick feedback versus for complex ones, get complex feedback? There's some tools that actually draw diagrams for you of what they do, right? So I think this will be the battle. I think there's stickiness in that. How much exactly? So in that sense, people should be doing subscriptions to get people locked in, right? instead of moving to usage-based pricing?
10:51Erik Torenberg:Well, I think it's the customers that don't want to do usage-based pricing because it's so hard to guarantee. It's so hard for it to get away from them, and you actually want guarantees, and you're willing to commit to pretty high spend in order to not have usage-based pricing. I think it's the model companies that want usage-based pricing. I think with consumers, it's frankly very hard to not have usage-based pricing just because the variability is so massive, right? If it's us coding versus somebody who does this as their full-time job, right? You just have a factor of 20 or so difference in usage.
11:20That costs a lot of money, right? I think for enterprises, we could see more flat fee pricing because we can average it out more.
11:27Erik Torenberg:You have a developer that's using it all day, you kind of know in a general sense how many hours a day they're programming and what that sort of looks like. The vibe quotas are harder. Before we leave OpenAI, I want to ask a broad question, which is if Sam Altman was sitting here and saying, hey, Dylan, I'll listen to anything you tell me to do, any advice you have, as long as it makes OpenAI more valuable, what would you tell him? I would say immediately launch a method for you to input your credit card into ChatGPT and agree that for anything it like agentically does for you, it'll take X cut and then launch that product because where it does shopping, right?
12:01Because like everyone knows that like Anthropic and OpenAI and all the other labs are buying RL environments of Amazon and of Shopify and of Etsy and of all the different ways to shop on the internet of airline websites, right? Now, just like, hey, integrate my calendar. I want to fly to there on Thursday. Make sure I don't miss a meeting. Cool, book, right? Do that integration like super well. Know my preferences on whether I like aisle or window, all this stuff, right? And just take a take rate. I think this will make them so much money the moment they launch it. And I think they're working on it already, but I'd like to hear how he thinks about it because he shifted his tone massively on like ads over the last six months, right?
12:38He used to be like, no way. And now he's like, maybe, you know, There's a way to do it without harming the user. And I think this is how you monetize the free user, right? So I think that's probably what I'd tell him slash ask him about. Like a whole line of questions around this.
12:53Erik Torenberg:Well, he's coming on the podcast in a few weeks, so we'll ask him. I want to shift to NVIDIA. NVIDIA is having a monster year. They're up almost 70%. What are the possible paths from here? How do you see it playing out? Depends like how piled you are on like the continued growth. But I think you guys have a good vantage point. We have a good vantage point of how fast revenue is growing for a lot of these companies, especially the code companies, but even many other applications. I think we can clearly see the demand side is accelerating, right? And then if you look at the training side, I think the race is on.
13:23That is upping hugely. Google's upping hugely. If you just look at, again, just OpenAI and Anthropic and the compute that they have and are getting this year from Google and Amazon for Anthropic and from Microsoft, CoreWeave, Oracle for OpenAI, 30 % of the chips are going to them, just those two companies. But that's actually like, okay, well, like 70 % of the stuff, like who's making off? Well, one third of it is like ads, right? Whether it be Bike Dance or Meta or many of the other people who are doing ads. So then it's still like, okay, well, where are the rest of these one third of the chips coming from?
13:55Well, they're like mostly uneconomic providers who I don't think it's like an obvious bet that they're going to keep raising bigger and bigger rounds. So what happens there? I think with the, you know, we talked about like coding, right? like earlier, actually the Quen Coder 3 model is actually super cheap if you're running it on-prem or if you're running it in the cloud with all these inference libraries. And so there's stuff like that as well. So I think the question is, how much does it keep growing? Because clearly, I think the first third is definitely skyrocketing, right, of open AI anthropic lab spend.
14:26The second third of ads is going to grow. It's not going to grow like crazy, but I think there's definitely an inflection point that could be hit with Gen AI ads. I know Meta's been experimenting with it a lot, but I could totally be convinced that there's going to be a huge inflection. Take right there, where you start showing me personalized ads. Every person that's an ad looks like me, and I'll be like, okay, yes. Except slightly better, so I feel better. I have no idea how this is going to scale. But if you ask the question, how much could it scale? How much value are we creating here? Can we create enough value to actually keep growing for a long time?
14:58If you just take AI software development, we know we can easily get about 15 % more productivity out of a developer. I don't think that's right. I think it's way higher. No, no. But the straight, like, I talk to a lot of enterprises, like a classical enterprise, straight up GitHub Copilot deployment, that gives you about 15%. We can do much more than that.
15:14Dylan Patel:But bro, like, you know how bad GitHub Copilot is? Like, how did they, look at the revenue ARR chart. It's so funny. It's so funny. If you look at the revenue ARR chart, it's like, Cloud Code in three months has surpassed that. Cursor, you know, easily surpassed that. And then like, even like companies like Replin or like, and Windsurf slash Cognition are like, gonna pass them. I'm like, it's like, you're preaching to the choir. What's going on? So look, let's assume we can get this to 100%. Yeah. So we can double the productivity of a developer, right? About 30 million developers worldwide, give or take.
15:44Yeah. Right, let's say 100K value add per developer. This might be a little high worldwide. The US is low, but worldwide is high. So it's$3 trillion. Yeah, yeah. Right, so we're probably building technology here which adds$3 trillion of GDP value. In theory, we could put that into GPUs because that's the main cost factor here. Just from a coding model. Just from a coding model. Ignoring every other use case. So at least in theory, the value generation is here to keep growing, right? How that translates to the industry is a much more complicated.
16:08Dylan Patel:I think we've already seen AI's value creation. So sort of there's like the whole like the famous like, oh, 300 billion problem or 200 billion problem now. It's 600 billion problem, I'm sure. So Q is going to put out like the$1.8 trillion problem, right? Soon enough. But like there is some like reality in that, of course. But, you know, it ignores that like infrastructure spend today is accounting for five years of revenue, not like one. and the revenue looks like this, not like flat line. But I think the main thing is that AI is already generating more value than the spend. It's that the value capture is broken.
16:45I legitimately believe OpenAI is not even capturing 10 % of the value they've created in the world already just by usage of chat. And I think the same applies to Anthropic and Cursor and whoever else you're looking at. I think the value capture is really broken Even internally, I think what we've been able to do with four devs in terms of automation, our spend on Gemini API is absurdly low, and yet we go through every single permit and regulatory filing around every single data center with AI. And we take satellite photos of every data center, and we're able to label our data set and then recognize what generators people are using, what cooling towers and the construction progress and substations.
17:28All this stuff is automated, and it's only possible because of Gen.AI. And we do it with very few developers, and then the value capture that I'm able to generate by selling this data, by consulting with it, is so high, but the companies making it, they get nothing out of it. There's a value capture challenge here that far out exceeds the creation. And as you get models like GPT-5 or open-source models continuing to drive it down, it's like the value capture is just harder and harder and harder for these companies because they're making 50 % gross margin on inference or less in many cases. In so many words, you're saying, we're getting commoditized and therefore you can't capture the value and thus you should temper your expectations of how much you can spend on GPUs.
18:16Well, no, I think there's still ways to inflect hugely on value capture, right? I mentioned the ads are a huge value capture.
Read the full transcript
18:25Erik Torenberg:But that needs to happen before we see a massive increase. No, I think the other thing is there's a lot of capital that's not been spent. The hyperscalers still can grow CapEx 20-30 % next year from what they're doing this year. In addition, companies like CoreWeave and Oracle, because they're tapping capital markets, can raise way more than 20-30 % CapEx. And then you go down the list further and it's like, oh, the largest infrastructure funds in the world, like Brookfield and Blackstone, well, actually, they're turning all of their eyes to investing even more into infrastructure, AI infra. And then you're like the sovereign wealth funds of the world, like the G42s or, you know, the Norway one or GIC in Singapore, like these people have barely started touching AI.
19:12And so I think there's a whole lot more CapEx that can come without it being necessarily like economically motivated day one. I'm also saying like economically motivated CapEx can only grow like so much, but there's so much other like, like where it's not clear from, you know, if you have a spreadsheet, you know, and you're basing it on real business that you should actually spend this much. But people will because they believe. I believe, I think you believe, like Infra, you know, people believe that this will be, you'll get profit out of it, but there's no like 100 % certain like, you know, way to argue it.
19:47Erik Torenberg:Yeah. How threatened is NVIDIA by custom Silicon? I think that's the biggest thing, right? Is when we look at orders from Google and from Amazon, right, especially, and Meta, their custom silicon is, not Microsoft, their custom silicon kind of sucks. But the other three, they're really upping their orders massively over the last year. You know, Amazon is making millions of Tranium. Google's making millions of TPUs. TPUs clearly are like 100 % utilized, right? Yeah. Trinium's not there, but I think Amazon will figure out how to do that, and Anthropic will. So I think that's the biggest threat to NVIDIA, is that people figure out how to use custom silicon more broadly.
20:35And this sort of becomes this sort of like, if AI is concentrated, then custom silicon will do better. And that's not even talking about OpenAI's silicon team and stuff, right? Like, if AI is really concentrated, then they'll do better, custom silicon. but if it gets dispersed broadly because there's all these open source models from China and there's all these open source software libraries from, you know, NVIDIA and China, and it makes the deployment costs like rock bottom, then potentially. Hear me out here. If Google's TPU is able to compete with NVIDIA, in theory it could do it on the open market.
21:12NVIDIA is worth more than Google these days. Shouldn't Google start selling the chips to everyone? I mean, in theory, they should be able to achieve a higher market cap.
21:18Dylan Patel:I absolutely think so. I think Google's even discussing it internally. I think it would require a big reorg of culture and a big reorg of how Google Cloud works and how the TPU team works and how the JAX software team and XLA software teams work. I totally think they could. It would just take them shaking themselves pretty hard to be able to do it. But I totally think Google should sell TPUs externally. Not just renting, but physically. It's kind of funny if a side hobby, in theory, has a higher company value potential as your main product. Than your entire business,
21:58Erik Torenberg:especially as you think about the degradation of search as a core business.
22:01Dylan Patel:Yeah, but I think if you were to ask Sergey, like, hey, do you think selling chips and racks is more valuable, or a cloud, or a Gemini? he'd be like, no, no, no, no, no. Like Gemini is going to be worth way, way, way more. It's just not yet today, right? And so I think like today you say NVIDIA is the most, again, it's like a whole concentration thing, right? If the world is super concentrated in terms of customers, then NVIDIA will not be the most valuable company in the world, right? But if it gets dispersed more and more, which arguably we're starting to see with a lot of these open source models getting better and better and better and with the ease of deploying them getting better, then you would see, I think you could argue, NVIDIA will remain the most valuable company in the world for a long period of time.
22:51Historically, no pun intended software has eaten the world in most markets, right? If you look at early networking days, Cisco was the most valuable company on the planet for a while. It's no longer, right? They're the guys that build services on top, like Google or Amazon or Meta eventually eclipsed. Which is why NVIDIA is making all these software libraries.
23:10Dylan Patel:and they're trying to commoditize inference you guys don't I think even have an inference API provider investment do you? Well we have all kinds of model providers I'm talking about a pure API provider investment is that correct? I think I talked to one of the team members maybe Radko or someone about why you guys didn't invest in a Together or a Fireworks and sort of the argument was just serving models alone without making them will sort of be commoditized. We have some in the stable diffusion ecosystem. With FAL, yeah. It's a little bit different dynamics there, I think. They tend to make much more compound models than the LM folks.
23:55Yeah, yeah.
23:57Dylan Patel:But you guys don't have one of these base 10 or any of these sort of API investments because you think, this is from someone on the infra team, that you guys think it'll get commoditized because the software NVIDIA is making, because VLM and HDLang, which is like, open source software coming out of Berkeley and now sort of has their own environments now and supported by many, this being commoditized means that API providers aren't necessarily worth a ton, right? Is sort of your argument, maybe. I think that's relevant to this whole thing, which is, you know, why, right? Like, why would you do this?
24:30Shifting gears, what about the silicon startups? What's your take on those? I mean, there's a ton of capital flowing into that. We've seen, I have not numbers, but probably billions being invested in chip startups.
24:44Dylan Patel:Yeah, for sure, for sure. I mean, whether you're looking at companies like, I think it's pretty impressive that a few companies like Etched and Revos and a number of other companies, Madax and others, have gotten the amount of funding they've had without even launching a chip, right? In the past, yes, silicon companies would make money or raise money, but they would at least launch a chip before they get a big round, but Etched and Revos have raised a lot of money without ever launching a chip publicly, which I think is, I mean, it speaks to, well, yes, silicon is super capital intensive if you're building a chip, especially an accelerator, which has so many moving pieces.
25:28Dylan Patel:And there's like 10 different AI accelerator companies out there, right, that are newish in the last few years. I think there's a lot more. That are like, yeah, that's fair. and then there's the old guard which continues to raise money like Grok and Cerebris and Salmonova and Tenztorin and so on and so forth or Graphcore getting bought out by SoftBank and SoftBank dumping money into this effort as well there's a lot of capital being invested to dispel sort of NVIDIA's top dollar or top position but it becomes challenging it's like how do you beat NVIDIA the hyperscalers I think are kind of lucky in that they can do mostly the same thing as NVIDIA they can just win on supply chain I'm using cheaper providers it's a margin compression exercise essentially and maybe for certain workloads like Meta for recommendation systems they can specialize more but for the most part it's like no we're targeting the same workloads we can just simplify supply chain or in-house a lot of it and compress margin and it'll be fine But in the case of these other companies, it's like, well, they don't have a captive customer.
26:40Dylan Patel:So now you have to contend with, well, I'm using the same ecosystem and either I can use some custom silicon provider who's going to take a margin anyways on top and that's going to compress what I can sell for or I can try and in-house everything. But then it's like, this is really hard, right? I'm going to do all the silicon design. I'm going to build all this different IP. I'm going to manage the supply chain on chips, on racks, on everything, right? Ends up being a huge effort in terms of team size. All in the end, like, hey, I make a 75 % gross margin as NVIDIA. AMD sells their GPUs for 50 % gross margin, and they have a hard time out engineering NVIDIA, and they're great at engineering, right?
27:25Dylan Patel:But yet they still take more silicon area, more memory to achieve the same performance, and they have to sell for less, so their margin gets compressed. That makes sense. Look, I think historically, if you look at it, typically, if a new entrants in markets didn't win by marginally improving on something existing. That happens sometimes, but more likely, they jumped on some kind of disruptive technology leap, right? Whereas, like, we have a different approach, we have different technology. Is that possible here? I mean, to some degree, maybe it's over-simplifying a little bit, but I think part of the reason why the transformer model won was because it runs so incredibly great on GPUs, right?
27:59Like a recurring neural network is similarly performing, it looks like, but it runs terribly on a GPU. So did we sort of pick the model for an architecture? And now it's hard to come up with an architecture that really...
28:11Dylan Patel:Well, it's hardware software co-design, right? Exactly. There's all this hype about neuromorphic computing, right? Theoretically, it's amazing and super efficient. It's like, okay, great. There's no ecosystem of hardware. There's no ecosystem of software. It would take tens of thousands of people who are the best AI today focusing on that to even prove out if it's worthwhile or not, right? On a hardware side, on a software side, on a model side. And so you look at Grok, Cerebris, Samanova, they all sort of over-indexed to the models that were leading at the time when they designed their chips.
28:44Dylan Patel:And so they made certain trade-offs, right? They put a lot more memory on chip. And NVIDIA was like, well, we're not going to do that. A lot faster, at least, right? Well, more, like if you compare the amount of memory of SRAM on NVIDIA's chips, it's much, much lower than... Yes, correct. They went SRAM instead of DRAM. But then they usually have less DRAM, so there's a trade-off there as well. Right, there's less DRAM, there's more SRAM, and because there's more SRAM on the chip, you have to have less compute on the chip. And so they ended up losing, right? Because the model sizes got too big and all this, right?
29:13Dylan Patel:And so you have this super weird dynamic where they bet on something that was actually better, right? Like, I have no doubt that Cerebris would run certain types of models better than NVIDIA or GroK or, hey, Dojo, right? Dojo runs certain, you know, Tesla's Dojo would run certain types of models way better than NVIDIA's chips because they're optimized to that. But then it's like, oh, well, actually, even in vision tasks, you use vision transformers now. So it's like, okay, cool. Gives model sizes grew and all these things. So it ends up being a, you know, catch 22 in that, like, you optimize for something.
29:45Dylan Patel:And so now, like, today you have this new age of AI accelerator companies. They're like, okay, we're going to optimize for transformers. But the time they started designing, they're like, okay, transformers are dense models that are this big. what's the best, you know, the hidden dimension is 8K and your batch sizes are this big and your sequence points are this big, so let's just make a super large systolic array so you can, you know, create the maximum efficiency and then it turns out, oh, look at DeepSeq or, you know, go look at what the labs are doing. Actually, their shapes are much smaller.
30:12Dylan Patel:Actually, you need to do a bunch of small matrix multiplies, not massive, massive, massive, you know, singular matrix multiplies per layer. And then it ends up, you know, oh, well, that chip you're designing for that is actually not super effective for that. And so the software is evolving constantly because of what works best on NVIDIA. And you see that with, you know, whether it be what DeepSeq's doing or Alibaba's doing or what the labs are doing internally. And you even see this like for Google, right? Like their open source Gemma models make different decisions because the shapes of a TPU are different than a GPU.
30:46And those, the GPU and the TPU are actually not that far apart, right? Like you would say, yes, they're very different, but Blackwell and TPUs are converging on similar designs, actually. Whereas to be NVIDIA, you can't just have the supply chain with. You don't have this captive customer. So now you need to do something that will give you 5x advantage in hardware efficiency for a certain type of workload and then pray the workload doesn't shift. Because NVIDIA is also optimizing their architecture generation. They've added a lot of stuff to make their chips way better for the existing models, but it's like they're taking large steps every year, every two years towards something, whereas you have to go way over there in left field and hope that models stay over there, right?
31:32Dylan Patel:Because you have to win by 5x because NVIDIA is going to have supply chain efficiency over you. They're going to have time to market over you in terms of a new process node or new memory or whatever technology, right? Even AMD, right? They got to 2 nanometer before NVIDIA. They had higher density HBM. they use 3D stacking, all these things on supply chain that should be better than NVIDIA, and yet they still lose. They're still the software angle. NVIDIA is fantastic. Yeah, and then there's software as well, right? But it's like, NVIDIA is going to have better networking than you. They're going to have better HBM.
32:04Dylan Patel:They're going to have better process node. They're going to come to market faster. They're going to be able to ramp faster. They're going to have better negotiations with whether it's TSMC or SK Hynix and the memory and silicon side or all the rack people or like copper cables, everything, they're going to have better cost efficiency. So you have to be like 5X better. But to be fair, If somebody had a viable competitor, which would even be marginally cost competitive, if my guess is many of the big consumers of GPUs would immediately shift some revenue there just to have a number tool, right? Just to turn into a tool.
32:30That's AMD today, right? And Microsoft stopped. Somewhat, yeah.
32:33Dylan Patel:I mean, like... There's still pretty limited traction, though, right? Sure, but... Yeah. Meta continues to buy from them, and Microsoft did buy a bunch, and then they stopped, because it's like, well, yes, they're, you know, AMD's giving you all these advantages, but it ends up still not being better on a performance per watt basis, and they have a way bigger software team, that are somewhat competitive on all these dynamics that I mentioned. So you can't just do the same thing as NVIDIA. And do it better, or try and execute better like AMD. You have to really leap forward in some other way. But the design cycle takes so long that models will shift.
33:07Because they're like, oh, what's the next generation of TPU and GPU look like? Okay, let's optimize for that. And the research path is great.
33:15Dylan Patel:Yes, neuromorphic computing could be the most optimal thing for us to do, but no one's working on that because you have to advance in the tech tree you've chosen. If you restart the tech tree, you're going to be like, well, this sucks. And so if it branches this way and you're over here, you're screwed. Because you have to be 5x better. Because the supply chain stuff means that 5x actually turns into a 2.5x and then NVIDIA can compress their margin a little bit if you're actually competitive and then that 2.5x becomes like a 50 % better. And then, yeah, so it's like it ends up being way too difficult to in the software stuff, right?
33:48Dylan Patel:Everything like takes your 5X and makes it like, oh, you're actually only 50 % better.
33:52Erik Torenberg:And defense supply chain for sure.
33:53Dylan Patel:Yeah, defense supply chain. And then like they get that, right? Like, so it's like, and Lutnik himself said we had to do this for rare earth minerals. And it's like interesting. China, there's like provinces in China that have like rules that say the H20 is not efficient enough to be deployed, which is like super bizarre because it's clearly the best AI chip China has. Huawei is still a little bit behind.
34:18Erik Torenberg:Well, what's interesting is that, you know, efficiency is just not, is so much less of an issue in China than here because they just have the power infrastructure to be able to support. So even if they're running less powerful chips, you know, you would imagine that it doesn't really matter because China has just such an infinite supply, infinite supply of power that, you know, they'd sort of be okay with it. So it's interesting. Which is, it's a big challenge in America, right? Like there have been, there have been companies that were like,
34:46Dylan Patel:they would, they've like, you know, Jensen keeps saying he couldn't give away H20 in America for free, but I've literally heard companies now say, yeah, no, I wouldn't because I only have this much power. How am I going in data centers ready to go over the next year? If I bought an H20, I'd literally have less compute capacity and then I'd lose, even if it was free. It doesn't make sense. Whereas China doesn't care. They can build these things. They have the muscle. I'm curious how this all shakes out. you know, China's posturing really hard. They even like put out something that was like, we're investigating to see if there's backdoors in the H20.
35:22Dylan Patel:It's like, there's no backdoor in the H20. Like chill. You know, it's like, you know, GPU, GPU is usually like firewall from the public internet anyways. Like you step through stuff before you get to the GPU clusters. So like a backdoor wouldn't even matter. I don't know. I think it'll be interesting to see because China can definitely deploy way, way, way more power to AI the moment they decide to. But there's these, like, there's, like, competing interests, right?
35:55Erik Torenberg:Because they want Huawei to be better than NVIDIA. Yeah, and then this is how NVIDIA argued to the administration. They're like, if we don't do this, actually, I think it's, like, a very, like, powerful argument that, like, for example, within Triton, which is a common ML library, anyway, like, ByteDance has open-sourced some stuff
36:12Dylan Patel:that plugs into this that is super awesome. And there's all these other libraries. It's not just models that China open sources. It's software for NVIDIA that Chinese companies open source. In a sense, by NVIDIA selling GPUs, as NVIDIA's argument again, they were able to stop Huawei from building up a software ecosystem and the Western ecosystem is better. But then on the flip side, it's like, again, if you believe the models deliver more economic value to society than the hardware, which I actually think they do, it's just there's a value capture problem today, then you're giving China way more by giving them H20s and soon a version of Blackwell that's cut down, like Trump said, right?
36:55Versus selling them the chips, right? The economic value derived from selling them the chips is not as large as being able to somehow sell them AI services.
37:02Erik Torenberg:So is China gatekeeping power for AI? I don't think so. I think, again, there's a lot of like...
37:12Dylan Patel:What we see is that even with H20 being sold into China and future versions of the chip, H20E and other chips, we still see Chinese companies like Alibaba renting GPUs outside of China because the GPUs they can get outside of China are just so much better on a dollar spend per performance basis. Renting them or even going through a Singaporean company that is effectively a Chinese company and building data centers and putting chips in them. So it's like, I don't think China's limiting the power per se. It's that it's, you know, you can only, if you can spend, like Chinese companies are growing their CapEx way more than U.S.
37:51Dylan Patel:companies on a percentage basis next year. The absolute dollar number is, you know, obviously the U.S. companies are spending more still on AI. The percentage basis, Chinese companies are growing more next year. And you still have the problem of like, well, dollars spent to AI output in tokens or in whatever is going to be lower because these chips are worse. So power is not the gating factor. It's always capital, right? At least today, right? Now China can spend a lot more capital if they wanted to. They're subsidizing the semiconductor industry to the tune of like$150,$200 billion a year through SOEs, through CapEx, that's not generating revenue, et cetera.
38:29Dylan Patel:So it's not like they couldn't do this to the AI ecosystem, right? Given Meta's CapEx is like$60 billion, right? And Google's CapEx is like$80 billion, right? They could totally spend way more than that on a single effort. They just haven't decided to. And I just think, for the US, our build-outs are constrained by power. Google has a ton of TPUs sitting, waiting for data centers to be powered and ready. As does Meta with GPUs. We posted about how Meta is now building these effectively tents. Isn't this to some degree also coupled to their unwilliness to sell them to a broader ecosystem? If they want to be confined in their own data centers and didn't ramp data center build out for their own hyper-sacrifices quickly enough, then yes, that constrains them.
39:16If they were on the open market, would we still be constrained?
39:19Dylan Patel:Yeah, for sure. Because companies like CoreWeave, why is CoreWeave valuable? It's really because they build infrastructure really fast. And their software is nice, I think, but a lot of their customers are bare-model. Just replace the GPUs whenever they're broken and networked it properly. They grew more aggressively. And they'll go anywhere. Jensen likes them as well, to be fair. Yeah, that's very important as well. Because it unconcentrates the ecosystem, which is better for NVIDIA. Having worked at Intel, I know exactly what Chip was going through his mind. Yeah, so I think what's really important is that CoreWeave doesn't care.
39:53Dylan Patel:They're like, oh, crypto data center, I will convert it to AI data center. They bought a company for$10 billion that's doing crypto mining, which is worth$2 billion a couple years ago, and it's not because their Bitcoin mining business is growing. It's because they have power data centers. Anywhere and everywhere, people are trying to build power data centers. And companies like CoreWeave and Oracle are moving to that. Actually, today, Google just bought 8 % of a crypto mining company called Terrowoof.
40:25Erik Torenberg:Not because they're getting into crypto mining.
40:27Dylan Patel:No, because they need the data centers. They want the power. They need the power. And it's like all the hyperscalers have said, screw off to my sustainability pledges because they need power as fast as possible. They're doing things that are not, they take a little bit longer to move the ship, but even if you didn't do it in your own self-built data centers, there's still a lot of challenges in the open market. There's a deficit.
40:56Dylan Patel:and that's constraining American chip buildouts heavily yes others could maybe do it a little bit faster like CoreWeave or others Oracle's got an open mind as well but it's still constraining US buildouts heavily even though the capital's been spent the chips are 60-80 % of the cost of the cluster depending on what chips you're getting so it's like they've already bought the chips They just can't put them anywhere because the data centers aren't ready. So it applies to Google, applies to Microsoft, applies to Meta, applies to a lot of folks.
41:28Erik Torenberg:I mean, it's really hard to build power infrastructure in the U.S.
41:32Dylan Patel:Power, grid interconnections, transmission, substations, all of this stuff. Like electrical contractors, electricians in Texas, if you're willing to be a travel electrician, it's like oil pay, right? It used to be that if you're physically adept, you could go make$100 ,000 in West Texas. but like who the fuck wants to do that? Now it's like, well, you could go like 200 miles away from Dallas and what's still a reasonable town and build a data center and work on the wiring within the data center and all this other stuff, the transmission stuff, and your pay is up like 2X now versus what it was just a few years ago.
42:09This labor problem is a challenge too. And it's, yeah, I think in China, they don't have any of these problems, but they just haven't spent the capital yet. But capital is an issue as well. because of the scale of what's being spent, right? Like NVIDIA's revenue this year is going to be like over$200 billion and next year it expects over$300 billion plus Google's going to spend like$50 billion on TPU data centers, right? And Amazon's going to spend tons and tons on Tranium data centers. It's like the scale of dollars is quickly growing to nation-state level stuff. And what's more important is being able to decide to spend the dollars and what's cost-effective.
42:49And so to some extent, China's still constrained by that, but they can smuggle chips in, they can build data centers outside of China, they can rent data centers outside of China and have the most cost-effective, you know, Blackwell chips or whatever, right? ByteDance is, you know, either the biggest or the second biggest customer of Google Cloud for a reason, right? And they're getting, you know, over, you know, they're getting many, many Blackwell from them, right? And the same with Oracle and the same with Microsoft and all these other companies are renting tons of chips to China anyways because it's more cost-effective to do that than build it yourself.
43:22So it's not like China has this mentality where we only have to, well, the government does, but the infrastructure companies don't, like Alibaba, Tencent, ByteDance, et cetera. So what's the end game for data centers? I mean, we need more power, we need more cooling. Will the end be, all data centers will be next to a nuclear reactor or lots of solar, you know, next to a deep level, like deep seawater that we use for cooling or something like that? or what's...
43:48Dylan Patel:I think that cooling is... The physical cooling of a data center, there's this whole narrative about AI uses so much power and it's not really... Farming alfalfa uses 100x the water of AI data centers. Even by the end of the decade, it'll be the same and alfalfa is worth very little. It's like cooling is not that... People have experimented with undersea data centers to reduce the cooling cost. That doesn't make sense. It's like 5-10 % savings, but then like... It's easier to get the water out of the ocean then, then put the data center into the ocean. It's like if you want to service it, like you're screwed, right?
44:23Dylan Patel:So like the same with power, it's like we talk a lot about like the power is not actually that expensive. It's just hard to build, right? And hard to get to the right place. And hard to get to the right space and convert it down to the voltages and all the stuff that chips need.
44:37Erik Torenberg:So it's less the magnitude of power and more where it is and how it moves.
44:40Dylan Patel:Well, the magnitude too, right? Like it's going to be... It's just a total worldwide energy consumption. AI data centers are still... It's a fraction of a percent. Yeah, yeah. Even by the end of the decade, the U.S. will be like 10 % of our power will be AI data centers, which is still like... Of electricity. Of our electricity. In terms of energy, that's even a smaller fraction, right? Oh, yeah, yeah, because you think about... By shifting to electric vehicles, we can probably make a bigger swing than with all the AI data centers we can build. But outside, it's like in Europe, that number's not moving up that fast and all these other countries.
45:13Dylan Patel:I think... We need to build a lot more power, but it's not like some crazy, crazy amount. It's just like doing it properly is the hard thing. And again, the cost of power, you go look at these deals people are signing, they're still signing, even though the price has skyrocketed from a few cents a kilowatt hour for these massive, massive purchases to 10. It's still, when you think about the full TCL, the cluster, the GPU cost, the networking, All of this stuff far outstrips the power. And same with cooling. What percentage is power? Like if I do a four-year amortized GPU data center, what percentage would be power?
45:54Dylan Patel:80 % of the cost of a GPU data center if you're building Blackwell is capital. It's the GPU purchases, it's the networking, it's the physical data center power conversion equipment. All of this stuff is like 80 % of the cost. And then 20 % is going to be your LAN and your power and your cooling. and your cooling towers and your backup power and your generators and all this stuff. It's like nothing, which is why it doesn't matter if you spend 10 % or 50 % more on that. Because at the end of the day, the expensive thing, right? Like this is why what Elon did would seem silly, right? They spent a lot more money on generators outside the data center and these mobile chillers to cool the water down for their liquid cooling instead of like the more cost-effective option because it got the data center up three months faster.
46:42Dylan Patel:And so that three months of additional training time is worth way, way, way more on a TCO basis. The performance you got out of the chips and the time to market and all this is way, way faster. And therefore, it was the right decision, even though this part of the data center bloomed at cost. Everything else is still there, and you're still paying for the chips. And if they were sitting idle, it's not worth it.
47:03Erik Torenberg:Just by bypassing the grid, bypassing anything to do with interconnect, anything to do with public utilities. Exactly, exactly. What's your take on Intel? Where is Intel going? I think the world, well, the U.S. needs Intel. I think the world needs Intel. I think the world needs Intel because, like, Samsung is doing worse than Intel on leading-edge process development, in my opinion, based on even on various customers in the industry having done test chips at, like, Intel versus Samsung. They think, I think the industry generally agrees that Intel is further along, you know, sort of the two-nanometer class process technology than Samsung is.
47:38But both are way behind TSMC.
47:41Dylan Patel:And TSMC is a monopoly in some extent. The number one question always people ask is, why is TSMC not making more money? Why are they only raising prices next year 3 % to 10 % depending on what it is? It's like TSMC is a monopoly. They could raise a lot more, but they're good Taiwanese people rather than dirty American capitalists. If TSMC was owned or was managed by Americans, I think most ownership is actually American in terms of the stock market. It's on the New York Stock Exchange and all this. They would have raised prices a lot more. And so there is this difficult, difficult thing to be done that like, hey, there's one island that controls all leading-edge semiconductors and not just all leading-edge, like the majority of trailing-edge production as well.
48:27Dylan Patel:Something needs to be done. Intel is behind, but not like absurdly so, right? Like if something were to happen to Taiwan, Intel would have the most advanced technology in the world, right? It's just, it's not economic. Can you keep Intel as one company if you want them to be competitive? I think the process of splitting it would take so much executive time and so much executive effort that you would have been bankrupt by then. Right? And that's the big challenge. I think Intel should be separate, right? But to properly split the company and for all the management time that's needed is absurd. And instead, what you need is you need Lip Bhutan, who's the CEO of Intel.
49:05Dylan Patel:There's a lot of drama going around about him because he's one of the greatest semiconductor investors ever, right? He's invested in so many different companies first. You know, he was on the board of like SMIC, which is China's TSMC effectively, which is like a big like drama or like some of the biggest tool companies in China. He's the first investor in them because, you know, it was a multipolar world there and he's making good investments. But like, you know, now like people are getting mad about that, but it's like, no, he recognizes the companies, like he understands the supply chain. He needs to not spend his time on splitting the company because then he never actually fixes the company, right?
49:39Dylan Patel:Intel's problem is that like, it takes them five to six years to go from design to shipping the product, in some cases more. And when they tape out a chip, right? Like, you know, you send the design to the fab, the fab brings back the chip. They go through 14 revisions in some cases, whereas like the rest of the industry goes through like one to three, right? Revisions, if they're good, of like send the design in, get the chip back, test it, send the design in for a public launch. And they'll launch a chip in three years. So, but if you look at Intel today, right, they still don't have a competitive entry on the AI side.
50:16And they won't. Can you, so what does it mean for their offering? I mean, they're still doing great on CPUs. They don't have a good AI chip product. Is it long-term sustainable positioning, right? I mean, as a standalone chip company?
50:30Dylan Patel:I mean, IBM still makes more money every launch off of mainframes. So it's not like x86 is dead. it's like you don't get the growth rates, but you could totally run this as a very profitable enterprise. And I think the same with PCs, right? There's some turmoil, there's some arm entry, there's some AMD competition. Well, I think it can be a very profitable business if it had one third the people or half the people working on it. And so Liputan, to fix Intel, needs to go into both the design company and lay off a shitload of people, but keep all the good people and make sure that they're designing fast and they're launching from design conception to launch is two to three years, not five to six.
51:10Dylan Patel:And that's on the design side and make that profitable. And then on the fabs, he has to do the same thing. There's all these people, like one of the heads of fab automation at Intel. I explicitly told Lip Bhutan because we have a couple ex-Intel people who are actually good in the company that worked on the fab side. And we're like, who's the worst people and friends? It's like, oh, this guy sucks. I explicitly told Lip Bhutan. He had never talked to the guy because it was like four layers down. The company has like absurd amounts of hierarchy. It's like four layers down. He goes and talks to the guy and he's out, right?
51:41Dylan Patel:It's like, he figures out like who's bad, right? And who's good. And he has to go in and he's still like, hey, the vast majority of the team at Intel is the one who led the world in production and process technology for 20 years, right? But there's a lot of like built up crap. So he has to go figure this out, right? He can't waste his time on like, oh, all this like structuring to split. I think it would be better if the company split. I just don't think he can spend the time to do that. And if the design side of the company is, you're not really going to get into AI, you have to make some money there, but the fabs, I think, could truly become a competitor.
52:18But they're going to go bankrupt by the time anything can happen, so they have to figure out how to get capital. So he has to figure out how to get capital, he has to figure out how to clean up all the crap, make the yields go up, make the product ship way faster. Like, all of these things are basic problems. I think the goals are completely correct. I mean, I think the big challenge, just reflecting back on my time there, right? I think the big challenge is that right now, if you look at Intel, right, they have essentially software, sort of the chip design making, and then there's, you know, the core manufacturing part, right?
52:47And they have three very different cultures. And it's very hard to get everything under one umbrella, right? And so I think that is the big challenge. I think you should even run the company separately, right?
52:56Dylan Patel:But, like, you can't physically separate them entity-wise because it's going to take so long to sever all these things because he doesn't have time, right? Like Intel is literally going to go bankrupt if they don't have a big cash infusion or they lay off like half the company, right? Which some could argue you need to lay off like 30 % of the company anyways, but there's a lot of bad things that happen if that happens, right? And they need to spend a lot more on building the next generation FAB even if they fix the FAB and they don't have money for that, right? So there's like a lot more important problems than like physically separating the company, even though I think long-term, yes, the fab has to be separate from the chips design software, right?
53:34Dylan Patel:Or chip design part of the company. Just like that's going to make each company much more accountable, be able to service their customers better, et cetera. It's just that's going to take too long and they're going to go bankrupt by then. Awesome. But I think, I hope, I hope, I pray someone does something, right? Like you get a big capital infusion, I don't know. The big hyperscalers are like muscled into like, oh, okay, wait, if TSMC eventually grows their margin to 75 % because of the monopoly, plus they intake all this stuff like co-package optics and power delivery and all this all of a sudden the cost is going to spike so we should actually just throw$5 billion at Intel each screw it and that could actually give Intel enough of a lifeline to potentially get to something and maybe be competitive that's the hope
54:18Erik Torenberg:Can we finish by finishing this game that we started when we gave Altman advice if Jensen was here, what advice would you have for him?
54:28Dylan Patel:if Jensen was here you know I think he has a massive massive balance sheet right Jensen does free cash flow is like ridiculous the tax cut the new Trump tax bill institutes something really incredible which is that you can depreciate all of the GPU cluster cost in year one which we put out like a note about how like the tax implications to Meta are like$10 billion a year. And across each of the major hyperscalers, it's massive. It's like, well, NVIDIA's going to spend tons and tons of cash or they're going to spend tens of billions of dollars of taxes. Why don't you get into the infrastructure game somehow?
55:12Now, this is obviously going to be crazy because now they're buying their own GPUs and putting them in data centers and doing stuff, and they're competing with their own customers, but they're already doing that anyways because their customers are trying to make chips.
55:25Dylan Patel:but they should accelerate the data center ecosystem with investments, right? Because really we think we can have very high degree of accuracy on what they're going to do next year in terms of revenue because it's just the number of data center watts that are being built, right? This is a harder thing to shift up and down, right? Now there's a little bit of share difference between how much is TPU versus GPU but it's like you have to accelerate the infrastructure and you need to spend all of this capital that you're building, right? Like, okay, do you want to go the route of like doing buybacks and dividends?
55:56Dylan Patel:Like, great. Like, you're a loser if you do that, right? Like, you can make more money by reinvesting and building a bigger company that's not just chips into the ecosystem or servers into the ecosystem, but actually like controlling the infrastructure end to end somehow. So I think there's something he could do there with this massive war chest. And there's a reason like NVIDIA has done some buybacks and they've done some dividends and increasing. but the cash on their balance sheet keeps growing and they're going to have north of$100 billion of cash on their balance sheet by the end of this year, I think.
56:27So it's like, what are you going to do with that? I think there's something moving into the infrastructure layer much more that they could do if he really wants to be the king of the world, right? Which I think he does. Sergey and Sindar? Cool.
56:45Dylan Patel:I think they should open up the Kimono on TPUs, right? Start selling them. Open up the software. Open source a lot more of the XLA software because there's open XLA and there's XLA, but the vast majority is closed source. Really, really open up the kimono on that and be a lot more aggressive, right? They're still pretty not aggressive on data centers. They're pretty not aggressive on a lot of elements of the company. The TPU team's next-gen designs are pretty not aggressive, partially because a lot of the TPU team has left to go to OpenAI, the best people that I knew. It was actually really annoying.
57:21Dylan Patel:I knew like four people or five people and they all went to OpenAI. And it's like, fuck, now I don't get as much. I met some other people, right? But it's like, I think they could be a lot more aggressive in many ways across the company. They don't have to be, right? But they could. Because AI, this ChatGPT, TakeRay, the shift of search queries, the monetizable ones, especially from two purchasing agents is going to really screw Google long-term if they don't get their act together. I think they've gotten their act together on DeepMind. There's still some inefficiencies, but Sergey works within DeepMind a lot and they're driving hard.
57:55Dylan Patel:They're still a little bit behind, but I think physical infrastructure, TPU, and how much money they could make and how much they could take the wind out of everyone else's sails if they start selling TPUs externally and reorg around building data centers much faster so that they do have the most compute in the world, because they did, but now there's certain companies that are going to surpass them potentially over the next few years if they don't really get their act together. So I think that's what I would say for them. Yeah, and also learn how to ship product. Zuck? I think Zuck, it remains to be seen what goes on with superintelligence, but they're trying to move super fast with the data centers.
58:40Screw it, we'll build tents instead of physical data centers because we only need these for five years anyways. The superintelligence moves, you could say,
58:48Dylan Patel:whatever you want, but like, you know, trying to buy like Thinky for like 30 billion or SSI for 30 billion didn't work out. So then they spent, you know, not even that much on hiring, not 30 billion on hiring all these people. So I think that he recognizes the urgency with the models, with the infrastructure. So I really think he needs to like, you know, if you read his website post about like AI, like I think, you know, he sees the vision, right? There's the wearables, is there's integrating AI into that. There's being your AI assistant to do all this purchasing and stuff. I think he sees the vision, but I think he also needs to focus on actually releasing that faster, but also the products that they do outside of their core IP every time they launch something is kind of mid, right?
59:34Dylan Patel:Metal Reality Labs is doing well, but I think they should go more explicit, like have a chat GPT competitor, have a cloud code competitor, just start releasing way more products because they're really just focused on their individual gardens rather than branching outside of it.
59:51Erik Torenberg:Do you think Apple should have that same sense of urgency or if Tim Cook was here, what would you tell them?
59:56Dylan Patel:The funny thing is some of their best AI people are now at Superintelligence. They're building an AI accelerator. They have AI models, but they're just way slower. They did mention on the last earnings call they're going to allocate more capital to this, but it's like, guys, Apple, You guys are going to lose the boat if you do not spend $50,$100 billion on infrastructure. You don't think the concierge will cut it? I think more and more you'll see people like, great, Apple has this walled garden, but they can only do so much to protect it. IDFA, they shut down data sharing to Meta, but Meta made better models and now they have way more data and way more power over the user than they ever did before.
1:00:38Dylan Patel:It was good that Meta kicked the crutch off of them, or Apple did. But the same applies to like AI. Like, yes, they have access to the text and they have access to this. But like, I think other people are going to be able to integrate user data and agents will be able to integrate all this user data and they'll start to lose control of what the user experience is as more and more gets disintermediated by AI being the interface rather than touch, rather than, you know, touchpad and keyboard. And I don't think they've truly realized what happens when the interface to computing is AI. They market it, but that's going to shift computing really heavily.
1:01:16Dylan Patel:They have great hardware, and their hardware teams are working on awesome stuff and form factors. But I just don't know if they get what is actually going to happen to the world in the next five years truly well enough, and they're not building fast enough for it.
1:01:29Erik Torenberg:What about Microsoft to that end?
1:01:32Dylan Patel:Microsoft has the same problem. I think they were super aggressive in 23 and 24. And then they pulled back heavily, right? Now, like, OpenAI is slipping through their grasps. There's that whole thing there. They cut back on data center investments heavily. They were going to be the largest infrastructure company in the world by, like, a factor of 2x, which would have been, you know, you could argue maybe that was, like, too much, and maybe it wouldn't have been economical. But, like, they're losing grasp on OpenAI. Their internal model efforts are failing spectacularly. Like, they're on LLM arena right now, and they're pretty decent there, but that's just a sycophantic model.
1:02:10It's a code name, but whatever. MAI is failing.
1:02:14Dylan Patel:Azure is losing a lot of share to Oracle and CoreWeave and Google and so on and so forth. Their internal chip effort is by far the worst of any hyperscaler. They're just mis-executed. How is GitHub not the highest ARR software code model? They only had the best IDE, the best source code repository, the best enterprise, Salesforce, the best model company as a relationship, and they were the first to market, right? It's like they had everything going for them. And there's just nothing, right? It's like GitHub Copilot is failing. Microsoft Copilot is still crap, right? Yeah, it's unusable. It's like, what is going on?
1:02:57Dylan Patel:You need to shake the crap out of the company. I think they win a lot because they have the best business-to-business relationship with so many enterprises. Best Salesforce on the planet. Yeah, but they end up not having the actual product to sell them, which is really scary. So they need to really work on product. Satya has done great on sales and stuff, but yeah.
1:03:16Erik Torenberg:If Elon was here, what advice would you give him?
1:03:19Dylan Patel:A lot of people at XAI are mad about the porn models, like porn stuff. It's fine, you're going to make a ton of money off of this. This is how you accelerate the revenue of that company. But he's losing a lot of talent and axing a lot of good projects. But Elon is a magnet to amazing talent and building stuff, so I won't bet against him, but it seems like since he left the administration and focused on stuff again. But I think, I don't know, I think he's focused on a lot of things, and I think, like, RoboTaxi, starting to look good, actually, again. Like, I haven't ridden one yet, but I have some friends who've ridden one.
1:03:48Dylan Patel:It, like, looks pretty decent. He could, like, not make these snap decisions, which often are the reason why he's amazing, but, like, some of these snap decisions are hurting him. I'm not sure if I can give Elon that much great advice because I think maybe it's just, like, get off Twitter and focus on like the products again, right? More. But he is working on that stuff a lot. Yeah.
1:04:10Erik Torenberg:I think that might be a good place to wrap. This was a great discussion. Dylan, thanks so much for joining us.
1:04:14Dylan Patel:Thank you for having me. Thank you.
1:04:17Erik Torenberg:Thanks for listening to this episode of the A16Z podcast. If you liked this episode, be sure to like, comment, subscribe, leave us a rating or review and share it with your friends and family. For more episodes, go to YouTube, Apple Podcasts and Spotify. Follow us on X at A16Z and subscribe to our Substack at a16z.substack.com. Thanks again for listening, and I'll see you in the next episode.
1:05:02Erik Torenberg:For more details, including a link to our investments, please see a16z.com forward slash disclosures.
From the publisher
As part of our summer replay series, we're revisiting one of our favorite conversations on the future of AI infrastructure.
SemiAnalysis founder Dylan Patel joins Erin Price-Wright, Guido Appenzeller, and Erik Torenberg to examine the rapidly evolving economics of AI hardware, from GPUs and custom silicon to data centers, power, and the global race for compute.
The conversation explores NVIDIA's competitive advantages, the rise of custom chips from Google, Amazon, and Meta, the economics of frontier AI models, and the infrastructure constraints shaping the industry's next phase. They also discuss AI startups, export controls, robotics, enterprise software, and why simply copying NVIDIA isn't enough to build a winning AI hardware company.
Whether you're building AI products, investing in infrastructure, or trying to understand where the industry is headed, this conversation offers a practical look at the forces shaping the future of compute.
Resources:
Follow Dylan Patel on X: https://x.com/dylan522p
Follow Erin Price-Wright on X: https://x.com/espricewright
Follow Guido Appenzeller on X: https://x.com/appenz
Learn more about SemiAnalysis: https://semianalysis.com/dylan-patel/
Stay Updated:
Find a16z on YouTube: YouTube
Find a16z on X
Find a16z on LinkedIn
Listen to the a16z Show on Spotify
Listen to the a16z Show on Apple Podcasts
Follow our host: https://twitter.com/eriktorenberg
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
