Dylan Patel — Deep dive on the 3 big bottlenecks to scaling AI compute

13 Mar 2026 · 2 h 31 min · 61 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Dwarkesh Podcast Episode Notes

Episode Title

Dylan Patel — Deep Dive on the 3 Big Bottlenecks to Scaling AI Compute

Guest

Dylan Patel, Founder of SemiAnalysis

---

Overview In this episode, Dylan Patel provides an in-depth exploration of the three major bottlenecks in scaling AI compute: logic, memory, and power. He also discusses the economics of labs, hyperscalers, foundries, and fab equipment manufacturers, offering insights into the intricacies of the semiconductor supply chain and the future of AI compute.

---

Key Themes and Discussions

  1. Three Bottlenecks to Scaling AI Compute
  2. Logic: The challenges in producing advanced chips at scale. The dependency on tools like EUV (Extreme Ultraviolet Lithography) from ASML.
  3. Memory: The anticipated incoming memory crunch, largely due to the increased demands from AI workloads. The cost dynamics and production difficulties of different types of memory (e.g., HBM vs. standard DRAM).
  4. Power: Although scaling power in the U.S. seems manageable, there are complexities in delivering sufficient energy to support increasing semiconductor production.
  1. Economic Insights
  2. The combined forecasted capital expenditures (CapEx) for major tech companies (Amazon, Meta, Google, Microsoft) reaching approximately $600 billion, reflecting future AI compute scaling.
  3. The difference in compute needs and spending patterns between companies like OpenAI and Anthropic versus traditional tech firms.
  1. Trends in AI Compute Demand
  2. Discussion about the rapid growth of compute requirements for inference at AI labs and the economic implications for memory and chip suppliers.
  3. The increasing costs of memory and how it affects hardware performance and overall AI infrastructure deployment.
  1. Future Projections
  2. Projections for AI compute capabilities indicate potential power needs reaching 200 gigawatts by 2030.
  3. Challenges related to labor and skills in executing large-scale infrastructure projects for AI, with a focus on the need for specialized workforce training.
  1. Risks and Opportunities in the Semiconductor Supply Chain
  2. The geopolitical risks associated with Taiwan's semiconductor dominance and implications for U.S. supply chains.
  3. The potential for China to outscale the West in semiconductors if current trends continue.

---

Key Takeaways

  • Bottlenecks in scaling AI compute are interconnected, with logic and memory being crucial for powering the next generation of AI workloads.
  • Economic Forces: The financial viability of AI companies will increasingly rely on their ability to secure compute resources efficiently and affordably amid rising prices.
  • Labor Shortages: The need for skilled labor in semiconductor manufacturing and AI infrastructure will create challenges as demand increases.
  • Future of AI: The path towards advanced AI will be shaped by the balance of supply capabilities and the ability to manage demand effectively across the compute landscape.

---

Sponsor Messages

  1. Mercury - A financial platform streamlining invoicing and tax processes.
  2. Labelbox - A tool for diagnosing performance issues in voice models using their new evaluation pipeline, EchoChain.
  3. Jane Street - A research-driven trading firm with extensive GPU and storage infrastructure.

---

Timestamps

  • 00:00:00 – Why an H100 GPU is worth more now than three years ago
  • 00:24:52 – Nvidia securing TSMC allocation early, Google facing pressure
  • 00:34:34 – ASML as a constraint for AI compute scaling by 2030
  • 00:55:47 – Discussion on utilizing older TSMC fabs
  • 01:05:37 – Predictions for China’s semiconductor capabilities
  • 01:16:01 – Anticipating a memory crunch
  • 01:42:34 – Power scaling in the U.S. and implications
  • 01:54:44 – Discussion on the feasibility of space GPUs
  • 02:14:07 – Insights on hedge funds and AI investments
  • 02:18:30 – Potential for TSMC to cut Apple out from N2
  • 02:24:16 – Risks involving Taiwan in semiconductor supply

---

Conclusion This episode provides a comprehensive overview of the key challenges and economic factors influencing the future of AI compute. Dylan Patel's insights into the bottlenecks and implications for the industry allow listeners to grasp the complexities of scaling AI infrastructure in the coming years.

---

For more information and to access the full transcript, visit [Dwarkesh Podcast](https://www.dwarkesh.com).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Big Tech's CapEx and AI Compute

0:45 to 2:30

Discussion on the combined CapEx of major tech companies and the implications for AI compute.

“how to think about the timeline around when that CapEx comes online.”

Understanding Compute Timelines

2:30 to 4:00

Exploration of when the CapEx comes online and the significance of funding by AI labs.

“all these other things that they're doing for further out into the future so that they can set up this super fast scaling, right?”

Hyperscaler Dynamics and Revenue Implications

4:00 to 6:30

Analysis of hyperscalers' compute capacities and the revenue growth of AI labs.

“You know, Dario, when he was on your podcast was very, very like conservative.”

Anthropic vs OpenAI: Compute Strategies

6:30 to 9:30

Comparison of compute acquisition strategies between Anthropic and OpenAI.

“Who built the spare capacity such that it's available for Anthropic and OpenAI to get last minute?”

The Costs of Acquiring Compute

9:30 to 12:00

Discussion on the costs and implications of acquiring compute under different conditions.

“And, you know, there's a tradeoff there.”

Depreciation Cycles and GPU Utility

12:00 to 14:01

Insight into GPU depreciation cycles and their impact on AI compute economics.

“Michael Burry was saying it's, you know, three years or less, right, is like sort of his argument.”

The Value Dynamics of AI Chips

14:01 to 17:01

Explore how advancements in AI models like GPT-5.4 affect the valuation of GPUs and computing resources.

“The price of a hopper would fall at a spot or a short-term contract rate as the new chips come out and the price per performance goes up.”

Compute Demand and Market Dynamics

17:51 to 19:39

Discuss the inconsistencies in AI compute demand and the implications for companies like OpenAI.

“banking services provided through Choice Financial Group and Column NA members FDIC.”

The Economics of AI Compute

19:40 to 24:42

Delve into how increasing compute costs are affecting the pricing strategies of AI models and cloud providers.

“So to give a specific example, suppose the better tasting apple costs$2 and then the shittier apple costs$1.”

NVIDIA's Market Leverage

24:43 to 28:00

Examine NVIDIA's strategic position in the semiconductor market and its impact on AI chip allocation.

“There's no way they can continue, Anthropoc can continue at the current pace without destroying demand.”
Show all 61 chapters

NVIDIA's Unique Position in AI Compute Supply

28:00 to 29:00

Explore how NVIDIA navigates supply chain dynamics to secure AI compute resources.

“And NVIDIA is a bit unique because, oh, yes, they have CPUs.”

Google's Delays and Strategic Decisions

29:00 to 30:50

Understand the strategic missteps by Google in the evolving AI compute landscape.

“And so in that case, there was a huge sort of like, OK, well, these guys are delaying, but NVIDIA is wanting more, more, more, more, more.”

Anthropic's Unexpected Compute Acquisition

30:50 to 33:10

Learn about Anthropic's proactive approach to securing AI compute resources.

“Google sold, I think, a million, was it the V7s, the Ironwoods, to Anthropoc.”

Shifting Bottlenecks in AI Compute

33:10 to 35:00

Discuss how the bottlenecks in AI compute have evolved over time.

“And you can see this from Google's Gemini ARRs, right?”

The Role of Semiconductor Fabs in AI Scaling

35:00 to 37:00

Examine the critical role of semiconductor fabs in the AI scaling process.

“It switches back from being power and data center as a major bottleneck to chips.”

Impact of EUV Tools on Chip Manufacturing

37:00 to 39:50

Understand how EUV tools influence the manufacturing capacity for AI chips.

“So to scale compute further, right, there's some different bottlenecks this year, next year.”

Projected AI Compute Growth by 2030

39:50 to 42:00

Explore the projections for AI compute growth and the implications for the industry.

“So when you think about what it's doing across a wafer, it's taking the wafer and it's scanning and it's stepping across, right?”

Scaling AI Compute Demand

42:00 to 43:20

Discussion on the projected AI chip demand and production capacity.

“And what does that imply about the – Sam Altman says he wants to do a gigawatt a week in 2030.”

Advancements in EUV Tools

43:20 to 46:01

Exploration of the evolution and capabilities of EUV tools in semiconductor manufacturing.

“When the 7nm started, so I don't know when that was exactly.”

Detailed Mechanics of EUV Machines

46:01 to 48:24

In-depth explanation of the components and functionality of EUV machines.

“And for simplicity's sake, we're kind of ignoring the advances for this podcast, the advances in overlay or throughput per tool.”

Complex Supply Chains in Semiconductor Production

48:24 to 51:32

Analysis of the intricate supply chains involved in producing semiconductor machinery.

“And then it's in this thing that is like basically collecting all the light and directing it into the lens stack, right?”

Challenges in Meeting AI Compute Needs

51:32 to 54:33

Discussion on the challenges faced by the semiconductor industry in meeting AI compute demands.

“And it has to be accurate to the level of single digit nanometers or even smaller because the entire system, the overlay, right?”

Exploring Semiconductor Supply Chain Alternatives

56:00 to 57:29

Learn about alternative approaches in semiconductor production when faced with bottlenecks.

“these are the alternatives that people fall back on.”

Performance Differences in Chip Design

57:30 to 59:28

Understand the nuances of performance comparisons between different chip architectures and designs.

“But then there's a similar amount for seven nanometer, right?”

Data Movement and Its Impact on Performance

59:29 to 1:02:10

Discover how data movement between chips affects overall computational efficiency.

“And most of this is not attributed to flops.”

Future of Semiconductor Production in China

1:02:11 to 1:06:06

Examine the potential advancements in China's semiconductor production capabilities by 2030.

“So when you look at inference at, let's say, 100 tokens a second for DeepSeq and Kimi K2.5, Hopper versus Blackwell, the performance difference is on the order of 20x.”

Long-term Perspectives on AI and Semiconductor Dominance

1:06:07 to 1:10:02

Contemplate the implications of AI advancements on global semiconductor leadership by 2035.

“chain indigenized rather than having random suppliers in Germany and Netherlands and whatever would mean that China would be ahead in its ability to produce mass flops.”

China's Semiconductor Dominance and AGI

1:10:02 to 1:10:44

Exploration of potential future scenarios regarding AGI and China's role in semiconductors.

“Should you expect a world where China is like dominating in semiconductors, which I think, I don't know, doesn't get asked enough in San Francisco.”

Estimating Future AI Compute Needs

1:10:44 to 1:11:33

Discussion on tracking data centers, compute capacity, and challenges in long-term forecasting.

“But the time lags for these things are relatively short, right?”

Competitive Landscape in AI Lab Compute

1:11:33 to 1:12:31

Analysis of competition between US and Chinese AI labs over compute capacity.

“I think Opus 4.6 and GPT 5.4 have really pulled away and made the gap a little bit bigger, but I'm sure some new Chinese models will come out.”

Capital Expenditure in AI Infrastructure

1:12:31 to 1:13:46

Insights into the capital expenditure trends in the US and the implications for AI growth.

“At some point, they end up getting to a point where the model performance should start to diverge more.”

The Divergence of US and China's AI Capabilities

1:13:46 to 1:15:49

Exploration of how infrastructure investments affect AI capabilities in the US versus China.

“If and when Anthropic 10x is revenue again, and I think our answer would be when, not if, then China doesn't have the compute to deploy at that scale.”

Memory Constraints in AI Development

1:15:49 to 1:17:57

Discussion on the implications of memory technology constraints on AI performance.

“But I don't know, like, I don't know what fast timelines means, right?”

Trade-offs in Memory Technology for AI

1:17:57 to 1:20:56

Examination of the trade-offs between using HBM and DDR in AI applications.

“And yet they don't because no one actually wants to use a slow model.”

Impact of Memory Prices on Consumer Electronics

1:20:56 to 1:24:00

Analysis of how rising memory costs affect consumer electronics pricing and market dynamics.

“What is the bandwidth difference between HBM and a normal DRM?”

Impact of Memory Prices on iPhone Costs

1:24:00 to 1:25:16

Learn how fluctuations in memory prices can affect iPhone pricing and margins.

“If you look at the bill of materials of an iPhone, what fraction of it is the memory?”

Trends in Smartphone Sales and Market Dynamics

1:25:16 to 1:26:48

Understand the implications of decreasing smartphone sales on the market.

“Actually, that's only a few hundred million phones a year, right?”

Shifts in Memory Demand and AI Impact

1:26:48 to 1:28:11

Discover how the shift in memory demand affects AI and consumer electronics.

“And Apple's volumes will not go down as much as like a low-end smartphone provider.”

Constraints in Memory Production and Industry Dynamics

1:28:11 to 1:30:16

Explore the challenges in memory production and the associated industry changes.

“And in fact, you've produced more memory for AI.”

Elon Musk's Fab Ambitions and Challenges

1:32:40 to 1:35:29

Examine the feasibility of Elon Musk's plans for rapid fab development.

“I interviewed Elon recently, and his whole plan is that I guess they're going to build this gigafab, terafab, some power of 10.”

Future of 3D DRAM and Manufacturing Challenges

1:35:29 to 1:38:00

Learn about the future potential of 3D DRAM and the manufacturing hurdles ahead.

“And some of these two other two companies aren't even that great.”

Retooling Fabs for New Process Nodes

1:38:00 to 1:39:24

Learn about the complexities and retooling needed in fabs to accommodate new semiconductor processes.

“You can't just like convert a logic fab to a DRAM fab or vice versa back and forth or a NAN fab to a DRAM fab in a short amount of time.”

Bottlenecks in AI Compute Scaling

1:39:24 to 1:41:38

Discover how supply chain limitations and tool availability affect AI compute scaling efforts.

“And for DRAM, it was in the mid-teens as well, or low-teens, and now it's trended towards the high-teens.”

Scalability of Power for AI Needs

1:41:38 to 1:43:19

Examine the potential to scale power generation to meet AI demands and the associated challenges.

“But all you're effectively doing is you're saying, ASML, you're dumb.”

Unlocking Additional Power Capacity

1:43:19 to 1:47:31

Learn about various methods and technologies that can unlock additional power capacity for data centers.

“So, I mean, right now we're at 30, right?”

Labor Constraints in Energy Scaling

1:47:31 to 1:51:55

Explore the labor challenges in scaling energy production and potential solutions to these issues.

“There's a lot of risks that people have to take.”

Modularization in Data Centers

1:52:00 to 1:54:38

Learn how modularization can streamline data center construction and operations.

“Humanoid robots maybe start to or robotics at least start to.”

Elon Musk's Space GPU Argument

1:54:39 to 1:55:38

Explore the challenges and considerations of building data centers in space.

“So speaking of big problems to solve, Elon Musk is very bullish on space GPUs.”

Permitting and Regulatory Challenges

1:55:39 to 1:56:40

Understand the complexities of permitting for data centers in the U.S.

“Because that's just what they thought they needed more of.”

Challenges of Space Data Centers

1:56:41 to 2:03:08

Discover the limitations and future potential of space-based data centers.

“So all of these things have plenty of room to be paid for.”

Power Density and AI Chips

2:03:09 to 2:05:59

Examine the effects of power density on AI chip performance and cooling.

“And so you want them deployed working on AI the moment they're done being manufactured.”

Understanding Scale-Up Domains in AI Hardware

2:06:05 to 2:08:41

Learn about the differences in scale-up architectures among NVIDIA, Google, and Amazon.

“And maybe it's at this point worth explaining what exactly a scale up is and what it looks like for NVIDIA versus Tranium versus TPUs.”

Challenges in Parameter Scaling for AI Models

2:08:41 to 2:12:34

Explore the factors affecting the scaling of AI models and the role of compute resources.

“So you can get the scale up to be hundreds or thousands of chips, but also have it not contend for resources when you're bouncing through chips.”

Resource Allocation Strategies in AI Development

2:12:34 to 2:14:06

Understand how compute resources are allocated in AI research and development.

“And then as you look to Google, Google does deploy the largest production model of any of the major labs, right, with Gemini Pro.”

Market Dynamics and Client Insights in AI Data

2:14:06 to 2:18:28

Gain insights into how clients use AI market data and the dynamics affecting their decisions.

“Why is Leopold the only person that is using your spreadsheets to make outrageous money?”

The Future of TSMC and Apple in AI Chip Production

2:18:28 to 2:20:00

Discuss the evolving relationship between TSMC and Apple amid rising AI demands.

“Can TSMC, if you're saying, look, the memory logic, et cetera, the N3 is mostly going to be AI accelerators, but then there's N2, which is mostly Apple now, and then in the future, I guess AI would also want to go on N2.”

Apple's Dominance in N2 Fabrication

2:20:00 to 2:21:05

Explore how Apple's control over N2 wafers impacts AI chip production.

“Yeah, I wonder if it's worth going to specific numbers.”

Huawei's Competitive Edge in AI Chips

2:21:05 to 2:22:38

Discuss Huawei's role in AI chip innovation and its advantages over competitors.

“Now with two nanometer, you've got AMD trying to make a CPU and a GPU chiplet that they used advanced packaging to package together in the same timeframe as Apple.”

The Future of Humanoid Robots and AI

2:22:38 to 2:26:28

Analyze the implications of humanoid robots needing local compute power.

“And so, you know, I mean, that's just moving to a process.”

Centralization of AI Intelligence

2:26:28 to 2:27:28

Understand the shift towards centralized compute for AI and robotics.

“because the power is really bad for robots, right?”

De-risking Taiwan's Semiconductor Supply Chain

2:27:28 to 2:29:48

Examine strategies to mitigate risks in Taiwan's semiconductor industry.

“I think Elon recognizes this, which is why he's like going to different places for his chips, right?”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Dwarkesh Patel:All right. This is the episode of My Roommate Teaches Me Semiconductors. It's also the send-off for this current set. Yeah. After you use it, I'm like, I can't use this again. I got to get out of here. No sloppy seconds for Dwarkesh. Okay. Dylan is the CEO of Semi Analysis. Dylan, the burning question I have for you, if you add up the big four, Amazon, Meta, Google, Microsoft, their combined forecasted CapEx that you published recently, this year is$600 billion. and given, you know, yearly prices of renting that compute, that would be like close to 50 gigawatts. Now, obviously we're not putting on 50 gigawatts this year.

0:40Dwarkesh Patel:So presumably that's paying for compute that is going to be coming online over the coming years. So I have a question about how to think about the timeline around when that CapEx comes online. Similar question for the labs where, you know, OpenAI just announced that they raised$110 billion. Anthropic just announced they raised$30 billion. And if you look at the compute that they have coming online this year, you should tell me how much it is, but like, is it not, is it not another four gigawatts total that they'll have this year? It feels like the cost to rent the compute that OpenAI and Anthropic will have this year to like sustain their compute spend at, you know,$10,$13 billion a gigawatt.

1:17Dwarkesh Patel:Those individual raises alone are like enough to cover their compute spend for the year. And then this is not even including the revenue that they're going to earn this year. So help me understand first, when is the timescale at which the big tech CapEx is actually coming online? And two, what are the labs raising all this money for if like the yearly price of a one gigawatt data center is like$13 billion?

1:41Dylan Patel:So when you talk about the CapEx of these hyperscalers, right, on the order of$600 billion, and you look at the cross the rest of the supply chain, gets you to on the order of a trillion dollars. A portion of this is, you know, immediately for compute going online this year, right? The chips and the other parts of CapEx that do get paid this year. But there's a lot of setup CapEx as well, right? So when we're talking about 20 gigawatts this year in America, roughly. Incremental. Incremental added capacity. A portion of this is not spent this year. A portion of that CapEx is actually spent the prior year.

2:17Dylan Patel:And so when you look at, hey, Google's got$180 billion, actually a big chunk of that is spent on turbine deposits for 28 and 29. A chunk of that is spent on data center construction for 27. A chunk of that is spent on, you know, power purchasing agreements and down payments and all these other things that they're doing for further out into the future so that they can set up this super fast scaling, right? And this applies to all the hyperscalers and other people in the supply chain. And so, you know, 20 gigawatts roughly deployed this year, a big chunk of that being hyperscalers, chunk not being.

2:51Dylan Patel:And all of these companies, their biggest customers are Anthropoc and OpenAI. Anthropic and OpenAI are in the, you know, two gigawatt and, you know, two and a half gigawatt and one and a half gigawatts roughly right now. They're trying to scale to much larger, right? If you look at what Anthropic has done over the last few months, you know, 4 billion, 6 billion revenue added. And if we just draw a straight line, hey, yeah, they'll add another $6 billion of revenue a month. People would argue that's bearish and that they should go faster. What that implies is that they're going to add$60 billion of revenue across the next 10 months, Right.

3:24Dylan Patel:And 60 billion dollars of revenue at the current gross margins that Anthropic had at least last reported by media would imply that they have roughly 40 billion dollars of compute spend for that inference for that 60 bill of revenue. That$40 billion of compute at roughly$10 billion a gigawatt rental cost means that they need to add four gigawatts of inference capacity just to grow revenue. And that's saying that their research and development training fleet stays flat, right? So, you know, in a sense, Anthropic needs to get to well above five gigawatts by the end of this year. And it's going to be really tough for them to get there, but it's possible.

4:02Dwarkesh Patel:Can I ask a question about that? So if Anthropic was not on track to have five gigawatts by the end of this year, but it needs that to serve both the revenue that's gone crazier than expected, and maybe it's going to be even more than that, plus the research and training to make sure its models are good enough for next year. Where is that going to come from?

4:20Dylan Patel:You know, Dario, when he was on your podcast was very, very like conservative. He's like, you know, I'm not going to go crazy on compute because if my revenue inflects at a different rate at a different point, I don't want to go bankrupt. You know, I want to make sure that we're being responsible with this scaling. But in reality, you know, he's definitely missed the pooch in terms of like going like OpenAI, which was let's just sign these crazy fucking deals. Right. And OpenAI is kind of got way more access to compute than Anthropic by the end of the year. And so what does Anthropic have to do to get the compute?

4:50Dylan Patel:Well, they have to go to lower quality providers that they would not have gone to before, right? You know, optimally, you know, Anthropic, at least historically, has had the best quality providers been like Google and Amazon. Whereas, you know, at least historically minded, you know, the biggest companies in the world now Microsoft and now they're expanding across the supply chain and going to other players that are newer. OpenAI has been, you know, a bit more aggressive on going to many players. Yes, they have tons of capacity from Microsoft. They have Google and Amazon as well, but they also have like tons with CoreWeave and Oracle.

5:20Dylan Patel:And they've gone to like random companies or, you know, one would think random companies like SoftBank Energy, who has never built a data center in their life. But, you know, they're building data centers now for OpenAI. So they've gone to and many others like Nscale and others that they're going and getting capacity from. And so there's this like conundrum for Anthropic because they were so conservative on compute because they didn't want to go crazy. Right. And in some sense, a lot of the financial freak outs in the second half of last year were like, oh, but I signed all these deals. They don't have the money to pay for them.

5:52Dylan Patel:OK, Oracle stock's going to tank. OK, Core Reef stock's going to take. OK, like, you know, all these companies stocks tanked and credit markets went crazy because people like the end buyer can't pay for this. Now it's like, oh, wait, they raised a ton of money. Okay, fine, they can pay for it. But in the sense, Anthropic was a lot more conservative. They were like, we'll sign contracts, but we'll be principled, and we'll purposely undershoot what we think we can possibly do and be conservative because we don't want to potentially go bankrupt.

6:17Dwarkesh Patel:The thing I want to understand is, so what does it mean to have to acquire compute in a pinch? Is it that you have to go with like NeoClouds? Is it that they have worse computers? Like in what way is it worse? And is it that you have to pay gross margins to a cloud provider that you wouldn't have otherwise had to pay to because they're coming in at the last minute? Who built the spare capacity such that it's available for Anthropic and OpenAI to get last minute? And like basically, what is the concrete advantage that OpenAI has gotten if they end up at similar compute numbers by 2027? Is it just like they're going to end this year with different gigawatts?

6:50Dwarkesh Patel:If so, how many gigawatts is Anthropik and OpenAI going to have by the end of this year?

6:54Dylan Patel:Yeah. So to acquire excess compute, I mean, yes, there is capacity at hyperscalers. And not all contracts for compute are long-term, right? Five years, right? There's compute that in 2023 or 2024, H100, 2025, that were signed at not five-year deals, right? OpenAI, the vast majority of their compute is signed at five-year deals. but they can, you know, there were, there were many other customers that had one year, two year, three year deals, six month deals on demand. And as these contracts roll off, who is the participant in the market most willing to pay price? Um, and in this sense, right, we've seen H100 prices inflect a lot and go up and people willing to sign long-term deals for, you know, as above$2 even, right?

7:38Dylan Patel:Like I've seen deals where certain AI labs, I'm going to be a little bit vague here for a reason, have signed at as high as$2.40 for two to three years for H100s, which if you think about the margin,$1.40 for Hopper when you release it or Hopper to build it across five years. And now two years in your signing deals that are two to three years that are at$2.40, those margins are way higher. And so now you can crowd out all of these other suppliers, whether it's Amazon had these or CoreWeave had these or Together AI or Nebius or whoever it is, right? These neoclouds are the firms that had a higher percentage of hopper in general because they were more aggressive on it, A.

8:22Dylan Patel:And B, they tended to sign shorter-term deals, not core weave, but the others tended to sign shorter-term deals. And so, hey, if I want hopper, there is some capacity out there. And then also, while most of the capacity at an oracle or core Weave is signed for a long-term deal in terms of Blackwell. Anything that's going online, this quarter's already sold. And in some cases, they're not even hitting all the numbers that they promised they would sell because there are some data center delays, not just those two, but like Nebius and all the other folks, Microsoft, Amazon, Google. But there is a lot of neoclouds, as well as some of the hyperscalers who have capacity they're building that they did not sell yet, or capacity that they were going to allocate to some internal use that is not necessarily super AGI focused that they may now turn around and sell.

9:06Dylan Patel:Or they may, you know, in the case of Anthropic, they don't have to have all the compute directly, right? Amazon can have the compute. They can serve Bedrock or Google can have the compute and serve Vertex or Microsoft can have the compute and serve Foundry and then do a revenue share with Anthropic or vice versa.

9:20Dwarkesh Patel:Basically, you're saying Anthropic is having to pay either this like 50 % markup in the sense of the revenue share or in the sense of last minute spot compute that they wouldn't have otherwise had to pay had they bought the computer early.

9:32Dylan Patel:Right. And, you know, there's a tradeoff there. but also at the same time you know for a solid like four months everyone was like OpenAI we're not going to sign deals with you like that sounds crazy right because you guys don't have the money now everyone's like yeah OpenAI we believed you the whole time we can sign any deal because you've raised all this money but in a sense Anthropic is constrained in that sense there are not that many incremental buyers of compute yet because Anthropic hit the capabilities here first where their revenue is mooning

10:03Dwarkesh Patel:That's interesting. Because otherwise, you're like, well, having the best model is an extremely depreciating asset that three months later you don't have the best model. But the reason it's important is that you can sign these deals and then lock in the compute in advance, get better prices. Doesn't this also imply, by the way, and maybe this is an obvious point, but there's at least until recently, people had made this huge point about, oh, what is the depreciation cycle of a GPU? and the bears, Michael Burrys or whatever, have said, look, people are saying that four or five years for these GPUs.

10:37Dwarkesh Patel:And in fact, if you, maybe it's because the technology is improving so fast or whatever, in fact, it makes sense to have two-year depreciation cycles for these GPUs, which increases the sort of like reported amortized capex in a given year. And so it makes it maybe financially less lucrative to building all these clouds. But in fact, you're pointing at like, maybe the depreciation cycle is even longer than five years because if we're using hoppers and then especially if ai really takes off and in 2030 we're like fuck we got to like get the seven nanometer fabs up and we got to like we got to go back to the a100s like turn on the a100s again uh then it's like actually the depreciation cycle is incredibly long and um uh so i think that's an interesting financial implication of what you're saying there's a few um strings to

11:23Dylan Patel:pull on there. One is, um, what happens to depreciation of GPUs? Right. Um, and, and I guess I didn't answer your prior question, which is like Anthropic, I think we'll be able to get to like five gigawatts ish, maybe a little bit more by the end of the year through themselves, as well as their product being served through Bedrock or through Vertex or through Foundry. Uh, I think they'll be able to get to five or six gigawatts, uh, which is way above their like initial plans, right? You know, and anyways, that's sort of like, and OpenAI will be roughly the same, maybe a little higher, actually a little bit higher based on our numbers.

11:59Dylan Patel:But anyways, the depreciation cycle of a GPU, right? Michael Burry was saying it's, you know, three years or less, right, is like sort of his argument. And there's sort of two ways and lenses to look at this. Like mechanically, in this, you know, there's a TCO model, right? Total cost of ownership of a GPU where we sort of project pricing out for GPUs and build up the total cost of a cluster. But there's a number of costs, right? There's your data center cost, right? There's your networking cost. There's your smart hands and people in the data center swapping stuff out. There's your spare parts, right?

12:30Dylan Patel:There's your actual chip cost. There's your server cost. All these various costs get lumped together and there's some depreciation cycles on it. You know, there's certain credit costs on it. And you get to, okay, that's how you build up, hey, an H100 costs$1.40 an hour to deploy at volume across five years if your depreciation is five years. And then if you sign a deal at$2 an hour for those five years, your gross margin is roughly 35%. It's a little bit above that. But if you sign it for$1.90, it's 35 % roughly. And then you assume at that fifth year, the GPU falls off a bus, right? It's dead.

13:03Dylan Patel:And in some cases, you know, sort of the argument people are making is, well, if you didn't sign a long-term deal because every two years NVIDIA is tripling, quadrupling the performance while only 2xing the price or 50 % increasing the price, then the price of an H100, sure, maybe the value in the market was$2 at 35 % gross margins in 2024. But in 2026, when Blackwell is in super high volume and deploying millions a year, you're actually now worth$1 an hour. And when Rubin in 27 is in super high volume, right, even though it starts shipping this year, is in super high volume next year, doing millions of chips a year deployed into clouds, you've got another 3x in performance and another 50 % or 2x in price.

13:44Dylan Patel:Actually, the hopper is only worth 70 cents an hour. And so the price of a GPU would continue to fall. That's like one lens. The other lens is what is the utility you get out of the chip, right? Because if you could build infinite Rubin or infinite of the newest chip, then yes, that's exactly what would happen. The price of a hopper would fall at a spot or a short-term contract rate as the new chips come out and the price per performance goes up. But because you are so limited on semiconductors and deployment timelines and all these things, you end up with actually what prices these chips is not, hey, what's the comparative thing I can buy today?

14:22Dylan Patel:It's actually what is the value I can derive out of this chip today, right? And in that sense, let's take GPT 5.4. GPT 5.4 is both way cheaper to run than GPT-4, has fewer active parameters. parameters, it's much smaller, right, in that sense of active parameters, plus because, you know, a sparser MOE versus GPT-4 being a coarser MOE. There's also been so many other advancements in training, RL, model architecture, et cetera, et cetera, data qualities, all these things that have made GPT-5.4 way better than GPT-4, and it's cheaper to serve. And so when you look at an H100, it can serve more tokens per GPU of 5.4 than if you had ran GPT-4 on it, right?

15:04Dylan Patel:So at some sense, it's producing more tokens of a model that is of higher quality. And so in some sense, obviously, GPT-4, what is the maximum TAM for its tokens? Maybe it was a few billion dollars, maybe it was tens of billions of dollars. Adoption takes time. For GPT-5.4, that number is probably north of$100 billion, but there's an adoption lag and there's competition. So other people are getting it. And there's the constant improvements that everyone else is having. So if improvements stopped here, the value of an H100 is now predicated on the value that GPT 5.4 can get out of it instead of the value that GPT 4 can get out of it and the margins and all that stuff that these labs are doing and they're in a competitive environment so their margins can't go to infinity.

15:44Dylan Patel:So you sort of have this like dynamic that is quite interesting in that an H100 is worth more today than it was three years ago. That's crazy.

15:52Dwarkesh Patel:And I mean, it's also interesting from the perspective of like, just take that forward. If we had actual AGI models developed, if we had like genuinely human on a server On a flop basis, an H100, these are such hand-wavy numbers about how many flops can the brain do. But on a flop basis, an H100 is estimated to 1E15 is how much some people estimate the human brain does in flops. Obviously, in terms of memory, the human brain has way more. H100 is like 80 gigabytes and brain might have petabytes.

16:23Dylan Patel:Oh, yeah, you've got petabytes? Name a petabyte of ones and zeros, bro. So name me a string.

16:31Dwarkesh Patel:Well, this is actually the point.

16:33Dylan Patel:No, we've just got the best sparse attention techniques ever.

16:36Dwarkesh Patel:Genuinely, right? In the sort of amount of information that is compressed, it might be petabytes, but it's extremely sparse MOE. But anyways, imagine if we had a human knowledge worker can produce six figures a year of value. And so if H100 can produce something close to that, if we had actual humans on a server, The value of an H100 is like it can repay itself in the course of like a couple of months. So as I've been going through everything to prep for taxes, I realized that I worked with over 50 different contractors last year, from cinematographers to audio technicians to editors. And I owed all of them$10.99.

17:12Dwarkesh Patel:In the past, I've just used a spreadsheet and a big folder of invoices to figure out who I need to collect tax forms from. But with so many contractors, this takes a bunch of time. And I've almost missed some people. This year, though, Mercury made my process way more straightforward. Whenever I pay somebody in 2025, I just hit a toggle to have Mercury request a double-9 from them. Because of that, everything that I needed to issue 1099s got sent directly to Mercury. I literally just clicked a button and Mercury generated and sent them all out. This is just one of the many things that I never would have assumed that a banking platform could just handle for me.

17:42Dwarkesh Patel:Mercury has a bunch of features like this, which are going to collectively save me multiple days this tax season. You can learn more at Mercury.com. Mercury is a fintech company, not an FDIC-insured bank. banking services provided through Choice Financial Group and Column NA members FDIC. So when I interviewed Dario, the point I was trying to make is not that I think this singularity is two years away and therefore Dario desperately needs to buy more compute, although the revenue is certainly there that he needs to buy more compute. But the point I was trying to make is that given what Dario seems to be saying, given his statements that we're two years away from a data center of geniuses, certainly not more than five years away, and data server geniuses should be earning trillions upon trillions of dollars of revenue.

18:26Dwarkesh Patel:It just does not make sense why he keeps making these statements about being more conservative on compute or to your point, being less aggressive than OpenAI on compute. And I guess that point got lost because then people were like roasting me about like, oh, this podcast was I try to convince this like multi-hundred billion dollar company CEO. Like, why don't you YOLO it, bro? But no, I was trying to say that internally his statements are inconsistent. anyway so it's good to iron it out yeah I think you know going back to like sort of the earlier

18:55Dylan Patel:view that if the models are so powerful the value of a GPU goes up over time as we approach closer and closer to you know let's say a point where right now only open ion and anthropic have that viewpoint as we approach further and further out actually everyone is going to even with open source models be able to like sort of like start to see that value skyrocket per GPU. And so in that sense, you should you should commit now to compute. But interestingly, in like an anthropic fashion, right, you know, there is a bit of a meme that they are they don't they have problems with commitment issues and they're like sort of polyamorous.

19:35Dylan Patel:Not Dario, but this is a bit of a meme.

Read the full transcript

19:39Dwarkesh Patel:Explains everything. By the way, so there's this interesting economics effect called Alkin-Allen, which is the idea that if you increase the fixed cost of different goods, one of which is higher quality and one of which is lower quality, that will make people choose the higher quality good on the margin. So to give a specific example, suppose the better tasting apple costs$2 and then the shittier apple costs$1. Okay, now suppose you put an import tariff on them. And so now it's$3 versus$2 for like great Apple, medium Apple, right?

20:15Dylan Patel:Is that because they both increased by$1 or should it be like 50 % increase?

20:18Dwarkesh Patel:No, no, because they both increased by$1. The whole effect is that if there's a fixed cost that's applied to both, the relative price, the price difference between them, the ratio changes. So previously it was like the more expensive one was 2x more expensive, now it's just 1.5x more expensive. So I wonder if applied to AI, that would mean that, look, if GPUs are going to get more expensive, there will be a fixed cost increase in the price of compute. As a result, that will push people to be willing to pay higher margins for slightly better models. Because the calculus is, I'm going to be paying all this money for the compute anyways.

20:55Dwarkesh Patel:I might as well just pay slightly more to making sure it's like the very best model rather than a model that's slightly worse. Right.

21:01Dylan Patel:So the hopper went from$2 to$3. And if a hopper can make a million tokens of Opus and it can make 2 million tokens of Sonnet, the price differential between Opus and Sonnet has decreased because the price of the GPU has increased by a dollar from$2 to$3. Interesting. I think that makes a ton of sense. Also, I think we just see all of the volumes are on the best models today, all the revenues on the best models today. and in a compute limited world, there's sort of two things that happen, right? A, companies that have locked up, you know, and don't have commitment issues, you know, have these five-year contracts for compute, they've kind of locked in a humongous margin advantage because they've locked in compute for five years at a price of what it transacted at five years ago or three years ago or two years ago, whatever it is.

21:51Dylan Patel:Whereas if you're now three years into that five-year contract and someone else's two-year contract or three-year contract rolled off, And now you're trying to buy that at, you know, modern pricing. When you're priced to the value of models, the price is going to be up a lot more. And so in a sense, like the person who committed early has better margins in general. And the percentage of the market that is in long-term contracts is much larger than the percentage of the market in short-term contracts that can be this sort of flex capacity that you add at the last second. And at the same time, right?

22:23Dylan Patel:So where does the margin go, right? Because models get more valuable. How much can the cloud players flex their pricing? Well, in fact, like if you look at CoreWeave, their average term duration is like over three years right now for like 90 % plus of their compute. It's over three years. And so they end up with this like conundrum of like, well, they can't actually flex price. But every year they're adding incrementally way more capacity than they had previously, right? Right. This year alone. Right. Meta is adding as much capacity as they had in the entire fleet of compute and data centers for all purposes for serving WhatsApp and Instagram and Facebook in 2022 and doing AI.

23:03Dylan Patel:Right. They're adding that alone this year. So in the same sense, you know, you talk about Meta's doing that CoreWeave and and Google and Amazon, all these companies are adding insane amounts of compute year on year on year. That new compute gets transacted at the new price. So in a sense, yes, you've locked in, as long as we're in a sort of a takeoff, right? Oh, OpenAI went from 600 megawatts to two gigawatts last year, from two gigawatts to six plus this year, and six to 12 next year, right? The incremental added compute is where all the cost is, not the prior long-term contracts. So then who holds the card is the infra providers for charging margin, right?

23:37Dylan Patel:So now the cloud players, the neoclouds, or the hyperscalers can charge the margin. Oh, they can't because, or they can to some extent, But then as you go upstream to, oh, well, who has access to all the memory and logic capacity? Well, it's NVIDIA for the most part. They've signed a lot of long-term contracts. They've got like$90 billion of long-term contracts today, and they're negotiating three-year deals with the memory vendors today. You've got obviously Amazon and Google through Broadcom and Amazon Directly and all these companies, sort of AMD. These companies hold all the cards because they've secured the capacity.

24:10Dylan Patel:and TSMC is not raising prices, but memory vendors are just like sort of to some extent raising a lot of price, right? So they're going to double or triple price again, but then they're also signing these long-term deals. So who is able to accrue all the margin dollars is actually, you know, potentially the cloud, potentially the chip vendors and the memory vendors until TSMC or ASML like break out and they're like, no, actually we're going to charge a lot more. But at the same time, do the model vendors get to charge crazy margins? I think at least this year, we're going to see margins for the model vendors go up a lot, right?

24:41Dylan Patel:Because they're so capacity constrained, they have to destroy demand, right? There's no way they can continue, Anthropoc can continue at the current pace without destroying demand.

24:51Dwarkesh Patel:Yeah. Let's get into logic and memory.

24:56Dwarkesh Patel:How specifically NVIDIA has been able to lock up so much of both? So I think according to your numbers, by 27, NVIDIA is going to have like 70 plus percent of N3 wafer capacity or something like that or around that area. And then I forget what the numbers were for memory at SK Hynix and Samsung and so forth. But if you look at, so think about how the NeoCloud business works and how NVIDIA works with that or how the RL environment business works and how Anthropic works with that. In both those cases, NVIDIA is purposely trying to fracture the complementary industry to make sure that they have as much leverage as possible.

25:34Dwarkesh Patel:So they're giving, you know, allocation to random neoclouts to make sure that there's not one person that has all the compute. Similarly, Anthropic or OpenAI, when they're working with the data providers, they say, no, we're going to just seed a huge industry of these things so that we're not locked into any one supplier for data environments. And I wonder why on the three nanometer process, that's going to be Tranium 3, that's going to be TPU v7, other accelerators potentially, and why is TSMC just giving it all up to NVIDIA rather than, you know, trying to fracture the market?

26:07Dylan Patel:Yeah, so I think there's a couple like points here, right? On three nanometer, you know, if we go back to last year, the vast majority of three nanometer was Apple, right? Apple is being moved to two nanometer, memory prices are going up, so Apple's volumes may go down, right? Because as memory prices go up, they have to either they cut margin or they move on. You know, there's some time lag because they have long-term contracts, But basically, Apple likely reduces demand slash moves to two nanometer faster, where two nanometers only capable of sort of mobile chips today. And in the future, AI chips will move there.

26:41Dylan Patel:So sort of Apple has that. And then Apple is also talking to third party vendors because they're getting squeezed out of TSMC a little bit. Because TSMC's margins on high performance computing, HPC, AI chips, et cetera, is higher than it is for mobile because they have a bigger advantage in HPC than they do in mobile. But anyways, when you look at what's TSMC running calculus here, actually, they're providing really good allocations to companies that are doing CPUs, right? So when you think about, hey, Amazon has Tranium and Amazon has Graviton, both of those are on three nanometer, Graviton being their CPU, Tranium being their AI chip.

27:18Dylan Patel:they're actually tsmc is much more excited to give allocation to graviton than they are to tranium because they view cpu business as more stable long-term growth right and as a company that is conservative and doesn't want to ride cycles of growth too hard you actually want to allocate to the uh the market that is more stable and lower growth rate first before you allocate all the incremental capacity to the fast growth rate market now that is that is the case generally And so when you look at like, hey, same for AMD, right? The allocations they get on, you know, their CPUs is like TSMC is much more excited about those than they are for GPUs.

27:59Dylan Patel:Likewise for Amazon. And NVIDIA is a bit unique because, oh, yes, they have CPUs. Yes, they make switches. Yes, they make networking. They make NVLink. They make all these different InfiniBand, Ethernet, all these different products, Nix. By and large, most of these things will be on three nanometer by the end of this year with the Rubin launch and all the chips that are in that family. the GPU being the most important one. And yet NVIDIA is getting the majority of supply, right? Part of this is because you look at the market and you like sort of like, you know, TSMC and others, like there are many ways that they forecast market demand.

28:34Dylan Patel:But also it's market signal, right? The market signaled, hey, we need this much capacity next year. We need this much. We need this much. We'll sign non-cancelable, non-returnable. We may even pay deposits, right? Things like this. NVIDIA just did it way earlier than Google or Amazon. And in some cases, Google and Amazon had stumbling blocks. You know, there was one of the chips got delayed slightly by a couple quarters, Tranium and all these sorts of things happened. And so in that case, there was a huge sort of like, OK, well, these guys are delaying, but NVIDIA is wanting more, more, more, more, more.

29:07Dylan Patel:And we are checking with the rest of the supply chain. Is there enough capacity? Right. So they're going to all the PCB vendors and they're saying, hey, is there enough Victory Giant? Is there enough PCB? This is like one of the largest suppliers of PCBs to NVIDIA and they're a Chinese company. All the PCBs come from China, sort of from them or many of them. And anyways, they're like, do you have enough PCB capacity? Great. Oh, hey, memory vendors, who has all the memory capacity? Okay, NVIDIA does. Great. So when you look at sort of in the same way, you know, who is AGI-pilled enough to buy compute in long timelines at levels that seem ridiculous to people who aren't AGI-pilled, but nonetheless, they're willing to pay a pretty good margin and sign it now because they view in the future that that ratio is screwed up.

29:50Dylan Patel:The same thing happens with the supply chain for semiconductors, right? But NVIDIA was, well, I don't think NVIDIA is quite AGI-pilled, right? You know, Jensen doesn't believe software is going to be automated fully and all these things, right?

30:00Dwarkesh Patel:Accelerated computing, not AI chips, right? It's AI chips, right? But that's what he calls it, right?

30:05Dylan Patel:Yeah, because, I mean, I think there's a broader term, right? AI is within that, but, like, physics modeling and simulations and, like—

30:11Dwarkesh Patel:Or really just, like, he's not embracing the sort of, like, main use case and—

30:14Dylan Patel:I think he's embracing it. But, like, I just don't think he's, like, AGI-pilled like Dario, right? Or Sam. But he's still way, way more AGI-pilled than Google was at Q3 of last year or Amazon was at Q3 of last year. And he saw way more demand, right? And the reason is pretty simple. You know, you can see all the data center construction. He's like, okay, I want to have this market share. You know, we sort of like have all the data centers tracked. And, you know, you can see, you know, there's a lot of data centers that you could say, well, they could be one or the other, right? And so to some extent, Google and Amazon, you know, Google especially, even though their TPU is just better for them to deploy, they have to deploy a crap load of GPUs because they don't have enough TPUs to fill up their data centers.

30:54Dylan Patel:They can't get them fabbed.

30:55Dwarkesh Patel:Wait, can I? So I have a question about that. Google sold, I think, a million, was it the V7s, the Ironwoods, to Anthropoc. and you're saying in general, there's this big bottleneck right now, this year, next year, I mean, I guess going forward forever now is going to be the, you know, logic memory, the stuff that like it takes to build these ships. And Google has DeepMind. This is the other third prominent AI lab. And if this is the big bottleneck, why would they sell it rather than just giving it to DeepMind?

31:23Dylan Patel:Right. So this is again, like a problem with like, you know, DeepMind people are like, this is insane. Why did we do this? Yeah. Right. But then Google Cloud people and Google executives saw a different thought process, right? And basically, you and I know the compute team. There's one guy from, both of them actually came from Google, the main people on the compute team at Thropic. They saw this dislocation. They negotiated a deal, and they were able to get access to this compute before Google realized. And so actually, the chain of events, at least from our data that we found was in early Q3, we saw over the course of like six weeks, we saw capacity on TPUs go up by a significant amount over the course of those six weeks.

32:11Dylan Patel:And it went up like multiple times in those six weeks, right? There were multiple requests. Google even had to go to TSMC and explain to them why they needed this increase in capacity because it was so sudden. But a lot of that capacity increase was for selling to Anthropic.

32:24Dwarkesh Patel:Yeah.

32:25Dylan Patel:Because Anthropic saw it before Google. and then Google had Danobanano and Gemini 3 which caused their user metrics to skyrocket and leadership at Google was like oh and then they started making the statement of we have to double compute every is it six months or I don't remember the exact number that they said but they really woke up a lot more and then they're like oh hey TSMC we want more we want more and it's like well sorry guys like we're sold out for next year we can work on next year we can maybe get like 5-10 % more for 26 but really we're going to work on 27 right it's sort of like you know there's this like information asymmetry of the labs in my mind, right?

32:58Dylan Patel:I don't know if this is exactly the narrative I've spun myself from seeing all the data in the supply chain on like wafer orders and like what's going on with the data centers that, you know, Anthropic signed and FluidStack signed and all this, like sort of, it's pretty clear to me that Google screwed up. And you can see this from Google's Gemini ARRs, right? They had next to nothing in Q1 to Q3, Q3 a little bit, right? Once they started inflecting, but Q4, they were at like 5 billion ARR, right? Exiting or something like this. So it's like, or 5 billion revenue for Q4 on an ARR basis. And so it's clearly like Google didn't see revenue skyrocket.

33:33Dylan Patel:And in a sense, right, Anthropic was not willing, you know, it was kind of had like a little bit of commitment issues before their ARR exploded, even though they have far more information asymmetry and see what's coming down the pipe. Google is going to be more conservative than Anthropic is, A, and B, Google had even less ARR. So they sort of were like, I think, just not willing to like sort of do it. And then they realized they should do it. And so now since then, Google has gotten absurdly AGI pilled, right? In terms of like what they're doing, they bought an energy company, they're buying putting deposits down for turbines, they're buying a ridiculous percentage of the powered land, they're going to utilities and negotiating long term agreements are doing this on the data center and power side, very, very aggressively, right?

34:21Dylan Patel:So, you know, I think Google woke up towards the end of last year, but it took them some time.

34:26Dwarkesh Patel:And how many gigawatts do you think Google will have by the end of next year? Buy my data. You charge for that kind of information. Yes, yes. I feel like every year, the bottleneck for what is preventing us from scaling AI compute keeps changing. A couple of years ago, it was co-host. Last year, it was power. This year, you'll tell me what the bottleneck is this year. But I want to understand five years out, what will be the thing that is constraining us from deploying the singularity?

34:51Dylan Patel:Yeah, I think the biggest bottleneck is compute. And for that, the longest lead time supply chains are not power or data centers. They're actually the semiconductor supply chain themselves, right? It switches back from being power and data center as a major bottleneck to chips. And in the chip supply chain, there's a number of different bottlenecks, right? There's memory. There's logic wafers from TSMC. There's fabs themselves. Construction of the fabs takes a couple years, two to three years, versus a data center takes less than a year, right? We've seen Amazon build data centers in as fast as eight months, right?

35:28Dylan Patel:So there's a big difference in lead times because of the complexity of the building, the fab that actually makes the chips. And then the tools, right? Those also have really long lead times. And so the bottlenecks, as we've scaled, have shifted from, hey, what is the supply chain currently not, what is it currently not able to do, which was COOS and power and data centers. But those were all shorter lead time items, right? COOS is a much more simple process of packaging chips together. Power and data centers are ultimately way more simple than the actual manufacturing of the chips. And so there's been some sliding of capacity across mobile or PC to data center chips, But that's been somewhat fungible, whereas COOS and power and data centers have sort of had to start anew as supply chains.

36:15Dylan Patel:But now there's sort of no more capacity for the mobile and PC industries, which used to be the majority of the semiconductor industry, to shift over to AI, right? NVIDIA is now the largest customer at TSMC. And NVIDIA is the largest customer at SK Hynix, the largest memory manufacturer, right? So it's sort of impossible for the scaling or the sliding of resources away from the common person, right, PCs and smartphones to shift any more towards the AI chips. And so now how do we scale the AI chip production? And that's the biggest bottleneck as we go to 2030 is those.

36:51Dwarkesh Patel:It would be very interesting if there's an absolute gigawatt ceiling that you can project out to 2030 based just on, hey, we can't produce more than this many EUV machines.

37:03Dylan Patel:Right. So to scale compute further, right, there's some different bottlenecks this year, next year. But ultimately, by 28, 29, the bottleneck falls to the lowest rung on the supply chain, which is ASML. right? ASML makes the world's most complicated machine, i.e. an EUV tool. And the selling price for those is$300,$400 million. And currently, they can make about 70. Next year, they'll get to 80. Even under very aggressive supply chain expansion, they only get to a little bit over 100 by the end of the decade. And so what does that mean? Okay, they can make 100 of these tools by the end of the decade and 70 right now.

37:40Dylan Patel:How does that actually translate to AI compute, right? We see all these numbers from Sam Altman and many others across the supply chain, gigawatts, gigawatts, gigawatts, right? How many gigawatts are we adding? And we see Elon saying, hey, the 100 gigawatts in space. A year. A year, right. The problem with any of these numbers or the challenge to these numbers is actually not the power, not the data center. We can dive into that. But it's manufacturing the chips, right? So a gigawatt of NVIDIA's Rubin chips, right? So Ruben is announced at GTC, I believe, the week this podcast goes live. And to make a gigawatt worth of data center capacity of NVIDIA's latest chip that they're releasing at the end of this year or towards the end of this year, you need a few different wafer technologies, right?

38:27Dylan Patel:You need about 55 ,000 wafers of 3 nanometer. You need about 6 ,000 wafers of 5 nanometer. And then you need about 170 ,000 wafers of DRAM memory. And so across these three different buckets, each of these requires different amounts of EUV, right? So when you manufacture a wafer, there's thousands and thousands of process steps where you're depositing material, removing them. But the sort of key critical step, which at least in advanced logic is like 30 % of the cost of the chip, is something that doesn't actually put anything on the wafer, right? You take the wafer, you deposit photoresist, which is like a chemical that basically chemically changes when you expose it to light.

39:07Dylan Patel:And then you stick it into the EUV tool, which shines light at it in a certain way. It patterns it, right? Because there is what's called a mask, which is a stencil effectively for the design. And so when you look at a wafer, you know, a leading edge three nanometer wafer has 70 or so masks, right? 70 or so layers of lithography, but 20 of them are the most advanced EUV, right? And that specifically, you know, if you think about, okay, well, if I need 55 ,000 wafers for a gigawatt, if I do 20 EUV passes per wafer, you then can do the math. That's like, okay, that's 1.1 million passes of EUV for a single gigawatt.

39:43Dylan Patel:So actually, it's pretty simple. And then once you add the rest of the stuff, it ends up being 2 million right across 5 nanometer and all the memory. You're at roughly 2 million EUV passes for a single gigawatt. These tools are very complicated. So when you think about what it's doing across a wafer, it's taking the wafer and it's scanning and it's stepping across, right? It's standing, stepping across, and it does this hundreds of times across the entire, or dozens of times across the whole wafer. And so when you're talking about, hey, how many EUV passes, that's the entire wafer is being exposed at a certain rate.

40:15Dylan Patel:A wafer, a EUV tool can do roughly 75 wafers per hour, and the tool is up roughly 90 % of the time, right? So in the end, you end up with, actually, I need about three and a half EUV tools to do the 2 million EUV wafer passes for the gigawatt. So 3.5 EUV tools satisfies the gigawatt. So it's funny to think about the numbers, right? Because we're talking about, oh, what's the gigawatt cost? It costs like$50 billion roughly, right? Whereas what does 3.5 EUV tools cost? That's like 1.2, right? It's actually like quite a lower number, which is interesting to think about like, oh, 50 gigawatts of economic sort of capex in the data center.

40:53Dylan Patel:And what gets built on top of that in terms of tokens is even larger, right? It might be$100 billion worth of AI value into the supply chain is held up by this$1.2 billion worth of tooling that simply just cannot expand its supply chain quickly.

41:06Dwarkesh Patel:And I think, so you had this article recently where you're saying over the last three years, TSMC has done$100 billion of CapEx. It's like 30, 30, 40. And if you think about, I mean, And a small fraction of that is sort of like being used by NVIDIA for the three nanometer that it's going to or, you know, previously four nanometer that it's using for its chips. But NVIDIA has turned that into what was what are its like earnings last quarter was like 40 billion and so 40 billion times four. So one hundred and sixty billion dollars. So NVIDIA alone is turning some small fraction of one hundred billion in CapEx.

41:43Dwarkesh Patel:It's going to be depreciated over many years, not just this one year, into one hundred and sixty billion dollars in a single year. And then it gets even more intense when you go down the supply chain to ASML, which is taking a billion dollars worth of machines to produce a gigawatt. And of course, those machines last for more than a year, right? So it's doing more than that. Okay, so now I want to understand, okay, well, how many such machines will there be by 2030 if you include not just the ones that are sold that year but have been compiling over the previous years? And what does that imply about the – Sam Altman says he wants to do a gigawatt a week in 2030.

42:14Dwarkesh Patel:When you add up those numbers, is that compatible with that? Right.

42:17Dylan Patel:That's completely compatible, right? Because if you think about TSMC and the entire ecosystem has something 250 to 300 EUV tools already. And then you stack on 70 this year, 80 next year, growing to 100 by 2030. You're at like 700 EUV tools by the end of the decade. 700 EUV tools, three and a half tools per gigawatt, assuming it's all allocated to AI, which it's not. But three and a half tools per gigawatt gets you to 200 gigawatts worth of AI chips for the data centers to deploy, right? So 200 gigawatts, Sam wants 50 gigawatts, right? 52 gigawatts a year. He's only taking 25 % share then, right?

42:51Dylan Patel:Obviously, there's some share given to mobile and PC, assuming that for some reason we're allowed to even have consumer goods still and we don't get priced out of them. But roughly, he's saying 25 % market share of the total chips fab. That's kind of very reasonable given this year alone, I think he's going to have access to 25 % of the Blackwell GPUs that are deployed, right? So it's not that crazy.

43:19Dwarkesh Patel:I find it surprising that, you know, when was the first, when did ASML start shipping EUV tools? When the 7nm started, so I don't know when that was exactly. But you're saying in 2030, they're going to be using machines that initially were shipped in 2020. So 10 years, you're using the same most important machine in this most technologically advanced industry in the world. I find that surprising.

43:41Dylan Patel:So ASML has been shipping EUV tools now for roughly a decade, but it only entered mass volume production around 2020. The tool is not the same. Back then, the tools were even lower throughput. There's various specifications around them called overlay. I was mentioning you're stacking layers on top of each other. You'll do some EUV. You'll do a bunch of different process steps, depositing stuff, etching stuff, cleaning the wafer, dozens of those steps before you do another EUV layer. There's a spec called overlay, right? Which is, okay, you did all this work. You know, you drew these lines on the wafer.

44:15Dylan Patel:Now I want to draw these dots, right? Let's just say I want to draw these dots to connect these lines of metal to, and then holes. And then the next layer up is another set of lines that goes perpendicular. So now you're connecting wires going perpendicular to each other. You have to be able to land them on top of each other. So it's called overlay. And overlay is a spec that's been improved rapidly by ASML. Wafer throughput has been improved rapidly by ASML. And also the price of the tool has gone up, but not as much as the capabilities of the tool, right? Initially, the EUV tools were like$150 million, and over time, they're now like$400 million, you know, as I look out to 2028.

44:50Dylan Patel:But the capabilities of the tools have more than doubled as well, right? Especially on throughput and overlay accuracy, which is the ability to stack, you know, accurately align the subsequent passes on top of each other, even though you do tons of steps between. And so this is, you know, ASML is improving super rapidly. I think it's also something noteworthy to say. ASML is, you know, maybe one of the most generous companies in the world, right? They have this linchpin thing. No one has anything competitive. Maybe China will have some EUV by the end of the decade, but no one else, you know, has anything even close to EUV.

45:27Dylan Patel:And yet they haven't taken price and margins up like crazy, right? You go ask some other folks that we talk to all the time, like, for example, Leopold, and they're like, let's have the price go up, right? Because they can. The margin is there. You can take the margin. Like, NVIDIA takes the margin. Memory players are taking the margin. But ASML has never risen the price more than they've increased the capability of the tool. And so, in a sense, they've always provided net benefit to their customer. It's not that the tool is stagnant. It's just that, like, these tools are old. Yes, you can upgrade them some, and the new tools are coming.

46:01Dylan Patel:And for simplicity's sake, we're kind of ignoring the advances for this podcast, the advances in overlay or throughput per tool.

46:07Dwarkesh Patel:So you say we're producing 60 of these machines this year and then 70, 80 over subsequent years. What would happen if ASML just decided to double its CapEx or triple its CapEx? What is preventing them from producing more than 100 in 2030? Why so confident that even five years out, you can be relatively sure what their production will be?

46:29Dylan Patel:So I think a couple factors here, right? ASML has not decided to just go YOLO, let's expand capacity as fast as possible, right? In general, the semiconductor supply chain has not, right? It's lived through the booms and busts. And we can talk a bit more about it. But basically, no one, you know, some players as of very recently have like woken up. But in general, no one really sees demand for 200 gigawatts a year of AI chips or, you know, trillions of dollars of spend a year. in the semiconductor supply chain. They're just like, they're not AI-pilled, right? They're not AI-pilled.

47:02Dwarkesh Patel:We're going to get to a trillion dollars this year.

47:05Dylan Patel:Yeah, I feel you, but I'm saying like, no one really understands this in the supply chain. Constantly, we're told our numbers are way too high. And then when they're right, they're like, oh yeah, but your next year's numbers are still too high. And it's like, but anyways, like ASML has sort of, their tool has four major components, right? It has the source, right? Which is made by Symer in San Diego. It has the reticle stage, which is made in Wilmington, Connecticut, right? It has the wafer stage and the optics, right, the lenses and such. And those two are made in Europe, right? And so when you look at each of these four, they're tremendously complex supply chains that A, they have not tried to expand massively, and B, when they try to expand them, the time lag is quite long, right?

47:53Dylan Patel:Right. And so, again, this is the most complicated machine that humans make, period. Right. At a volume, any sort of volume. But like, let's talk about the source specifically. Right. What does the source do? It drops these tin droplets. It hits it three subsequent times with the laser perfectly. So the first one hits this tin droplet, expands out. It hits it again. So it expands out to this perfect shape. And then it blasted at super high power. and the tin droplets get excited enough that they release EUV light, 13.5 nanometer. And then it's in this thing that is like basically collecting all the light and directing it into the lens stack, right?

48:29Dylan Patel:Then you have the lens stack, which is Carl Zeiss, right, as you mentioned, and some other folks, but Zeiss being the most important part of it. They also have not tried to expand production capacity because they don't see any, you know, they're like, oh, yeah, yeah, like we're growing a lot because of AI. We're growing from 60 to 100, right? It's like, no, no, no, no, no. We need to go to like a couple hundred, but it's fine, whatever. Each of these tools has, you know, I think 18 of these lenses effectively. Mirrors, they are multilayer mirrors, which are perfect layers of molybdenum and ruthenium, if I recall correctly, stacked on top of each other in many layers.

49:05Dylan Patel:And then the light bounces off of it perfectly. But it's not just like, you know, like when we think about a lens, you know, it's like in a shape and it focuses the light. This is a, this is like a mirror that's also a lens. And so it's pretty complicated. Any defect in this perfect layer of stat in this in these like super thinly deposited stacks will mess it up. Any curvature issues like there is a lot of challenges with scaling the production. It's quite artisanal. Right. In the sense. Right. Because you're not making tens of thousands of these a year. You're making hundreds. You're making thousands.

49:35Dylan Patel:Right. You know, talk about 60 tools a year. 18 of these per tool you end up with. You know, you're still in the hundreds of tools. or a thousand, you're at the thousand number roughly for these lenses and projection optics. So then you step forward to the reticle stage, which is also something really crazy. This thing moves at, I want to say, 9Gs. Like it will shift 9Gs because as you step across a wafer, the tool will go, and the wafer stage is complimentary. It's the wafer part. So you line these two things up, you're taking all the light through the lenses that's focused. and here's the reticle, here's the wafer and you're passing, the reticle's moving one direction, the wafer's moving the other direction as it scans a 26 by 33 millimeter section of the wafer and then it stops, it shifts over to another part of the wafer and does it again.

50:28Dylan Patel:And it does that in just seconds, right? And each of them are moving at nine Gs in opposite directions. So each of these things is like a wonder and marvel of like chemistry, fabrication, you know, sort of like mechanical engineering, optical engineering, because you have to align all these things and make sure they're perfect. All these things have crazy amounts of metrology because you have to perfectly test everything because if anything is messed up, the yield goes to zero, right? Because this is such a finely tuned system. And by the way, it's so large that you're building in the factory in Heindhoven, Netherlands, and they're deconstructing it and shipping it on many planes to the customer site, and then you're reassembling it there and testing it again.

51:12Dylan Patel:And that process takes many, many months. So like, it's just, there's so many steps in the supply chain, right? Whether it's Zeiss making their lenses and projection optics, or CIMR, which is an ASML-owned company, making the EUV source. And each of these has its own complex supply chain, right? ASML's commented, their supply chain has over 10 ,000 people in it, right? Like individual suppliers. Yes. And it might not be directly, it might be through like, hey, you know, Zeiss has so many suppliers and XYZ company has so many suppliers, but if you just think about like, okay, you're talking about two physically moving objects that are like this large and this large, the size of a wafer, right?

51:50Dylan Patel:And it has to be accurate to the level of single digit nanometers or even smaller because the entire system, the overlay, right? Layer to layer variation has to be on the order of three nanometers, right? And so if the overlay is three nanometers, that means each individual part, the accuracy of its physical movement has to be even less than that, right? It has to be sub one nanometer in most cases because the error of these things stacks up, right? And so there's no way to just snap your fingers and increase production, right? It's things simple as power, right? The US going from 0 % power growth to 2 % power growth, even though China's already at 30, was like so hard for America to do, right?

52:34Dylan Patel:And that's a really simple supply chain with very few people in the supply chain, right? Who make difficult things. And there's, you know, probably what? 100 ,000 electricians slash people who work in the supply chain of electricity or more in the US. And, you know, when you look at, oh, ASML employs like so few people. Carl Zeiss probably employs like less than a thousand people working on this. And all of those people are like super, super specialized. So, you know, you can't just train random people up for this, like in the snap of a finger. You can't just get your entire supply chain to get galvanized, right?

53:08Dylan Patel:NVIDIA has had to do a lot to get the entire supply chain to even deliver the capacity they're going to make this year. Even though when you go talk to Anthropic, they're like, well, we're short of TPUs, we're short of training, we're short of GPUs. When you go talk to OpenAI, they're like, we're short of these things, right? So OpenAI and Anthropic, they know they need X. NVIDIA is not quite as AGI-pilled and they're building X minus one. And you go down the supply chain, everyone's doing minus one. And in some cases, they're doing like divided by two, right? Because they're not AGI-pilled, right?

53:38Dylan Patel:And so you end up with the time lag for this whip to react, right? The sort of AI-pilledness and desire to increase production is so long. And then once they finally understand, hey, we need to increase production rapidly, right? And they think they understand, oh, AI means we have to go from 60 to 100. In addition to the tools all just getting better and faster, the source getting higher power from 500 watts to 1 ,000 and all these other aspects of the supply chain advancing technically plus increase of production. They think they're actually increasing production a lot. But if you float through the numbers of, hey, what does Elon want?

54:16Dylan Patel:He wants 100 gigawatts a year in space by 2028, is it? Or 2029? And, you know, Sam Altman wants 50 gigawatts, 52 gigawatts a year by the end of the decade. And you look at, you know, probably Anthropic needs the same. And then, you know, Google needs that. You know, you go across the supply chain. It's like, wait, no, the supply chain can't possibly build enough capacity for everyone to get what they want on the side of compute.

54:41Dwarkesh Patel:Real conversations are full of fits and starts and pauses and interruptions. I mean, just listen to this episode. At least superficially, voice models have gotten pretty good at handling these kinds of things. But at a deeper level, interruptions can throw off a model's understanding and degrade the quality of its responses. And it's not always clear why. Labelbox realized that this was a huge bottleneck for their customers. So they built an evaluation pipeline called EchoChain to help you diagnose and fix your voice model's specific failure modes. EchoChain starts by feeding conversations into your voice model.

55:09Dwarkesh Patel:It then injects interruptions at specific intervals and classifies any failures into one of three different modes. One, did it acknowledge a correction but keep the old plan? Two, did it adapt briefly but then slide back to old assumptions? Or three, did it abandon the old task entirely? This is extremely useful information because LabelBox can get your model the exact data it needs to fix whatever issue is preventing it from being a viable and competent voice model. So if you want to ensure that your voice model stays performant in real conversations, you should reach out to LabelBox. Go to labelbox.com slash thwarkash.

56:08Dwarkesh Patel:these are the alternatives that people fall back on. And I want to ask you a question about whether we can imagine a similar thing happening in the semiconductor supply chain. So if EUV becomes a bottleneck, well, what if we just went back to 7 nanometer and do what China is doing currently in producing 7 nanometer chips with multi-patterning with DUV machines? And if you look at a 7 nanometer chip like the A100, there's been a lot of progress, obviously, from the A100 to the B100 or B200. But how much of that progress is just numerics? And then if you just told constant, say FP16 from A100 to B100, the B100 is a little over one petaflop.

56:53Dwarkesh Patel:And then A100 is like 300 teraflops. And so you have basically 3x holding numerics constant. You have a 3x improvement from A100 to B100. And then some of that is the process improvement. Some of that is just the accelerator design improving, which we could replicate again in the future. And so then it just seems like actually it's like very small effect from the process improving from seven nanometer to four nanometer. So I don't know, say we have, I don't know the numbers offhand, but let's say there's like 150k wafers per month of three nanometer and then eventually similar amounts for two nanometer.

57:30Dwarkesh Patel:But then there's a similar amount for seven nanometer, right? So if you have all those old wafers and then there's maybe a 50 % haircut because the process, the bits per wafer area are like, what is it, 50 % less or something? Then it's like, it doesn't seem like that bad to just bring on seven nanometer wafers and then, oh, that gives you another 50 or another 100 gigawatts. Yeah, tell me why that's naive.

57:54Dylan Patel:Yeah, so I think, you know, we potentially do go crazy enough that this happens because we just need incremental compute and the compute is worth the higher cost power, etc., of these chips. But it's also unlikely to some extent, to a large extent, because of, I think, just comparing, you know, some of these are like not fair comparisons, right? For example, you know, from A100, which is 312 teraflops, to Blackwell, which is like 1 ,000-ish of FP16, or maybe it's 2 ,000 and then Rubin is like 5 ,000 or so FP16. It's not a fair comparison because these chips have vastly different, you know, design targets, right?

58:39Dylan Patel:At A100, that is what NVIDIA optimized for was FP16, BFlood16, numerics. When you look at Hopper, they didn't care as much about that. They cared about FPA. When you look at Rubin, they don't care about FP16 and BF16 as much. They care mostly about FP4 and 6, right? And so numerics are what they've designed their chip for. And so there's a couple like, you know, okay, let's just say, let's redesign. Let's make a new chip design on 7 nanometers. Sure, we can do that. And then it's optimized for the numerics of the modern day. The performance difference is still going to be much larger than the flops different you mentioned, right?

59:19Dylan Patel:Often it's easy to boil things down to flops per watt or flops per dollar. But that's actually not a fair comparison, right? And so this is where sort of you can bring in, hey, let's look at Kimi K1 or DeepSeek. When you look at Kimi or Kimi K 2.5, sorry, and DeepSeq, when you look at these two models and you look at their performance on Hopper versus Blackwell on, you know, very optimized software, you get vastly different performance, right? And most of this is not attributed to flops. A lot of this is – or numerics, right? Because those models are actually 8-bit. So it's not like Blackwell's and Hopper, they're both optimized for 8-bit and Blackwell's not really taking advantage of its 4-bit there.

1:00:02Dylan Patel:you know the performance gulf is is actually much larger and you know the way you can sort of compare them and think about them is sure it's one thing to you know shrink process technology and make the transistor smaller and each chip has x number of flops but you forget the big gating factors these models don't run on a single chip they run on hundreds of chips at a time right if you look at deep seek's production deployment which is well over a year old now they were running on 160 gpus right and that's what they serve production traffic on and so they split the model across 160 gpus every time you cross the barrier of a chip to another chip there is an efficiency loss because you now have to transmit over you know high speed electrical surities and there is a latency cost there's a power cost there's a there's all these dynamics that hurt as you shrink and shrink and shrink the process node you've increased the amount of compute in a single chip Now, in-chip, movement of data is at hundreds or at least tens of terabytes a second, if not hundreds of terabytes a second.

1:01:04Dylan Patel:Whereas between chips, you're on the order of a terabyte of second, right? And so this movement of data between chips that are super close to each other physically, and then you can only put so many chips close to each other physically, so you have to put chips in different racks. the order of data between that is on the order of hundreds of gigabits a second, right? 400 gig or 800 gig a second. So 100 gigabytes a second, roughly. And so you've got this like huge ladder of like, oh, on chip, I can communicate at super fast speeds. Within the rack, I can communicate at order of magnitude speeds.

1:01:35Dylan Patel:Outside the rack, I can communicate at an even order of magnitude lower than that. And as you break the bounds of chips, you end up with this performance loss. Anyways, the reason I explain this is because when you look at Hopper versus Blackwell, Even if both of them are using, you know, a rack worth of chips, the hopper is significantly slower because the amount of performance that you have leveraged to the task within that, you know, within each domain of, hey, tens of terabytes a second of communication between these transistors or these processing elements. And, you know, terabytes a second between these processing elements is much, much higher.

1:02:09Dylan Patel:And therefore, the performance is much higher. So when you look at inference at, let's say, 100 tokens a second for DeepSeq and Kimi K2.5, Hopper versus Blackwell, the performance difference is on the order of 20x. Not 2 or 3x like the flops performance difference indicates, even though those are on the same process node. Makes sense, yeah. There's just differences in networking technologies and what they've worked on. And so you can translate some of these back. But when you look at like Rubin, what they're doing on 3 nanometers, some of these things are just not possible to do all the way back on A100, even if you make a new chip for...

1:02:40Dylan Patel:Interesting. 7 nanometer. There's just like certain architectural improvements you can port. There's certain ones you cannot. And so the performance difference is not just going to be the difference in flops. It's in some senses cumulative between the difference in, you know, flops per chip, networking speed between chips, how many flops are on a chip versus a system, memory bandwidth on a single chip and on an entire system.

1:03:02Dwarkesh Patel:All of these things compound. Can I ask you a very naive question? So this year, last year, the B200 has now two dyes on a single chip so you can get that bandwidth on a single chip without having to go through enemy link or infinite band. And then next year, Ruben Ultra will have four dyes on one chip. What is preventing us from just doing that with an old, like how many dyes could you have a single chip and still get these tens of terabytes a second?

1:03:27Dylan Patel:Yeah. So even within Blackwell, there are differences in performance when you go, when you're communicating on the chip versus across the chips. Those bounds are obviously much smaller than when you're going, you know, out of the entire chip, but each die versus, you know, within the package. And so anyways, when you scale, you know, the number of chips up, there is some performance loss. It's not just perfect, but it is way better than different entire packages. Now, how large can advanced packaging scale? the way nvidia is doing it is co-auth the way uh you know google and with broadcom and media attack and you know amazon tranium all these chips are doing is called co-auth but actually you can go and look back at what um what tesla did with dojo right dojo uh which they canceled and restarted anyways dojo was a chip that was the size of an entire wafer they had 25 chips on it um and there were some trade-offs, right?

1:04:24Dylan Patel:They couldn't put HBM on it. But the positive side of it was that they had 25 chips on it. And so to date, it is still probably the best chip for running convolutional neural networks. It's just not great at transformers because the, you know, the sort of the shape of the chip, the memory, the arithmetic, all these various specifications of it are just not well suited for transformers. They're well suited for CNNs. And anyway, so, So Dojo chips were optimized around that. They made a bigger package. But at the same time, as you make packages bigger and bigger and bigger, you have other constraints, right?

1:04:58Dylan Patel:Networking speed, memory bandwidth, cooling capabilities, all of these things start to rear their heads. It's not simple. But yes, you will see a trend line of more chips on the package. And yes, you're going to be able to do that on 7 nanometer. In fact, that's what Huawei did with their Ascend 910C or D. they put they put they were initially just one and then they did two and they're focusing on scaling the packaging up because that is an area where they can advance faster than sort of process technology where they can't shrink but at the end of the day that's still you know that's something that you can do on the leading edge chips too right anything you do on seven nanometer you can also probably do on three nanometer in terms of packaging um so if we're if you end up in this

1:05:38Dwarkesh Patel:world in 2030 where the west has the most advanced process technology but it has not ramped it up as much. Whereas China, I don't know if you think by 2030, they would have EUV and I don't know, two nanometer or whatever, but they are semiconductor pills. So they're producing in mass quantity. Basically, I'm wondering what the year is where there's a crossover, where our advantage in process technology has faded enough and their advantage in scale has increased enough. And also their advantage in like having one country that has the entire supply chain indigenized rather than having random suppliers in Germany and Netherlands and whatever would mean that China would be ahead in its ability to produce mass flops.

1:06:21Dylan Patel:Yeah. So to date, China still does not have entire indigenized semiconductor supply chain, right? But were they in 2030? Yeah. By 2030, it's possible that they do. But to date, right, all of China's 7 nanometer and 14 nanometer capacity uses ASML DUV tools, right? And the amount that they can ship and import from ASML is large. But the point being that the vast majority of ASML's revenue, especially on EUV, all of it is outside of China. So the scale advantage is still in the favor of, let's call it the West plus Taiwan, Japan, et cetera.

1:06:59Dwarkesh Patel:They're trying to make their own DUV and EUV tools, right?

1:07:01Dylan Patel:They're trying to do all these things. The question is how fast can they advance and scale up production as well as quality. And to date, we haven't seen that. Now, I'm quite bullish that they're going to be able to do these things over the next five to 10 years, right? Really scale up production, really kick it into high gear. They have more engineers working on it. They have more desire to throw capital at the problem.

1:07:24Dwarkesh Patel:So by 2030, do they have fully indigenized EUV?

1:07:27Dylan Patel:I think for sure, for sure. EUV, yes.

1:07:29Dwarkesh Patel:And fully indigenized EUV by 2030?

1:07:30Dylan Patel:I think they'll have working tools. I don't think that they'll be able to manufacture a bunch yet, right? You know, there's sort of having it work and then there's production hell, right? And ultimately, like, ASML had EUV working in the early 2010s at some capacity, right? Now, the tools were not accurate enough. They were not scaled for high volume manufacturing, reliable enough. And then they had to ramp production. And that all took time. Production hell takes time, right? Which is why it took another five to seven years to get EUV into mass production at a fab rather than just working in the lab.

1:08:07Dwarkesh Patel:So how many DUV tools do you think anybody will manufacture in 2030?

1:08:11Dylan Patel:ASML?

1:08:12Dwarkesh Patel:No, China.

1:08:13Dylan Patel:Oh, that's a great question.

1:08:18It's a bit of a challenge to look into this supply chain, especially.

1:08:24Dylan Patel:We try really hard. but you know in some instances they're like buying stuff from Japanese vendors and if they want a fully indigenized supply chain they need to not buy these lenses or buy these projection optics or stages from Japanese vendors they need to build it internally so it's really tough to say where they'll be able to get to like I honestly think it's like a shot in the dark but it's it's probably not unlikely that they'll be able to do you know on the order of a hundred DUV tools a year whereas ASML is doing hundreds of DUV tools a year currently. You know, no one's made a process node.

1:08:59Dylan Patel:No company has a process node where they make a million wafers a month, right? Elon says he wants to do it and China's obviously going to do it, right? And I don't think the, you know, TSMC is trying to do that. The memory makers may get there as well, right, to the million wafers a month, but not in a single fab. It's sort of mind-boggling to think of that scale and challenging to see the supply chain galvanize for that. So I'm not sure, you know, I don't want to doubt, you know, China's capability to scale.

1:09:29Dwarkesh Patel:Right. I guess this is an interesting question that I think it might, you know, at some point in time analysis, we'll do the deep dive on this. But I think this question of like, by when would China be able, like indigenized Chinese production could be bigger than the rest of the West combined? If you just add up like all the input of the input of your model, when they'll have UV machines at scale, when they'll have UV machines at scale. Because I think there's this like question around if you have long timelines on AI, by long meaning 2035, which is not that long in the grand scheme of things.

1:10:02Dwarkesh Patel:Should you expect a world where China is like dominating in semiconductors, which I think, I don't know, doesn't get asked enough in San Francisco. We're just like thinking on timescale of like, you know, weeks. And then if you're outside of San Francisco, you're not thinking about AGI at all. And so this question of like, OK, what if we have AGI? What if you have this transformational thing that is commanding tens of trillions of dollars or hundreds of trillions of dollars of economic growth and, you know, token output and so forth? But then it happens in 2035. And what does that imply for the West versus China?

1:10:31Dwarkesh Patel:I think it's just like, I don't know, the semi-analysis has got to ride the definitive model on this.

1:10:36Dylan Patel:Yeah. So I think it's really challenging when you move time scales out that far, right? Like what we tend to focus on is like we're tracking every data center, we're tracking every fab, we're tracking all the tools and we're tracking where they're going. But the time lags for these things are relatively short, right? We can only make reasonably accurate estimates for data center capacity based on land purchasing and permits and turbine purchasing and all these things. And we know where all these things are going and that's what the data we sell is. But as you go out to 2035, things are just so radically different and your error bars get so large, it's kind of hard to make an estimate.

1:11:15Dylan Patel:But at the end of the day, like, you know, there is if takeoff or timelines are slow enough, right, then certainly China, I don't see why they wouldn't be able to catch up drastically. Right. You know, in some sense, we've got like this valley, right, of where, you know, call it three to six months ago, Chinese models were or maybe even now Chinese models are competitive as they've ever been. I think Opus 4.6 and GPT 5.4 have really pulled away and made the gap a little bit bigger, but I'm sure some new Chinese models will come out. But as we move from, hey, these companies are selling tokens where they provide the entire reasoning chain and all that to selling automated white-collar work, automated software engineer, send them the request, they give you the result back, and there's a bunch of thinking on the back end that they don't show you.

1:12:02Dylan Patel:The ability to distill out of American models into Chinese models will be harder, A. B, as the scale of the compute that the labs have, right? OpenAI exited the year with roughly 2 gigawatts last year. Anthropic will get to 2 plus gigawatts this year. And by the end of next year, they'll both be at like 10 gigawatts of capacity. China is not scaling their AI lab compute nearly as fast. And so at some point, you know, when you can't distill the learnings from these labs into the Chinese models, plus this compute race that OpenA Anthropik, Google, et cetera, Meta are all racing on. At some point, they end up getting to a point where the model performance should start to diverge more.

1:12:44Dylan Patel:And then all of this CapEx that's being spent on data centers and all that, Amazon, 200 billion, Google 180, so on and so forth. All these companies are spending hundreds of billions of dollars of CapEx. there's nearly a trillion dollars of CapEx being invested in data centers in America this year, roughly, right? You end up with, okay, well, what's the return on invested capital here? You and I would think that the return on invested capital for data center CapEx is very high. And at least if we look at Anthropics revenues in January, they added like 4 billion. In February, which is a shorter month, they added like 6.

1:13:21Dylan Patel:We'll see what they can do in March and April, given compute constraints are what's bottlenecking their growth, right? The reliability of cloud code is actually quite low because they're so compute constrained. But if this continues, then the ROIC on these data centers is super high. And at some point, the US economy starts growing faster and faster over the next, you know, this year and next year because of all this capex and all this revenue that these models are generating and downstream supply chain versus China doesn't have that yet, right? they have not built the scale of infrastructure to then invest in model to invest in models to get to the capabilities to then deploy these models at such scale right because when you look at like anthropics hey they're at call it 20 billion arr of that you know the margins are sub 50 percent at least last reported by the information so then you know you're at okay that's like 13 14 billion dollars of compute that it's running on rental cost wise which is actually like$50 billion worth of CapEx that someone laid out for Anthropic to generate their current revenue.

1:14:22Dylan Patel:And China has just not done this. If and when Anthropic 10x is revenue again, and I think our answer would be when, not if, then China doesn't have the compute to deploy at that scale. And so there is some sense of like, oh, we're in fast takeoff-ish, right? It's not like we're talking about Dyson Sphere by X date. It's more like the revenue is compounding at such a rate that it does affect the economic growth. And the resources these labs are gathering are going so fast that, you know, and China hasn't done that yet. So in that case, the US and the West is actually diverging. The flip side is actually these infrastructure investments have middling returns.

1:15:01Dylan Patel:Maybe they're not as good as hoped. You know, maybe Google is wrong for wanting to take free cash flow to zero and spend$300 billion on CapEx next year. Maybe they're just wrong. And, you know, people on Wall Street who are bearish and people who don't understand AI are correct, right? And in which case, then the US is building all this capacity, it doesn't get really great returns. And China is able to build the fully vertical indigenized supply chain, not, you know, US, Japan, Korea, Taiwan, Southeast Asia, you know, Europe, all these countries together building this like less vertical supply chain.

1:15:37Dylan Patel:And in a sense, at some point, China is able to scale past us if AI takes longer to get to certain capability levels than, you know, I would say the vast majority of your guests on this podcast believe.

1:15:48Dwarkesh Patel:It's like fast timelines, U.S. wins, long timelines, China wins. Right.

1:15:51Dylan Patel:But I don't know, like, I don't know what fast timelines means, right? Like, I, like, don't think you have to believe in AGI to have the timelines where the U.S. wins.

1:16:00Dwarkesh Patel:Okay, let's go back to memory, because I think this is maybe people on Wall Street and people in the industry are understanding how big this is, but maybe generally people don't understand how big a deal this is. So we've got this memory crunch, as you're talking about. And earlier I was asking about, oh, could we solve for the EUV tool shortage by going back to 7 nanometers? So let me ask a similar question about memory. HBM is made of DRAM, but has 3 to 4x less bits per wafer area than the DRAM it's made out of. is it possible that accelerators in the future could just use commodity DRAM and not HBM?

1:16:34Dwarkesh Patel:And so just we can make much more capacity out of the DRAM we get. And the reason I think this might be possible is, look, if we're going to have agents that are just going off and doing work and it's not a synchronous chatbot application, then you don't necessarily need extremely high fast latency kinds of things anymore. And so maybe you can have the low bandwidth because the reason you stack DRAM into stacks and make HBM is for higher bandwidth. And so is it possible to go to HBM accelerators and basically have the opposite of cloud code fast, like have cloud code slow and do that?

1:17:16Dylan Patel:I think at the end of the day, the incremental purchaser who's willing to pay the highest price for tokens also ends up being the one that's like less price sensitive. And, you know, the compute should be allocated in a capitalistic society towards the goods that have the highest value. And the private market determines this by willingness to pay. And so to some extent, sure, Anthropik could actually release a slow mode, right? They could release Claude's slow mode and have an increase in tokens per dollar by a significant amount. They could probably, like, reduce the price of Opus 4.6 by, you know, 4x, 5x and reduce the speed by another – by maybe just like 2x.

1:17:54Dylan Patel:Like the curve on inference throughput versus speed is there already just on HBM. And yet they don't because no one actually wants to use a slow model. And furthermore, on these agentic tasks, you know, it's great that the model can run at this time horizon of hours. That's kind of like, OK, well, if the model was just running slower, that hours would become a day. Right. Or vice versa. Right. If the model is running faster, that hours becomes hour. and yet no one really wants to move to that day-long wait period because the highest value tasks also have some time sensitivity to them, right? And so I struggle to see, you know, yes, you could use DDR, but then there's a couple like things that are challenging with this, right?

1:18:38Dylan Patel:You could use regular DRAM. One is you're still limited, you know, one of the like core constraints of chips, even though they're sort of like a chip is like a certain size, all of the IO escapes on the edges of the chip, right? So oftentimes what you see is the left and the right of the chip are HBM, the IO from the chip to the HBM is on the sides, and then the top and bottom are IO to other chips, right? And so if you were to change from HBM to DDR, then all of a sudden this IO on this edge would have significantly less bandwidth, but it had significantly more capacity per chip. Yeah. Because, and so, yes, you're making less, you know, the metric that you actually care about is bandwidth per wafer, not bits per wafer.

1:19:31Dwarkesh Patel:Because the thing that is constraining the flops is just getting in and out the next matrix. And for that, you just need more bandwidth.

1:19:39Dylan Patel:Yeah, getting out the weights and getting in and out the KV cache.

1:19:42Dwarkesh Patel:Right.

1:19:42Dylan Patel:And so in many cases, these GPUs are not running at full memory capacity. Yes, it's obviously like a system design thing, you know, model hardware, software co-design of, hey, how much KV cache do I do? How much do I keep on the chip? How much do I offload to other chips and call when I need it for tool calling or whatever? How many chips do I paralyze this on? Obviously, the search space of this is very broad, which is why we have InferenceX, which is like an open source model, searches all the optimal points on inference for a variety of eight different chips and models. Anyways, the point is you're not always necessarily constrained by memory capacity.

1:20:22Dylan Patel:You can be constrained by conflops. You can be constrained by network bandwidth. You can be constrained by memory bandwidth. Or you can be constrained by memory capacity. There's sort of like four. If you're really to simplify it down, there's like four constraints. And each of these can break out into more. but in this case if you switch to ddr yes you produce 4x the bits per dram wafer but all of a sudden the constraints shift a lot and your system design shifts a lot you go slower yes is the market smaller okay maybe possibly but also now all of a sudden all these flops are wasted because they're just sitting there waiting for memory it's like great i don't need all that capacity because i can't really increase batch size because then the kv cache is going to take even longer to read and And so you never, you can, yeah.

1:21:02Dwarkesh Patel:Interesting. What is the bandwidth difference between HBM and a normal DRM?

1:21:07Dylan Patel:Yeah. So an HBM stack of HBM4, let's just talk about like the stuff that's in Rubin because that's what we've been indexing on, is 2048 bits across connected in an area that's like 13 millimeters wide. So 2048 bits and it transfers memory at around 10 giga transfers a second. So HBM, a stack of HBM4 is 2048 bits on an area that's 13 millimeters wide, roughly, or 11. And that's the shoreline that you're taking on the chip. And in that shoreline, you have 2048 bits transferring at 10 gigatransfers per second. You multiply those together and you divide by eight bits to bytes. You're at roughly two and a half terabytes a second per HBM stack, right?

1:21:46Dylan Patel:When you look at DDR, in that same area, it's maybe 64 or 128 bits wide. And that DDR5 is transferring at anywhere from 6.4 gigatransfers a second to maybe 8 ,000 gigatransfers a second. So your bandwidth is like significantly lower, right? It's 64 times 8 ,000 divided by 8. You're at 64 gigabytes a second. And even if you take a generous interpretation of 128 times 8 gigatransfers, you're at 128 gigabytes a second for the same shoreline versus 2.5 terabytes a second. There's an order of magnitude difference in bandwidth per edge area. And if your chip is a square or it's 26 by 33, right, is the maximum size for a chip, individual die, you only have so much edge area.

1:22:31Dylan Patel:And then on the inside of that chip, you put all your compute. There's things you can do to try and change, right, more SRAM, more caching, blah, blah, blah. But at the end of the day, you're very constrained by bandwidth. Interesting.

1:22:41Dwarkesh Patel:So then there's a question of, like, where can you destroy demand to free up enough for AI? and I guess the picture is especially bad because as you're saying, if it takes 4X more wafer area to get the same byte for HBM, you had to destroy 4X as much consumer demand for laptops and phones and whatever in order to free up one byte for AI. So yeah, what does this imply for the next year or two of, sorry for the run on question. I think on your newsletter, you said 30 % of the capex in 2026 of big tech is going towards memory.

1:23:15Dylan Patel:Yes.

1:23:15Dwarkesh Patel:That's insane, right? Yeah. Like of the 600 billion or whatever, you're saying 30 % is going just to...

1:23:23Dylan Patel:And, you know, obviously there's some level of like margin stacking that NVIDIA does. And so if you separate out, you know, and you apply their margin to the memory and the logic. But at the end of the day, yeah, like a third of their capex is going to memory.

1:23:34Dwarkesh Patel:That's crazy. Okay. So what is the question I'm trying to ask? It's something like, yeah, what is this... Basically, what should we expect over the next year or two as this memory crunch hits?

1:23:41Dylan Patel:Yeah. So memory crunch will continue to be harder and harder. And prices continue to go up. And this affects different parts of the market differently, right? Gets to sort of the like, are people going to hate AI more and more? Yes, because now smartphones and PCs are not going to get incrementally better year on year. And in fact, they're getting incrementally worse.

1:24:00Dwarkesh Patel:If you look at the bill of materials of an iPhone, what fraction of it is the memory? Like how much more expensive does an iPhone get if the memory is 2x more expensive or whatever it has to be?

1:24:09Dylan Patel:So I believe an iPhone has 12 gigabytes of memory. Each gig cost, used to cost roughly$3 or$4. So it's 50 bucks. But now the price of memory is like triple. Let's call it if it's now, it's 12 bucks per gig for DDR. So now you're talking about$150 versus$50, right? A hundred dollar increase in cost on Apple. Also, Apple has some margin. They're not just going to eat the margin. So now that's a hundred dollar cost increase. That's just on the DRAM. The NAND also has the same sort of like market. So in fact, you know, it's probably$150 increase on the iPhone. Apple has to either pass it on to the consumer, A, or B, they have to eat it.

1:24:46Dylan Patel:I don't see Apple reducing their margin too much. Maybe they eat a little bit. But at the end of the day, that means the end consumer is paying$250 more for an iPhone. And now that's on like, hey, what is last year's memory pricing versus today's? Now, there is some lag for Apple to have to feel the heat because they have tended to have, you know, three, six or a year long contracts for a lot of memory. But at the end of the day, Apple gets hit pretty hard by this. but they won't really adjust until the next iPhone release. But that's the high end of the market. Actually, that's only a few hundred million phones a year, right?

1:25:19Dylan Patel:Apple sells, what, two, three hundred million phones a year? The bulk of the market is this mid-range low-end, right? Used to be 1.4 million smartphones were sold a year. Now we're at like 1.1. But our projections are we maybe get down to like 800 million this year. And next year are like 600 or 500 million. Because, and we look at like, you know, there's some data points out of China from some of our analysts in Asia and Singapore and Hong Kong and Taiwan. They've been tracking this and they see Xiaomi and Oppo are cutting low-end and mid-range smartphone volumes by half. Because yes, it's only a$150 price increase on a$1 ,000 smartphone or$150 bomb increase on$1 ,000 iPhone, where Apple has some larger margin.

1:26:01Dylan Patel:But if we look at the smaller phones, the percentage of the bomb that goes to memory and storage is much larger. and the margins are lower. So there's less capacity to even eat the margins. And they have like generally tended not to do as long-term agreements on memory. And why this is like a big deal is if smartphone volumes, let's say half, the halving will frankly happen in the low and mid range, not in the high end. So it's not like the bits released are halving, right? You know, currently consumers more than half of memory demand, even if you half the smartphone volumes because of the shape of the halving, right?

1:26:38Dylan Patel:It's like low-end gets cut by more than half, high-end gets cut by less than half because you and I will buy, you know, the high-end phones that cost north of$1 ,000, we'll buy them even if they get a little bit more expensive. And Apple's volumes will not go down as much as like a low-end smartphone provider. And the same applies to PCs. And what this does to the market is quite drastic, right? DRAM gets released, goes to AI chips who are willing to do longer term contracts, willing to pay higher margins, et cetera, et cetera, because at the end of the day, the margin that they extract is much larger from the end user or whatever.

1:27:13Dylan Patel:And so this probably leads to people hating AI even more, right? Because they're going to start being like, today you already see all the memes on PC subreddits and PC Twitter, gaming PC Twitter is like cat dancing videos. And it's like, this is why memory prices is doubled and you can't get a new gaming GPU, right? Or you can't get a new desktop. And it's gonna be even worse when memory prices double again, especially DRAM. Another dynamic that's quite interesting is it's not just DRAM, it's also NAND. NAND is also going up in price. Both of these markets have expanded capacity very slowly over the last few years.

1:27:48Dylan Patel:NAND almost zero, but smartphones, the percentage of NAND that goes to phones and PCs is larger than the percentage of DRAM that goes to phone and PC. So as you destroy demand, you unlock, you know, mostly for the DRAM purposes, you unlock more NAND that gets allocated and can sort of go to other markets. And so the price increases of DRAM will be larger than those of NAND because you've released more from the consumer. And in fact, you've produced more memory for AI.

1:28:16Dwarkesh Patel:Sorry, but the NAND is, maybe you just explained it and I missed it. Is it because SSDs are being used in large quantities for data centers or?

1:28:23Dylan Patel:They are, but not as large quantities as DRAM.

1:28:27Dwarkesh Patel:Okay, but you're saying they will also increase because they're using some quantity, but there's not as much in need as there is for HBM. Makes sense. One thing I didn't appreciate until I was reading some of your newsletters is that basically the same constraints that are preventing logic scaling over the next few years, it's quite similar to what's preventing us from producing more memory wafers. In fact, like literally the same exact machine, this EUV tool is needed for memory. So I guess, yeah, maybe there's a question that somebody could be asking right now, like, well, why can't we just make more memory?

1:29:00Dylan Patel:Is that somebody you?

1:29:03Dwarkesh Patel:Yeah, who knows?

1:29:04Dylan Patel:So I think the constraints, as I was mentioning earlier, are not necessarily EUV tools today or next year. They become that as we get to the latter part of the decade. but currently, right, the constraints are more so they physically just haven't built fabs, right? So over the last three to four years, these vendors have just not built new fabs. That's because memory prices were really low, their margins were low. And in fact, they were losing money in 2023 on memory. So they're like, oh, we're not building new fabs. And then like the market slowly recovered over time, but never really got amazing until last year.

1:29:40Dylan Patel:You know, in 2024, we were like banging on the drums that like, hey, reasoning means long context, which means large KV cache, which means you need a lot of memory demand. And we've been talking about that for like a year and a half, two years. And people who understand AI like went really long memory then, right? You know, and so you've seen that sort of like dynamic, but now it finally played out in pricing. It took so long for what was obvious, right? Hey, long context, KV cache gets bigger. You need more memory and accelerators, half their cost is memory. So, of course, they're just going to start, you know, they're going to start like going crazy on it.

1:30:13Dylan Patel:It took a year for that to actually reflect in memory prices. Once memory prices reflected, then it took another six months, three months for the memory vendors to start building fabs. And those fabs take two years to build. And so we don't have really meaningful fabs that you can even put these tools in until late 27 or 28, right? And so instead what you've seen is like some really crazy stuff to get capacity, right? Right. Micron bought a fab from a company in Taiwan that makes lagging edge chips. Right. Hynyx and Samsung are doing, you know, some pretty crazy things to try and expand capacity at their existing fabs that also have like very large knock on effects in the economy.

1:30:56Dylan Patel:And so, hey, why can't we build more capacity is like there's nowhere to put the tools. Right. And it's not just EUV. There's other tools involved in DRAM and logic. Right. Like logic, you know, N3, 30 % or so of the cost, you know, 28 % of the cost is EUV of the wafer, of the final wafer. When you look at like DRAM, it's in the teens. And it's going up, but it's in the teens. So it's as much smaller percentage of the cost as DRAM or is EUV. These other tools are also bottlenecks, although their supply chains are not as complex as ASMLs. And so you see Applied Materials and LAM Research and all these other companies also expanding capacity a lot.

1:31:34Dylan Patel:And anyways, you don't have anywhere to put the tool because the most complex building that people make is fabs. And fabs take two years to build.

1:31:42Dwarkesh Patel:You can think of Jane Street as a research lab with a trading desk attached. Their infrastructure team has built some of the biggest research clusters in the world with tens of thousands of high-end GPUs and hundreds of thousands of CPU cores and exabytes of storage. This compute is part of how Jane Street surfaces all the hidden patterns that are embedded in incredibly noisy market data. Even beyond the noise, the nature of the signal changes constantly in reaction to things like pandemics and elections and new regulations, and even changes in sentiment. There's this unremitting game of trying to figure out whether your old models still reflect the real world, and if not, what to do about it.

1:32:15Dwarkesh Patel:If you're interested in working on this sort of thing, Jane Street is hiring ML researchers and engineers. They're also accepting applications for their summer ML internship program, with spots in London, New York, and Hong Kong. And if you happen to find yourself at GTC, which is happening the week after this episode drops, Jane Street's GPU performance team is giving a talk. Go to jainestreet.com slash thwarkash to learn more. I interviewed Elon recently, and his whole plan is that I guess they're going to build this gigafab, terafab, some power of 10. And they're going to build the clean rooms.

1:32:50Dwarkesh Patel:I won't even ask you about the dirty rooms thing, but let's say they build the clean rooms. and for okay i have a couple questions one do you think this is the kind of thing that elon co could build much faster than people are conventionally building it but this is not about building the end tools this is just about building the facility itself how complicated is it to just build a clean room and do it extremely fast is this something that like elon with this move fast thing could do much faster if that's what we're bottlenecked on this year or next year and two um does that even matter if in two years your view is that we're not bottlenecked on clean room space where we're bottlenecked on the tooling.

1:33:26Dylan Patel:So I think, you know, as with any complex supply chain, it takes time and constraints shift over time. And even if something isn't any longer a constraint, that doesn't mean that market no longer has margin, right? So for example, energy will not be a big bottleneck as we get to, you know, a couple of years from now. But that doesn't mean energy is not growing super fast and there's no margin there. It's just like, it's not the key bottleneck. And in the space of fabs, right, clean rooms are the biggest bottleneck this year and next year. And as we get over time, 29, you know, 28, 29, 30, there will be still constraints there.

1:33:57Dylan Patel:The thing about Elon is I think he's had a tremendous capability to garner physical resources and really smart people to build things. And the way he's able to recruit really amazing people is just try and build the craziest stuff, right? In the case of AI, that's not really worked because everyone's trying to build AGI. Everyone's very ambitious. But in the case of like, we're going to make, you know, we're going to go to Mars and we're going to make rockets that land themselves, or we're going to make fully autonomous cars that are electric, right? Or we're going to make human aid robots, right?

1:34:25Dylan Patel:Like these are methods of recruiting the people who think that's the most important problem in the world to work on that problem because he's the only one trying really hard. In the case of semiconductors, I want to make a fab that's a million wafers per month. No one has a fab that big. That's what he stated, right? He wants to make a million wafers a month. You know, it's possible that he's able to recruit a lot of really awesome people and get them on this heroicly, you know, this crazy task of trying to build a fab that does a million wafers per month. Step one is to build the clean room. And I think that he probably can do, right?

1:34:52Dylan Patel:I think, you know, there's some mindset, you know, his, his mindset around like delete things. It can be dirty. It's fine. Probably not right. Or actually, I think 100 % it's not right. You like need the fab to be very clean. I think the entire air, the entire, all of the air in the fab gets replaced like every three seconds. It's like that fast. And there's so few particles per, but I think he can build the clean room. It'll take a year or two, maybe initially it won't be super fast, but then over time we'll get faster and faster at it. But then the really complex part is actually developing a process technology and building wafers.

1:35:23Dylan Patel:And I don't think he can develop that quickly. I think that has a lot of built up knowledge. It's, again, like the most complicated, like integration of very expensive tools and supply chain that's done is a TSMC or an Intel or a Samsung. And some of these two other two companies aren't even that great. And they're like tremendously complex.

1:35:42Dwarkesh Patel:How surprised would you be if in 2030, there just happened to be some total disruption? We're not using EUV. We're using something that has much better effects, is much simpler to produce. We can produce in much bigger quantities. I'm sure as an industry insider, that sounds like a totally naive question. But do you see what I'm asking? What probability should we put on, oh, something totally out of the left field comes out and none of this is relevant?

1:36:07Dylan Patel:Something that's very simple and easy to scale, I have very, very low probability for. There are a number of companies working on effectively like particle accelerators or synchotrons that generate light that's either 13.5 nanometer like EUV or even x-ray, like even narrower wavelength, like 7 nanometer or whatever wavelengths of light to then use in lithography tools. But those things are like massive particle accelerators that are then generating this light. It's a very complicated thing to build. So there's a couple of companies, and I think that that could be a big disruption to the industry beyond what EUV is.

1:36:38Dylan Patel:I don't necessarily think that like we're going to just magically build something new that is like direct write and super simple and can be manufactured at huge volumes. Although there are some attempts to do things like this.

1:36:50Dwarkesh Patel:Yeah. Because I ask because if you think about Elon codes in the past, rocketry was this thing that was thought to be. I mean, it is incredibly complicated.

1:36:59Dylan Patel:Look, I'm just a naive yapper compared to Elon, right? What have I built? So maybe it's possible, right? Yeah, yeah.

1:37:05Dwarkesh Patel:In order to be able to build more memory in the future, could we build 3D DRAM the way we do 3D NAND and then go back to DUV?

1:37:15Dylan Patel:This is the hope. Currently, everyone's roadmap for 3D DRAM is that you'll still use EUV because you want to have that tighter overlay. Because now when you're doing these subsequent processing steps, you want it to be, you know, everything is vertically stacked. You have more layers on top of each other and you want the pitches to be tighter and all these things. So generally people are still trying to do an EUV, but what 3D would do is it would take the, you know, hey, a single EUV pass, how many bits can it make, right, if you do this sort of like calculation? And that number would go up drastically if you go to 3D DRAM.

1:37:45Dylan Patel:That is the hope. But right now everyone's roadmap is sort of like you go from current, it's called a 6F cell to a 4F cell. And then finally 3D DRAM, like by the end of the decade or early next decade. So there's still like a lot of R &D and manufacturing and integration to be done. I wouldn't call that out of the cards I think it's very much likely going to happen it also is going to require a huge retooling of fabs right the breakdown of tools in a fab are very different right actually the lithography tool is the only thing that isn't like that different but the number of them relative to different types of chemical vapor deposition or atomic layer deposition or dry etch or different kinds of etch chambers with different chemistries all of these things you have all these different kinds of tools for different process nodes.

1:38:31Dylan Patel:You can't just like convert a logic fab to a DRAM fab or vice versa back and forth or a NAN fab to a DRAM fab in a short amount of time. And in the same way, existing DRAM fabs require a lot of retooling just to go from 1B or 1 alpha to 1 beta to 1 gamma process nodes because now they have to add EUV and change the chemistry stacks for when you're using EUV in terms of deposition and etch and the EUV tool has to be there. And furthermore, like when you change to 3D DRAM, there's going to be an even larger shift. And so there's a lot of retooling of these fabs that needs to happen. in terms of the tools.

1:39:00Dylan Patel:And so that would be a big disruption. That would make EUV demand generally lower. But as we've seen across time, EUV demand as a percentage of wafer costs has trended up initially, or lithography, right? Lithography initially, I want to say in like 2014-ish era, was like 16 % of the wafer cost, 17%. And it's gone to 30 over the last 15 years. And for DRAM, it was in the mid-teens as well, or low-teens, and now it's trended towards the high-teens. and before we get to 3D DRAM, it'll likely cross into the 20s percentage range. But then if we get to 3D DRAM, it tanks again in terms of the total end wafer cost as a percentage of EUV.

1:39:39Dwarkesh Patel:Yeah. I guess you care less about the percent of cost and more about how much it bottlenecks.

1:39:44Dylan Patel:Right, but the percentage of cost is sort of... A proxy, yeah, yeah.

1:39:46Dwarkesh Patel:Yeah. So if you're Jensen or Sam Waltman or whoever who stands to gain a lot from scaling up AI compute, there's these stories that they'd go to TSMC and say, hey, why can't we actually Y and Z? But I think the point you're making here is it doesn't really matter in some sense what TSMC does. And in fact, even if you have Intel and Samsung building more foundries, in the long run, you're still going to be bottlenecked by ASML and other toolmakers and other material makers. So first, is that correct interpretation? And second, then why should basically Silicon Valley people be going to the Netherlands to try to pitch ASML?

1:40:24Dwarkesh Patel:Like right now, should they be trying to pitch ASML to make more tools so that like in 2030, they can have more AI compute.

1:40:29Dylan Patel:You know, it's a funny dynamic we saw in 2324 and 2025. People who saw the energy bottleneck before others asymmetrically went to, you know, Siemens, Mitsubishi, and of course, GE Vernova and bought up turbine capacity. And now they're able to charge excess amounts for deploying these turbines places because of energy. And in the same sense, this could be done for EUV, except ASML is not just going to trust any random bozo who wants to buy EUV tools in the sense that like, you know, these turbines are much cheaper than EUV tools and there's many more of them produced, right? Especially once you like get to like industrial gas turbines or like, you know, not just combine cycle, but like the cheaper, smaller, et cetera, less efficient ones.

1:41:12Dylan Patel:People put down deposits for these. So in a sense, someone could do this, right? Someone should go to the Netherlands and be like, I'll pay you a billion dollars. you give me the right to purchase 10 EUV tools two years from now, right? And I'm first in line two years from now. And then over those two years, you then go around and wait for everyone to realize, oh crap, I don't have enough EUV tools. And then you try and sell your option at some premium. But all you're effectively doing is you're saying, ASML, you're dumb. You weren't making enough margin on these. I'm going to make a margin. And the question is like, well, will ASML even agree to this?

1:41:47Dylan Patel:Right? And I'm like, I don't think so.

1:41:49Dwarkesh Patel:Right? So there's a world where they at least get the demand signal from that to increase production.

1:41:53Dylan Patel:Potentially. Potentially. I agree.

1:41:55Dwarkesh Patel:But it sounds like you're saying, oh, they couldn't even increase production if they wanted to, given the supply chain.

1:41:59Dylan Patel:Right, but that's exactly the market in which if they can't increase production, just like TSMC cannot increase production that fast, and yet demand is mooning, then the obvious solution is to arbitrage this because you and I know demand is way higher than they're projecting and their capability to build. So then you arbitrage this by locking up the capacity and then sort of doing like a forward contract and then trying to sell it at a later date once other people realize actually shit, everything is fucked and we don't have enough capacity. And then you'll have like this insane margin that ASML and TSMC should have been charging.

1:42:30Dylan Patel:But the thing is, I don't know if ASML and TSMC will ever agree to this.

1:42:34Dwarkesh Patel:Okay, let me ask about power now. So it sounds like you think power can be arbitrarily scaled. Not arbitrarily, but yes. But beyond these numbers. And I think, if I'm remembering correctly, your blog post on the power, how AI lives are increasing power, you were like, well, you were implying that Giavanova and Mitsubishi and Simons could produce and gas turbines was like 60 gigawatts a year. And then there's other sources, but they're like less significant than the turbines. And so in only a fraction of that goes to AI, I assume. So, yeah, if in 2030 we have enough logic and memory to do 200 gigawatts a year, do you just think that these things are on a path to ramp up to more than 200 gigawatts a year?

1:43:18Dwarkesh Patel:Or what do you see?

1:43:19Dylan Patel:Yeah. So, I mean, right now we're at 30, right? Or 20. So this is critical IT capacity, by the way, right? This is an important thing to mention. When I'm talking about these gigawatts, I'm talking about critical IT capacity, server plugged in, that's how much power it pulls. But there's losses along the chain, right? There is loss on the transmission. There's losses on the conversion. There's losses on cooling, et cetera. And so you should gross this factor up from 20 gigawatts for this year or 200 gigawatts by the end of the decade to some number 20, 30 percent higher. And then you have capacity factors, right?

1:43:52Dylan Patel:Turbines don't run at 100 percent. In fact, if you look at PGM, which is the largest grid, I think, in America, sort of the Midwest, sort of Northeast kind of area-ish, not the full Northeast. But anyways, PJM, they rate in their models for like, hey, turbines, how much capacity? We want to have excess, you know, roughly 20 % capacity. In addition, in that 20 % excess capacity, we're running all the turbines at 90 % because they are derated some for reliability. Oh, things go down, maintenance, et cetera, et cetera, et cetera. So then in reality, the nameplate capacity for energy is always way higher than the actual end critical IT capacity because of all of these factors.

1:44:30Dylan Patel:But it's not just turbines, right? If you're just making power from turbines, that's simple, boring, easy. Humans and capitalism is far more effective. And so the whole point of that blog was, yes, there's only three people making combine cycle gas turbines, but there's so much more we can do. We can do aeroderivatives. We can take airplane engines and turn them into turbines as well. And there's even new entrants to the market, like Boom Supersonics trying to do that. And they're working with Crusoe. And also there's all the other ones that already exist in the market. there's um there's medium speed reciprocating engines right engines that spin in circles right so sort of like any diesel engine right there's like 10 people who make engines that way right so cummins you know you know at least i'm from georgia and we we you know people used to be like oh man you got a cummins engine in there um you know like you know regarding ram trucks but it's like well actually auto automobiles manufacturing is going down these companies all have capacity and could scale and convert that to for data center power right stick all these reciprocating engines yes It's not as clean as Combine Cycle.

1:45:30Dylan Patel:Maybe you can convert them from diesel to gas if you want. But at the end of the day, these spinning engines, oh, what about ship engines, right? All of these engines for these massive cargo ships, those are great. Nebius is doing that for a data center for Microsoft in New Jersey, right? They're running these ship engines to generate power. Oh, there's, you know, Bloom Energy's doing fuel cells. We've been like very positive on them for like a year and a half now because they have like such a capability to increase their production. And their payback period for production increase is like very fast, even if the cost is a little bit higher than combine cycle, which is like the best cost and efficiency.

1:46:05Dylan Patel:You know, and then there's solar plus battery, which as these cost curves continue to come down, those can come online. There's wind. And, you know, of course, the derating of those, you know, hey, when you put on a wind turbine, you might say, oh, I'm only going to expect 15 percent of the maximum power because things just oscillate. But you add batteries. There's all these things. And then the other thing is that like the grid is scaled for, you know, hey, we are not going to cut off power at peak usage, which is like the hottest day in the summer. But in reality, that's a load spike that is 10, 15, 20 percent higher than the average.

1:46:37Dylan Patel:Well, if you just put enough utility scale batteries or you put peaker plants that only run a small portion of the year, then all of a sudden, you know, and those could be gas, they could be industrial gas turbines, they could be combine cycle, they could be any of the other sources of power I mentioned, they could be batteries, then all of a sudden you've unlocked 20 % of the U.S. grid for data centers because most of the times that capacity is sitting idle and it's really only there for that peak, right? Which is a day or two, right? And it's a few hours of like maybe a few days of the full year is that peak.

1:47:08Dylan Patel:And so you just have enough capacity to absorb that peak load and all of a sudden you've transferred all. And today data centers only 3-4 % of the power of the U.S. grid. And by 28, there'll be 10%. But if you can just unlock 20 % of the U.S. grid like this, it's not that crazy. And the U.S. grid is terawatt level, not hundreds of gigawatts level. So we can add a lot more energy. It's not easy. I'm not saying it's easy. These things are going to be hard. There's a lot of hard engineering. There's a lot of risks that people have to take. There's a lot of new technologies people have to use. But Elon was the first to do this behind the meter gas.

1:47:43Dylan Patel:And since then, we've seen an explosion of different things that people are doing to get power. And they're not easy, but people are going to be able to do them. And the supply chains are just way more simple than chips.

1:47:56Dwarkesh Patel:Interesting. So I guess he made the point during the interview that the specific blade for the specific turbine he was looking at, the lead times for that go out beyond 2030. And your point is that... That's great. There's so many other ways to make energy. Just be inefficient. It's fine. Right. So you're like right now, I guess, combined cycle gas turbines have capex of$1 ,500 per kilowatt. And you're saying you could just it would make sense to have either technologies that are much more expensive than that or other things are getting cheap enough to that to make it competitive.

1:48:24Dylan Patel:Exactly. Exactly. You know, it can be as high as$3 ,500 per kilowatt even. Right. So it could be twice as much as the cost of combined cycle. And the total cost of the GPU, you know, you know, on a TCO basis has gone up a few cents per hour. Right. Right. Again, because we've been talking about hopper pricing,$1.40 now becomes, you know, oh, the power price doubles. OK, the hopper that was$1.40 is now$1.50 in cost. Right. It's like, oh, I don't care because the models are improving so fast that the marginal utility of them is worth way more than that 10 cent increase in energy. Okay.

1:49:00Dwarkesh Patel:And then so you're saying 20 % of the grid, so one terawatt about, 20 % of that can just come online from utility-scale batteries increasing what you'd be comfortable putting on the grid.

1:49:11Dylan Patel:The regulatory mechanism there is not easy, by the way.

1:49:13Dwarkesh Patel:But that's 200 gigawatts, if that hypothetically happens. But you're saying on just from the different sources of gas generation you mentioned, the different kinds of engines and turbines, combined, how many gigawatts could they unlock by the end of the decade?

1:49:27Dylan Patel:Yeah, so we're tracking in some of our data where there's over 16 different manufacturers of power generating things just from gas alone, right? So, yes, there's only three turbine manufacturers for a combined cycle, but we're tracking 16 different vendors and we have all of their orders and things like that. And it turns out there is just hundreds of gigawatts of orders to various data centers. As we get to the end of the decade, we think like something like half of the capacity that's being added will be behind the meter. and when we look at like a lot of this is actually behind the meter is almost always more expensive than grid connected but there's just a lot of problems with getting grid connected and you know permits and interconnection queues and all this sort of stuff so it ends up being even though it's more expensive people are doing behind the meter and then what they're doing behind the meter with ranges widely right it could be reciprocating engines it could be ship engines it could be aerosol derivatives it could be combined cycle although combine cycle is not that great for behind the meter it could be bloom energy fuel cells it could be solar plus battery, right?

1:50:26Dylan Patel:Like it could be any of these things.

1:50:28Dwarkesh Patel:You're saying any of these individually could do like tens of gigawatts?

1:50:32Dylan Patel:Any of these individually will do tens of gigawatts and in a whole they will do hundreds of gigawatts.

1:50:36Dwarkesh Patel:Okay.

1:50:37Dylan Patel:So that alone should more than... I mean, it's going to take, I mean, like electrician wages probably double or triple again, right? And like there's going to be a lot of new people entering that field and there's going to be a ton of people who make money, but it is something that I don't, like I don't see that as the main bottleneck, right?

1:50:52Dwarkesh Patel:So right now in Abilene, the 1.2 gigawatt data center that Caruso is building for OpenAI, I think they have like 5 ,000 people working there, or at peak they did. And if you turn that into 100 gigawatts, and I'm sure things will get more efficient over time, but that would be like 400K people it would take to build 100 gigawatts. And if you think about the U.S. labor force of how many electricians there are, how many construction workers there are. Yeah, I guess there's like 800K electricians. I don't know if they're all substitutable in this way. There's millions of construction workers. But if we're in a world where we're adding 200 gigawatts a year, are we going to be crunched on labor eventually?

1:51:34Dwarkesh Patel:Or do you think that is actually not a real constraint?

1:51:37Dylan Patel:So labor is a humongous constraint in this. People have to be trained. Likewise, we probably start importing the highest skilled labor in this way, right? Because now it makes sense that, you know, hey, a really high skilled electrician in Europe who was working on destroying power plants now comes to America and is building data center, you know, high voltage electricity, you know, power moving across the data center, right? Something like this, right? Humanoid robots maybe start to or robotics at least start to. But the main factor is going to be for reducing the number of people is modularizing things and making them in factories in Asia, unfortunately, but, you know, at least for America, but, you know, Korea, Southeast Asia, in many ways, China as well.

1:52:21Dylan Patel:But, you know, these areas are going to do, are going to ship more and more built out sections of the data center and those will be shipped in, right? Maybe today you, you know, you currently ship servers in or a rack in, and then you plug that into, you know, different pieces that you're shipping from different places. But now you'll ship it to a factory and integrate the entire, you know, hey, maybe this is a two megawatt block. And this block goes from, you know, high voltage power to the, you know, the voltage power that you, the voltage and maybe DC that you deliver to the rack instead of being AC and high voltage, right?

1:52:59Dylan Patel:Or something like this, right? Or cooling, you take, you ship a fully integrated thing that has a lot of the cooling subsystems already put together. Or, because plumbers are also a big constraint here, or furthermore, you take, instead of just a single rack, and now you have people wiring up all these racks of power and electricity and blah, blah, blah, blah, blah, you take a skid and you put an entire row of servers, and that is shipped from the factories. And today, a single rack may be 120, 140 kilowatts, but as we get to next generation, NVIDIA, Kyber, and things like that, it's almost a megawatt.

1:53:34Dylan Patel:And then in addition, if you do an entire row, it'll have the rack, it'll have the networking and it'll have the cooling and the power racks all integrated together. So now when you come in, actually you have much less stuff to cable, whether it be networking with a fiber, whether it be the power, right? There's fewer power, that power things to connect. And then there's fewer plumbing things to connect, right? And so this drastically can reduce the amount of people working in data centers. And therefore the capability to build these will be much larger. And along the way, there will be, you know, new things mean, you know, some people move faster to new things, some people move slower, right?

1:54:09Dylan Patel:Crusoe and Google have been talking a lot about this modularization, as has people like Meta and, you know, many others, right, have been talking a lot about this modularization. And others are going to be slower to doing it. But at the end of the day, you know, and people who move faster to new things may have more delays, or people who are slower have labor problems. So there will always be dislocations in the market, because this is a very complex supply chain. At the end of the day, it's still simple enough that we will be able to solve it through capitalism and human ingenuity on the time scales that are required.

1:54:38Dwarkesh Patel:Yeah. Okay. So speaking of big problems to solve, Elon Musk is very bullish on space GPUs. If you're right, that power is not a constraint on Earth. I guess the other reason that would make sense is that even you can, there is enough, there'll be enough gas turbines or whatever to build it on Earth. I think Elon's next argument then is like, You can't get the permitting to build hundreds of gigawatts on Earth. Do you buy that argument?

1:55:02Dylan Patel:Land-wise, it's pretty – America's big. Data centers don't take that much space. You can solve that. Permitting-wise, air pollution permits are a challenge, but the Trump administration has made it much easier. You go to Texas and you can skip a lot of this red tape. And so Elon had to deal with a lot of like this complex stuff in Memphis and then building a power plant across the border. and all these things for Colossus 1 and 2. But at the end of the day, there's a lot more you can get away with in the middle of Texas, right?

1:55:32Dwarkesh Patel:Given that Elon lives in Texas, why didn't he just go to Texas?

1:55:34Dylan Patel:I think it was partially like they over-indexed on grid power for a temporary period of time, right? Because that's just what they thought they needed more of.

1:55:42Dwarkesh Patel:But you said an Illumina refinery connected to the grid there.

1:55:45Dylan Patel:It was an appliance factory that was idled. But I think they may have indexed more to what was grid power. They may have indexed more to like water access and gas access because actually I think they bought that knowing that the gas line was right there and they were going to tap it. Same with water. It was a whole host of different constraints. It was probably an area where electricians and things like that were easier to find. But at the end of the day, I'm not exactly sure why they chose that site. I bet Elon would have chosen somewhere in Texas if he could have like gone back. But yeah, because of the regulatory faces he's challenged, challenges he's faced.

1:56:20Dylan Patel:It's ultimately like permitting is a challenge, but America is a big place and there are 50 states and things will get done. And there are a lot of small jurisdictions where you can just transport in all the workers that you need for a temporary period of six months to a year, depending on the type of contractor. It can be even three months for depending on the type of the contractor that's coming in and put them in temporary housing, pay out the butt because labor is very cheap relative to the GPUs and the power or not the power, but the GPUs and the like the networking and so on and so forth and the end value of the tokens it's going to produce.

1:56:52Dylan Patel:So all of these things have plenty of room to be paid for. And so I think it's fine. And also people are diversifying now. Australia, Malaysia, Indonesia, India, these are all places where data centers are going up at a much faster pace, but currently still 70 % plus of the AI data centers are in America. And that continues to be the trend. And so I think people are figuring out how to build these things and permitting. Like I just like ultimately like permitting and red tape in middle of nowhere, Texas or middle of nowhere, Wyoming or middle of nowhere like New Mexico is probably a hell of a lot easier than sending stuff into space.

1:57:30Dwarkesh Patel:Right. Well, other than the fact that the economic argument makes less sense once you consider the fact that energy is a small fraction of the cost of ownership of a data center. What are the other reasons you're skeptical?

1:57:41Dylan Patel:Yeah. So obviously power is free in space, basically.

1:57:44Dwarkesh Patel:That's the reason to do it.

1:57:45Dylan Patel:Yeah, that's the reason to do it. But then there's all the other counter arguments, right? Which is because even if power costs double, you're still at a fraction of the total cost of the GPU. The main challenges is. And what we've seen that disperses, right, we have ClusterMax, which rates all the NeoClouds and we test them. We test over 40 cloud companies, including the hyperscalers and NeoClouds. What differentiates some of these clouds the most outside of software is their ability to deploy and manage failure, right? GPUs are horrendously unreliable. Even today, 15 % of Blackwells or so that get deployed have to be RMA'd.

1:58:20Dylan Patel:You have to take them out. You have to maybe just plug them in and plug them back in. But sometimes you have to take them out and ship them to NVIDIA or rather their partners who do these RMAs and such.

1:58:28Dwarkesh Patel:What do you make of Elon's kind of argument that once you have the initial – after initial phase, they actually don't fail that much? Sure.

1:58:34Dylan Patel:But now you've done this. You've tested them all. You deconstructed them, put them on a spaceship, fucking put them into space, and then put them online again. That's months, right? And if your argument is that, hey, GPUs have a useful life of X years, if a GPU has a useful life of five years and it takes three additional months, probably six, let's say six additional months, then that is 10 % of your cluster's useful life. And because we're so capacity constrained, that compute is most valuable, theoretically, in the first six months you have it because we're more constrained now than in the future because that compute now can contribute to a better model in the future or can contribute to revenue now, which you can use to raise more money to get better, you know, all these sorts of things.

1:59:18Dylan Patel:Now is always the most important moment. And so you've delayed your compute deployment by six months potentially. And the thing that separates these clouds is we see clouds that take six months to deploy GPUs today on earth, right? We see clouds that take a lot less than six months. And so the question is, where does space get in there? I don't see how you would test them all on Earth, deconstruct them and ship them and shoot them into space and it not take longer than just putting them in the spot that you were testing them.

1:59:44Dwarkesh Patel:Yeah. So the question I wanted to ask is the topology of space communication. So right now, Starlink satellites talk to each other at 100 gigabits per second. and you could imagine that being much higher with optical inter-satellite laser links that are optimized for this. And that actually ends up being like quite close to the InfiniBand bandwidth, which is like 400 gigabytes a second, right?

2:00:09Dylan Patel:But that's per GPU, not per rack.

2:00:11Dwarkesh Patel:I see, okay.

2:00:12Dylan Patel:So multiply that by 72. Also like that was Hopper when you go to Blackwell and Rubin, that two X's and two X's again.

2:00:19Dwarkesh Patel:All right. But how much compute is happening per, like during inference, are the different scale-ups still working together or is it just happening, it's a batch within a single scale-up?

2:00:31Dylan Patel:A lot of models fit within one scale-up domain, but many times you split them across multiple scale-up domains. I think that you really have to, as models become more and more sparse, at least this is like the general trend, then you want to ping just a couple experts per GPU. and if leading models today have hundreds if not thousand experts then you'd want to run this across hundreds of chips or thousands of chips even as we continue to advance into the future and so then you end up with this problem of well now you need to you know need to connect all these

2:01:06Dwarkesh Patel:satellites together comms wise as well okay so that would be tough because i was imagining if there's a world where you could like do a batch inference for a batch on a single uh a scale up then maybe it's more plausible but if not then it's it's yeah i mean networking these ships

2:01:22Dylan Patel:together is a problem and and you can't just make this the satellite infinitely large right like there are a lot of challenges with physics to making a satellite really big right so then these inner that's why you need these inner interconnects between the satellites those interconnects are more expensive than the you know a cluster like 20 of the cost or 15 of the cost is networking all of a sudden now you're making it like space lasers instead of like pretty simple like lasers that are manufactured in millions of volumes with, you know, pluggable transceivers. And those things are very unreliable as well.

2:01:52Dylan Patel:More unreliable than the GPUs, by the way. Across the life of a cluster, you have to unplug, clean it all the time, right? Unplug, replug it just for random reasons. These things are just not as reliable. So you've got that problem as well. Like you've got a more expensive, complicated space laser to communicate instead of this pluggable optical transceiver that's been in super high volume. Okay.

2:02:11Dwarkesh Patel:So all in all, what does that imply for space data centers?

2:02:13Dylan Patel:So space data centers effectively are not limited by, you know, hey, we have this energy advantage. It's actually just limited by the same contended resource. We can only make 200 gigawatts of chips a year by the end of the decade. So what are we going to do to get that capacity? It doesn't matter if it's on land or in space. It doesn't really matter, right? Because you can build that power. And I think human capabilities and capacity could get to the period where we're adding a terawatt a year globally of various types of power. At some point, we do cross the chasm where space data centers make sense, but it's not this decade, right?

2:02:52Dylan Patel:It is much further out once you have energy constraints actually being a big bottleneck, once you have space, land permitting be a much bigger bottleneck as it subsumes more and more of the economy. And chips are no longer the bottleneck. because chips are the biggest bottleneck. And so you want them deployed working on AI the moment they're done being manufactured. And so there's a lot of things people are doing to increase that speed faster and faster, whether it be modulizing data centers or even modulizing racks where you actually put the chip in at the data center, but only the chip and everything else is already wired up and ready to go at the data center.

2:03:29Dylan Patel:So there's things like this that people are doing to decrease that time that you cannot do in space. And at the end of the day, all that matters in a chip-constrained world is get these chips working on producing tokens ASAP in a world, you know, maybe 2035, once the semiconductor industry and ASML and Zeiss and all these other suppliers, land research applied materials, fab manufacturers, like pendulum swings and are able to make enough chips. And really, we're optimizing every dial. And like, it makes sense to optimize the 10 % of energy costs or 15 % of energy costs, or as we move to ASICs potentially, and NVIDIA's margins aren't 70 plus percent, maybe that energy cost is 30 % of the cluster and fab construction, all this.

2:04:10Dylan Patel:Like these are the things, our data center construction, these are the things to optimize. But that's not a, you know, Elon doesn't win by doing, you know, 20 % gains. Elon never wins that way. Elon wins when he swings for the fences and does 10X gains, right? That's what SpaceX is about. That's what Tesla was about. That's what all of his success has been about, right? It's not been about these chasing the 20%. So I think space data centers will eventually be a 10x gain, potentially, as Earth's resources get more and more contentious. But that's not this decade.

2:04:40Dwarkesh Patel:Yeah. I mean, I think just to drive some intuition about how much land there is on Earth, obviously the chips themselves, especially if you move to a world where you have racks that have megawatts, megawatt-y charts, like literally it's not even a random factor.

2:04:52Dylan Patel:That's the other thing, right? The power density, you know, if chips and manufacturing is the constraint, right now roughly it's one watt per millimeter squared for AI chips and such. one easy way is to pump that to two watts per millimeter squared. Now, you may not get 2x the performance. You may only get 20 % more performance. And that requires much more exotic cooling, right? It requires more complicated cold plates and very complicated liquid cooling, or maybe it requires things like immersion cooling. But in space, higher watts per millimeter is very difficult, whereas on Earth, these are solved problems.

2:05:22Dylan Patel:And one of these things enables you to get a lot more tokens. Maybe it's 20 % more tokens per wafer that's manufactured. And that's a humongous way.

2:05:32Dwarkesh Patel:So a millimeter, you mean of dye area?

2:05:34Dylan Patel:Yeah, of dye area. Square millimeters of dye area.

2:05:36Dwarkesh Patel:I mean, it would be better for space because if you can run more watts per millimeter would be the chip runs hotter and the hotter the chip. I guess this is a question of computer chip engineering. But it cools to the power of forth by Stefan Boltzmann's law. So if you can run a very hot chip because it allows a lot to go. No, no, but you can't run it hotter.

2:05:51Dylan Patel:You can only run it denser. And the problem is getting the heat out of that dense area means you have to move away from standard air cooling and liquid cooling to more exotic forms of liquid cooling or even immersion to get to higher power densities. And that's more difficult in space than it is on Earth.

2:06:08Dwarkesh Patel:Yeah. And maybe it's at this point worth explaining what exactly a scale up is and what it looks like for NVIDIA versus Tranium versus TPUs. Yeah.

2:06:20Dylan Patel:So earlier I was mentioning how communication within a chip is super fast. Communication within chips that are in the same rack is fast, but it's not as fast. And then, you know, it's on the order of terabytes and then communication very far away is on the order of gigabytes, hundreds of gigabytes, right? So this order of magnitude, as you get further distance compute and maybe across the country, it's on the order of gigabytes a second, right? Scale up domain is this like tight domain where the chips are communicating on the order of terabytes a second. And so for NVIDIA, previously, this meant an H100 server had eight GPUs, and those eight GPUs could talk to each other at terabytes a second.

2:06:58Dylan Patel:With Blackwell and VL72, they implemented rack scale up, and that meant all 72 GPUs in the rack could connect to each other at terabytes a second speed. And the speed doubled gen on gen, but also the most important innovation they did was going from 8 to 72 in the domain. When we look at Google, their scale up domain is completely different, right? It has always been on the order of thousands, right? With TPU V4, they had pods the size of 4 ,000 chips. With V8, they have pods, or V7, they have pods in the 7 ,000, or sorry, 8 ,000, 9 ,000 range. And what's relevant here is that it's not the same as NVIDIA.

2:07:34Dylan Patel:It's not like for like. Google has a topology that's a Taurus, right? So every chip connects to six neighbors. Rather than NVIDIA, the 72 GPUs connect all to all, right? So they can send terabytes a second to each other, to any arbitrary other chip in that pod of scale up. Whereas Google, you have to bounce through chips, right? So this means if TPU one needs to talk to TPU 76, then it has to bounce through various chips. And there is always some blocking of resources when you do that. So because that one TPU is only connected to six other TPUs. And so there's a difference in topology and bandwidth.

2:08:07Dylan Patel:And there are trade-offs and advantages of both, right? Google gets to have a massive scale up domain, but then they have the trade-off of you have to bounce across chips to get to from one chip to another. You can only talk to six direct neighbors. And so there is like this trade-off. And Amazon has mutated their scale-up domain. They're somewhere in between NVIDIA and Google effectively, where they're trying to make larger scale-up domains. They try and do all-to-all to some extent, which is what Switch is, which is what NVIDIA does. But also to some extent, they use Taurus topologies like Google does.

2:08:37Dylan Patel:And as we advance forward to next generations, All three of them are moving more and more towards a dragonfly topology, which means there's sort of like there is some fully connected elements and there's some elements that are not fully connected. So you can get the scale up to be hundreds or thousands of chips, but also have it not contend for resources when you're bouncing through chips. Related question.

2:08:58Dwarkesh Patel:I heard somebody make the claim that the reason that parameter scaling has been slow and only now are we getting bigger and bigger models from OpenAI and Anthropic is that original GPT-4 is over a trillion parameters. And only now are models starting to approach that again. And I heard a theory. The reason is that NVIDIA's scale-ups have just not had that much memory capacity. And so what was the claim exactly? If you have, say, one 5T model running at FP8, so that's 5 trillion gigabytes. Yeah. And then you have the KV cache. Let's say it's like... Just call it the same size. Okay, let's say it's the same size for one batch.

2:09:50Dwarkesh Patel:So you need 10 gigabytes, sorry, 10 terabytes to be able to run...

2:09:54Dylan Patel:A single forward pass, yeah.

2:09:56Dwarkesh Patel:And then only with the GB200 and VL72 do you have an NVIDIA scale-up that has 20 terabytes. And before that, they were much smaller. Whereas Google, on the other hand, has had these huge TPU pods that are not all-to-all, but still have, I think, hundreds of terabytes of capacity in a single scale-up. So does that explain why parameter scaling has been slow? I think it's partially the capacity and bandwidth, but also as you build a larger model, the ability to deploy it is slower, right?

2:10:24Dylan Patel:Like in terms of like, hey, what is the inference speed for the end user? That's kind of irrelevant. What's really relevant is RL. And what we've seen with these models and allocation of compute at a lab is sort of there's a few main ways you can allocate compute. You can allocate it to inference, i.e. revenue. You can allocate it to development, i.e. making the next model, and you can allocate it to research. And in development specifically, you split it between pre-training and RL, right? And so when you think about, hey, what exactly is happening? Well, the model, the compute efficiency gains you get from research are so large, you actually want most of your compute to go to research, not to development.

2:11:03Dylan Patel:because, you know, all these researchers are generating new ideas, trying them out, testing them, and continuing to march along this and push the Pareto optimal curve of scaling laws further and further and further. And at least what we've seen empirically is, like, model cost gets 10x cheaper every year or even more than that, which at the same scale gets 10x cheaper, or to get to reach new frontiers, it costs the same amount or more, right? So you don't want to allocate too many resources to pre-training and post-training. well, you actually want to allocate most of your resources to research.

2:11:36Dylan Patel:And then in the middle is this sort of this like development period. If you pre-train a 5 trillion parameter model, now you have to spend all this time. How many rollouts do you have to do in these RLs? And these rollouts for a trillion parameter model versus a 5 trillion parameter model are five times larger, which then means it takes, if you wanted to do as many rollouts, maybe the larger model is more sample efficient. Let's say it's 2x more sample efficient. Okay, great. Now you need two and a half X much time of RL to get the model smarter. Or you could RL the smaller model for two X the time and you'd still have a 25 % difference in the big model, which is two X more sample efficient and doing X number of rollouts versus the small model, which is a trillion parameters, although it's less sample efficient, is doing twice as many rollouts.

2:12:21Dylan Patel:It's still done faster. And so you get the model faster, sooner, and you've done more RL. And then you can take that model to help you build the next models, help your engineers train and do all these research ideas. And so this feedback loop is actually weighed towards smaller models in every case, no matter what your hardware is. And then as you look to Google, Google does deploy the largest production model of any of the major labs, right, with Gemini Pro. it is a larger model than gpt uh 5.4 it's a larger model than opus and and so you end up with yes google does this because they have a unipolar set of compute right almost all tpu um whereas anthropic is dealing with h100s h200s blackwell traniums tpus of various generations right and and uh open ai is dealing with mostly nvidia right now but going towards uh having amd and Tranium as well, the fleets of compute, like Google can just optimize around a larger model and they can leverage a thousand chips in a scale-up domain to get the RL speed much faster so that you can actually have this feedback loop be fast.

2:13:33Dylan Patel:But at the end of the day, in isolation, you almost always want to go with a smaller model that gets RL'd faster and gets deployed into research and development so you can build the next thing and get more compute efficiency wins. And then this compounding effect of, oh, I made a smaller model that I RL'd more that I then deployed into research and development earlier. And I spent less compute on the training itself because I was able to allocate more compute to the research. This like compounding effect of being able to do the research faster and faster and faster is potentially a faster takeoff.

2:14:03Dylan Patel:And that's all these companies want is fastest takeoff possible.

2:14:06Dwarkesh Patel:Okay. Spicy question. You know, you're explaining you make the semi-analysis sells these spreadsheets and you're always like, ah, six months ago or a year ago, we told people the memory crunch or now you're telling people the clean room crunch and then the future, the tool crunch. Why is Leopold the only person that is using your spreadsheets to make outrageous money? What is everybody else doing?

2:14:29Dylan Patel:I think there are a lot of people making money in many ways. I think obviously Leopold jokes that, you know, he's the only client of mine that tells me our numbers are too low. Everyone else tells me our numbers are too high, almost ad nauseum. You know, whether it's a hyperscaler saying, hey, that other hyperscaler, their numbers are too high. You know, and we're like, nah, that's it. And they're like, no, no, no, no, it's impossible, blah, blah, blah. And then you're like, finally have to convince them through all these facts and data when we're working with hyperscalers or AI labs that in fact, no, that number isn't too high.

2:14:58Dylan Patel:That's correct. But eventually, like sometimes it's like six months later, it takes them to realize or a year later. I think other clients like on the trading side also use our data, right? We sell data to a lot of, you know, I think roughly 60 % of my business is industry. So AI labs, data center companies, hyperscalers, semiconductor companies, you know, the whole supply chain across AI infrastructure. But then like 40 % of our revenue is like hedge funds, right? And, you know, I'm not going to comment on who our customers are, but I think a lot of people use the data. It's just how do you interpret it?

2:15:33Dylan Patel:And then what do you like view as beyond it. And I will say Leopold is pretty much the only person who tells me my numbers are too low always. And sometimes he's too high. Sometimes I'm too low. Right. But in general, I think other people are, you know, doing that. And you can check certain you can you can look across the space at hedge funds and look at their 13Fs and see actually they own maybe not exactly what Leopold does, because it's always like a question of like, what is the most constrained thing? What's the thing that's going to be that's most outside of expectations? And that's what are really trying to exploit is inefficiencies in the market.

2:16:06Dylan Patel:And in a sense, what our data shows is like making the market more efficient by making the base data of what's happening more accurate versus like, but in a sense, I think many, many funds do trade on information that is out there. And it's not, I don't think Leopold's the only person. I think he has the most conviction on the entire, in the entire, like about the AGI takeoff though, right? Right.

2:16:33Dwarkesh Patel:I mean, but the bets are not about like what happens in 2035. The bets that you're making that are at least exemplified by public returns we can see for different funds, including Leopold's, are about what has happened in the last year. And the last year stuff could be predicted using your spreadsheets, right? So it's like, it's less about, it's about buying like the next year of spreadsheets.

2:16:52Dylan Patel:Just spreadsheets. There's reports. There's API access to the data. There's a lot of data. But anyways, you know, I think.

2:16:57Dwarkesh Patel:Do you see what I mean? Like, it's not about some crazy singularity thing. It's about like, oh, do you buy the memory crunch?

2:17:02Dylan Patel:A simple one, though, is like you only buy the memory crunch if you believe AI is going to take off in a huge way. And the memory crunch, a lot of it was predicated on like, you know, at least for like people in the Bay Area who think about infrastructure, it's like obvious. KV cache explodes as context lenses go longer. So you need more memory. and then you do the math and you also have to have a lot of supply chain understanding of like what fabs are being built and what data centers are being built and how many chips and all these things. And so we track all these different data sets like very tightly.

2:17:31Dylan Patel:But at the end of the day, it takes, you know, someone to fully believe that this is going to happen. Like I think a year ago, if you told someone memory prices were quadruple and smartphone volumes are going to go down 40%, you know, over the year or two after that, people were like, you're crazy. That never happened. except a few people do believe that and those people did trade memory right and and people did i don't think like leopoldo is the only person buying like memory companies i think there are a lot of people buying memory companies he of course sized and positioned and did things in a better ways than some um maybe most right i i don't want to comment on whose returns or what um but he certainly did well um but other people also did really well right um you're trying to be like Like, wow, you've made me diplomatic for the first time ever.

2:18:19Dylan Patel:No, no, you're fine. I think it's hilarious, right? I'm being a diplomat, you know, whereas usually I'm, like, spicy. Yeah.

2:18:25Dwarkesh Patel:Okay. Maybe some rapid fire to close out. Can TSMC, if you're saying, look, the memory logic, et cetera, the N3 is mostly going to be AI accelerators, but then there's N2, which is mostly Apple now, and then in the future, I guess AI would also want to go on N2. So can they kick out Apple if NVIDIA and Amazon and Google say, hey, we're willing to pay a lot of money for N2 capacity?

2:18:59Dylan Patel:So I think the challenge with this is chip design timelines take a long while. And so that's more than a year. And the designs that are on 2 nanometer are more than a year out. And so what would really happen is NVIDIA and all these others will be like, hey, we're going to prepay for the capacity. And you're going to expand it for us. and then Apple would be, and maybe TSMC takes a little bit of margin, but not a ton, they're not going to kick Apple out entirely, right? What they're going to do is when Apple orders X, they may say, hey, we project you only need Y or X minus one. And so that's what we're going to give you is X minus one.

2:19:30Dylan Patel:And then that flex capacity Apple's kind of screwed on. Whereas traditionally Apple's always over-ordered by like 10 % and cut back by 10 % over the course of the year. And some years they hit the entire 10%, just, you know, volumes vary, right? Based on the season and macro, blah, blah, blah, blah, blah. And so I don't think TSMC would kick out Apple. I think Apple will become a smaller and smaller and smaller percentage of TSMC's revenue and therefore be less relevant for TSMC to cater to their demands. And TSMC could eventually start saying, hey, you've got to pre-book your capacity for next year for two years out and you have to prepay for the CapEx because that's what NVIDIA and Amazon and Google are doing.

2:20:07Dwarkesh Patel:Yeah, I wonder if it's worth going to specific numbers. I don't have any of them on the hand of like how many N2 wafers or what percentage of N2 does Apple have its hands on versus over the coming years versus AI?

2:20:21Dylan Patel:Yeah, I mean, this year, Apple has the majority of N2 that's going to get fabricated. There's a little bit from AMD. They are trying to make some AI chips and CPU chips early. There's a little bit. But for the most part, it's Apple. and as we go forward to the year after that, Apple still gets closer to like half of it as other people start ramping. But then it falls drastically, right? Just like for N3, they were half. We'll see. And when I say N2, that includes A16, which is a variant of N2. Over time, those nodes will be the majority. And what's also interesting is traditionally Apple's been the first to a process node.

2:21:00Dylan Patel:Two nanometers is actually the first time they're not. Well, besides Huawei, right? Huawei back in 2020 and before was the first with Apple, but they were both making smartphones. Now with two nanometer, you've got AMD trying to make a CPU and a GPU chiplet that they used advanced packaging to package together in the same timeframe as Apple. and this is a big risk for AMD that causes potential delays potentially because it's a brand new process technology it's hard but at the end of the day this is this is a bet that they want to do to you know scale faster than NVIDIA and try and beat them as we move forward actually when we move to the A16 node the first customer there is not even Apple it's AI and as we move forward that will become more and more prevalent not only will Apple not be the first to a node They will also not be the majority of the volume to the new node.

2:21:48Dylan Patel:And then they'll just be like any old customer. And because the scale of TSMC's CapEx keeps ballooning, but Apple's business is kind of not growing at the same pace, they become a less and less relevant customer. And they also will just cut their orders because things in the supply chain are kicking them out, whether it be packaging or materials or DRAM or NAND. These things are increasing in cost. They can't pass on all the cost to customers likely because the consumer is not that strong. And you end up with like this conundrum where they are just not Apple TSMC's best bud like they have been historically.

2:22:20Dwarkesh Patel:Do you think if Huawei had access to 3 nanometer, they would have a better accelerator than Rubin?

2:22:25Dylan Patel:Potentially, yeah. I think Huawei, they were the first with a 7 nanometer AI chip as well. They were the first with a 5 nanometer mobile chip, but they were the first with a 7 nanometer AI chip. The Huawei Ascend was like two months before the TPU and like four months before NVIDIA's, I want to say, was it V100 or A100? A100, I think. And so, you know, I mean, that's just moving to a process. No, that doesn't imply software. It doesn't imply hardware design, all these other things. But Huawei is arguably the only company in the world that has all the legs, right? Huawei has cracked software engineers.

2:23:03Dylan Patel:Huawei has cracked networking technologies. That's, in fact, their biggest business historically, right? And they have cracked AI talent. But furthermore, beyond NVIDIA, they actually have better AI researchers. And furthermore, beyond NVIDIA, they have their own fabs. And furthermore, beyond NVIDIA, they have their own end market of selling tokens and things like that. And Huawei tends to be like they're able to get the top, top, top talent. NVIDIA is as well, but not as in much concentration. And Huawei has a bigger pool in China. it's very arguable that Huawei, if they had TSMC, would be better than NVIDIA.

2:23:38Dylan Patel:And there are areas where China has advantages outside of areas that NVIDIA can't access as easily, right? Around not just scale, but also like some things around certain optical technologies China's actually really good at. So there's certain, I think it's very reasonable that if in 2019 that Huawei was not banned from using TSMC, Huawei had already eclipsed Apple as the biggest TSMC customer. And Huawei has huge share in networking and compute and CPUs and all these things. They would have kept gaining share and they'd likely be TSMC's biggest customer. Wow. That's crazy.

2:24:16Dwarkesh Patel:I've got kind of a random final question for you. So the other part of the Elon interview was robots. And so if humanoids take off faster than people expect, if by 2030, there's millions of humanoids running around, which each need local compute. Any thoughts on what that implies? What would be required for that?

2:24:36Dylan Patel:You know, there's a lot of like difficulties with like the VLMs and all these things that people, VLAs that people are deploying on robots. But to some extent, you don't need to have all the intelligence in the robot. And it would be much more efficient to not do that, right? Because in the server, in cloud, you can batch process and all these things. So what you may want to do is, hey, a lot of the planning and longer horizon tasks are determined by a much more capable model in the cloud that runs at very high batch sizes. And then it pushes those directions to the robots who then interpolate between each subsequent action or is given like, hey, pick up that cup.

2:25:14Dylan Patel:And then the model on the robot can pick up the cup. And it's like, as it's picking up, it's like, oh, you know, in fact, this, you know, you know, things like weight and all these things might have to be and like force may have to be like determined by the model on the robot, but not everything needs to be like, you know, hey, pick up the robot, you know, this, right? Or like, hey, that's a headphone. Actually, I'm the supermodel in the cloud. I know that this headphones are, you know, Sony XM6s, which is not a Dworkesh ad spot, but you know.

2:25:41Dwarkesh Patel:I'm like, why is this guy plugging this thing so hard? He's like on the table. He's like on his neck when we're entering Satya together. Like, is he getting paid by Sony?

2:25:51Dylan Patel:Unfortunately not. Unfortunately not. But anyways, like, you know, it might say, hey, the headband is soft and this is the weight of it and all these things. And then the model on the robot can be less intelligent and take these inputs and do the actions. And it may get told by the model in the cloud every second, every 10 times a second, maybe, you know, depends on the hertz of the action. But a lot of that can be offloaded to the cloud because otherwise, if you do all of the processing on the device, I believe it would be more expensive because you can't batch. Two, you couldn't have as much intelligence as you do in the cloud because the models will just be bigger in the cloud.

2:26:22Dylan Patel:And three, we're in a semiconductor shortage world and any robot you deploy needs leading edge chips because the power is really bad for robots, right? You need it to be low power and efficient. And all of a sudden you're taking power and chips that would have been for AI data centers and you're putting them in robots. So now that 200 gigawatts gets lower if you're deploying millions of humanoids.

2:26:43Dwarkesh Patel:I think this is very interesting because something people might not appreciate about the future is how centralized in a physical sense intelligence will be. Where right now with humans, your compute, like there's 8 billion humans and their compute is on their heads, on their person. And in the future, even with robots that are out physically in the world, I mean, obviously knowledge work will be done in a centralized way from data centers with huge, like hundreds of thousands of instances or maybe millions of instances. But even for robotics, the future you're suggesting is one where there's like more centralized thinking and centralized computation that's driving millions of robots out in the world.

2:27:23Dwarkesh Patel:And so I think that just like, yeah, that's an interesting fact about the future that I think people might not appreciate.

2:27:28Dylan Patel:I think Elon recognizes this, which is why he's like going to different places for his chips, right? He signed this massive deal with Samsung to make his robot chips in Texas because he thinks, you know, like I personally think he thinks that, you know, Taiwan risk is huge. and because of that and the centralization of resources in Taiwan, him having his robot chips in Texas and also being a separate supply chain that is not as constrained by... No one's making AI chips really on Samsung besides NVIDIA's new LPU that they're launched. They're launching it next week, but we're recording it the week before.

2:28:03Dylan Patel:It's coming out this week.

2:28:04Dwarkesh Patel:This episode's coming out Friday.

2:28:05Dylan Patel:Oh, this episode's coming out before. Sick. So they're launching this new AI chip next week, which is built on Samsung, but that's like sort of a recent development from NVIDIA. And then that's the only other AI demand there, whereas on TSMC, everything is competing. So he gets this like both geopolitical diversification, but also supply chain diversity for his robots. And he's not as competing as much with the like willingness to pay of infinity of the data center of geniuses. Okay, final question.

2:28:36Dwarkesh Patel:On Taiwan, if we believe that tools are the ultimate bottleneck, how much of Taiwan's place in the ASM conductor supply chain could we de-risk simply by having a plan to airlift every single process engineer at TSMC out when things come to, if they get blockaded or something? Or do you actually still need to ship out the EUV tools, which would be multiple plane loads per single tool and would not be practical?

2:29:04Dylan Patel:If you ship out all the process engineers and assuming it's like hot enough that you destroy the fabs, no one has all the fabs in Taiwan now, which is a big risk, right? You know, these tools actually use a lot of semiconductors, which are manufactured in Taiwan. So it's like a, it's like a, you know, a snake eating its own tail sort of like meme, because you can't make the tools without the chips from Taiwan, which you can't use without the tools in Taiwan. You know, there's obviously some diversification there, but, and they don't use super advanced chips in lithography tools. But at the end of the day, there is some tail eating the dragon.

2:29:36Dylan Patel:Just shipping out all the engineers and blowing up the fabs means China has a stronger semiconductor supply chain than the rest of the world, right? In terms of verticalization, now that you've removed Taiwan and now you've got all the know-how, but you've got to replicate it in, let's say, Arizona or wherever for TSMC. And it's going to take a long time to build all the capacity that TSMC has had built over the years. And so you've drastically slowed US and global GDP, not just growth, you've shrank the GDP massively, and you've got a lot bigger problems, and your incremental ability to add compute goes to almost zero, right?

2:30:16Dylan Patel:Instead of hundreds of gigawatts a year by the end of the decade, let's say by the end of the decade, something happens to Taiwan. Now you're at maybe like 10 gigawatts across Intel and Samsung or 20 gigawatts. It's like nothing. But now all of a sudden, you've like really caused some crazy dynamics in AI. Of course, you have all the existing capacity, but that existing capacity pales in comparison to the capacity that's being expanded.

2:30:36Dwarkesh Patel:Yeah. Okay. Dylan, that was excellent. Thank you so much for coming on the podcast.

2:30:40Dylan Patel:Thank you for having me and see you tonight.

From the publisher

Dylan Patel, founder of SemiAnalysis, provides a deep dive into the 3 big bottlenecks to scaling AI compute: logic, memory, and power.

And walks through the economics of labs, hyperscalers, foundries, and fab equipment manufacturers.

Learned a ton about every single level of the stack. Enjoy!

Watch on YouTube; read the transcript.

Sponsors

* Mercury has already saved me a bunch of time this tax season. Last year, I used Mercury to request W-9s from all the contractors I worked with. Then, when it came time to issue 1099s this year, I literally just clicked a button and Mercury sent them out. Learn more at mercury.com.

* Labelbox noticed that even when voice models appear to take interruptions in stride, their performance degrades. To figure out why, they built a new evaluation pipeline called EchoChain. EchoChain diagnoses voice models’ specific failure modes, letting you understand what your model needs to truly handle interruptions. Check it out at labelbox.com/dwarkesh.

* Jane Street is basically a research lab with a trading desk attached – and their infrastructure backs this up. They’ve got tens of thousands of GPUs, hundreds of thousands of CPU cores, and exabytes of storage. This is what it takes to find subtle signals hidden deep within noisy market data. If this sounds interesting, you can explore open positions at janestreet.com/dwarkesh.

Timestamps

(00:00:00) – Why an H100 is worth more today than 3 years ago

(00:24:52) – Nvidia secured TSMC allocation early; Google is getting squeezed

(00:34:34) – ASML will be the #1 constraint for AI compute scaling by 2030

(00:55:47) – Can't we just use TSMC's older fabs?

(01:05:37) – When will China outscale the West in semis?

(01:16:01) – The enormous incoming memory crunch

(01:42:34) – Scaling power in the US will not be a problem

(01:54:44) – Space GPUs aren't happening this decade

(02:14:07) – Why aren't more hedge funds making the AGI trade?

(02:18:30) – Will TSMC kick Apple out from N2?

(02:24:16) – Robots and Taiwan risk



Get full access to Dwarkesh Podcast at www.dwarkesh.com/subscribe

More from Dwarkesh Podcast

All 94 episodes
Dylan Patel — Deep dive on the 3 big bottlenecks to scaling AI computeDwarkesh Podcast · 2 h 31 min
Listen in VO