Dylan Patel: NVIDIA's New Moat & Why China is "Semiconductor Pilled”

5 Feb 2026 · 1 h 17 min · 39 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Summary of The MAD Podcast Episode with Dylan Patel

Episode Overview In this episode of *The MAD Podcast*, host Matt Turck speaks with Dylan Patel from SemiAnalysis about the evolving landscape of AI chips, the strategic shift at NVIDIA, implications of China's semiconductor policies, and the broader geopolitical context surrounding AI technology and infrastructure.

Key Themes Discussed

  • NVIDIA's strategic shift and its impact on the AI chip market.
  • The growing specialization of AI inference tasks.
  • The geopolitical implications of semiconductor technology, especially concerning China.
  • The potential for a capital expenditure (CapEx) bubble in the AI sector.
  • The environmental impact of AI technologies, including energy consumption and water usage myths.

Key Points and Discussions

  1. NVIDIA's Strategy
  2. Acquisition of Grok:
  3. Transition from a "one chip can do it all" approach to a diversified portfolio strategy.
  4. Recognition that different workloads in AI require specialized chips for efficiency and performance.
  5. CUDA Moat:
  6. Discussion on whether NVIDIA's CUDA remains a competitive advantage amid the rise of open-source solutions.
  7. Importance of networking and software ecosystems in maintaining NVIDIA's lead.
  1. AI Chip Wars
  2. Specialization in AI Inference:
  3. The emergence of startups focusing on niche AI workloads (e.g., Grok for decoding, Cerebras for large model training).
  4. Companies like AMD and newer startups (Etched, Cerebras) are diversifying the landscape but face challenges against NVIDIA's dominance.
  1. Geopolitical Context
  2. China's Semiconductor Landscape:
  3. Description of China as "semiconductor pilled" due to its aggressive inward push for semiconductor self-sufficiency.
  4. Impact of local government policies incentivizing domestic chip production, despite challenges from US restrictions on advanced chip technology sales.
  5. The Threat of Huawei:
  6. Huawei's capability and vertical integration pose a long-term threat to US semiconductor dominance.
  1. CapEx Bubble
  2. Debate on Infrastructure Investment:
  3. Concerns regarding whether current spending is irrational or a necessary investment for future AI capabilities.
  4. Dylan suggests that while significant spending is occurring, it is driven by genuine demand for improved AI models.
  1. Environmental Considerations
  2. Energy Consumption:
  3. The claim that AI is "killing the grid" and consuming excessive water resources is challenged.
  4. Comparison of AI data center water usage to hamburger production, framing AI's water footprint as relatively minimal.
  1. Future of AI Models
  2. Software Innovations:
  3. Predictions on advancements in AI models and their integration into business processes.
  4. The potential for tools like Claude Code to enhance productivity by automating routine tasks for non-developers.
  1. Broader Implications
  2. Job Market Impact:
  3. Shifts in job roles as AI tools become more capable, potentially reducing the need for junior analysts and entry-level programming roles.
  4. Cultural and Societal Changes:
  5. Analysis of how AI's integration into various industries will reshape work culture, productivity, and possibly the nature of tech-related employment.

Conclusion This episode provides a comprehensive look at the intersection of AI technology, semiconductor strategies, and the geopolitical landscape. Dylan Patel's insights shed light on the future of AI and its implications for industries and economies worldwide.

Links and Resources

  • [Dylan Patel on LinkedIn](https://www.linkedin.com/in/dylanpatelsa/)
  • [SemiAnalysis](https://semianalysis.com)
  • [Matt Turck on LinkedIn](https://www.linkedin.com/in/turck/)
  • [FirstMark Capital](https://firstmark.com)

---

This markdown file aims to encapsulate the essence of the podcast episode, highlighting the key discussions and insights shared by the guests and host.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

AI and Economic Warfare

0:45 to 1:32

Dylan Patel discusses the implications of AI on global dominance and economics.

“power grid can actually handle the AI boom, and the geopolitical chess match playing out between the US and China.”

NVIDIA's Acquisition of Grok

1:32 to 2:20

An exploration of NVIDIA’s acquisition strategy and its market implications.

“It's very clear we're not sure where AI models are headed in terms of, you know, over the next few years, what happens to the architecture.”

The Future of AI Models

2:20 to 3:52

Discussion on the evolution of AI models and the need for specialization in chips.

“Doing autoregressive tokens in a single stream super fast.”

Challenges for NVIDIA

3:52 to 6:12

Examining NVIDIA's competitive landscape and the threats to its market position.

“Which has put a lot of memory on the chip and not necessarily, in the case of Cerebrus and Grok, no memory off chip.”

Regulatory Concerns in Tech Acquisitions

6:12 to 8:00

Dylan shares insights on antitrust issues and the complexities of tech acquisitions.

“Now, in the case of a large company buying a startup, I'm completely fine with it.”

CUDA: A Lasting Competitive Advantage?

8:00 to 9:00

Discussion on the relevance of CUDA and its impact on the AI chip market.

“When you look across the market, there is only a few companies who have successfully created a chip architecture, software to run the models accurately, run the models accurately, right?”

The Future of AI Consumption

9:00 to 11:54

Insights into how AI will be consumed and the role of open-source solutions.

“Currently or have raised such as Etched, Maddox, Positron, these new age of AI companies.”

Cost Management in AI Workloads

14:03 to 16:58

Explore how NVIDIA's KV cache management can significantly reduce inference costs.

“And so if you think about, oh, it just worked for nine hours on one task, one refactor, huge value.”

AMD vs. NVIDIA: The Semiconductor Race

16:59 to 20:38

A discussion on AMD's position in the semiconductor market and its competitive strategies against NVIDIA.

“great if i need 100 people to develop it like google and you know so on and so forth did then that's much harder do you think amd can catch up i think amd will be caught up at times and very behind at other times.”

The Rise of Startups in AI Hardware

20:39 to 23:08

Insights into how new startups are positioning themselves in the AI hardware space and their potential challenges.

“But the world where they win is a multi-silicon kind of world where any given customer uses a range of different GPUs?”
Show all 39 chapters

Geopolitical Impacts on Semiconductor Industry

23:09 to 28:00

Analyzing how geopolitical factors and local policies affect the semiconductor market in China.

“Actually, in some quarters last year, it was even north of 20, I think.”

China's Specialization in Supply Chains

28:00 to 29:06

Explore the extreme specialization of Chinese cities in various industries.

“It almost sounds like more like the U.S.”

Global Semiconductor Supply Chain Insights

29:06 to 30:55

Learn about the intricate global supply chain of the semiconductor industry.

“It's really sick or semiconductors in general, but like, you know, like in Japan, they like focus on a few different types of chemicals and they're the best at it.”

China's Semiconductor Verticalization

30:55 to 32:05

Understand China's efforts to create a vertical supply chain in semiconductors.

“of chemicals, a hundred percent share from Japan, right?”

China's Tech Gaps and Strengths

32:05 to 33:19

Discuss the technological gaps China faces in the semiconductor race.

“The the fabs would shut down without foreign supply, you know, and you go down or you go across the stack.”

The Competitive Landscape in AI

33:19 to 34:27

Examine the competitive dynamics between NVIDIA and Huawei in AI.

“or many American companies and their tools.”

Huawei's Market Influence and Challenges

34:27 to 36:24

Analyze Huawei's position in the market and its challenges in chip manufacturing.

“And I think that's directly the result of what Google's been doing.”

NVIDIA's Strategy in China

36:24 to 38:22

Discover NVIDIA's strategic maneuvers regarding its presence in China.

“but, like, for other markets, I don't know, UAE, Middle East, Europe, are NVIDIA and Huawei already head-to-head in deals?”

The AI Revenue Landscape

38:22 to 39:45

Explore the projected growth of AI revenue and its global distribution.

“they'll figure out how to sell chips to China.”

The Semiconductor Industry's Complexity

39:45 to 42:03

Understand the vast complexities and significance of the semiconductor industry.

“I think$100 billion is end of this year.”

The Complexity of Semiconductor Supply Chains

42:03 to 43:14

Explore the intricacies of semiconductor supply chains and the implications of government subsidies.

“Now, obviously like Google designed semiconductors, but it's like, oh wait, no, but their cost of search would be like 10x higher if they didn't have TPUs.”

Impact of COVID on Chip Manufacturing

43:15 to 44:34

Discuss how the pandemic influenced semiconductor manufacturing and automotive industries.

“TSMC is literally making chips for NVIDIA and Apple and AMD and others in Arizona today.”

U.S. Semiconductor Investments and Optimism

44:35 to 46:33

Assess the necessity and future of U.S. investments in semiconductor manufacturing.

“If that didn't happen, we wouldn't even have the chips act.”

Public Perception of AI and Its Impacts

46:34 to 48:38

Delve into societal reactions to AI technology and its perceived consequences.

“And then like there are a couple of people who booed.”

Evaluating AI CapEx and Market Dynamics

48:39 to 50:56

Analyze the current state of capital expenditure in AI and its implications for future growth.

“what you were saying earlier about the rate of revenue increase and therefore implied demand that you expect for this year?”

The Future of Programming with AI Assistance

50:57 to 53:14

Explore how AI is changing software development and the role of programmers.

“Ultimately, the CapEx that Microsoft spent in 2024 for OpenAI is what results in 2025 for OpenAI or CoreWeaver or whoever is what results in their models being so good this year.”

Challenges in Energy Production for Data Centers

53:15 to 56:00

Discuss the energy infrastructure challenges facing data centers in modern society.

“He built it in a week and he didn't type a single line of code, right?”

The Challenges of Equipment and Labor

56:00 to 57:02

Explore the complexities of building infrastructure and equipment for energy.

“I think ultimately that's the biggest problem is the equipment and the labor.”

Water Consumption Myths in AI

57:02 to 58:14

Debunk misconceptions about AI's water usage and its impact on resources.

“This other cool poster just last week or two weeks ago that was about the water consumption.”

AI Water Usage vs. Hamburger Production

58:14 to 59:14

A surprising comparison reveals how AI's water usage is minimal.

“Cause, cause you know, I've heard that argument from some like vegetarian people before or some Hindus or like, I'm Hindu myself.”

Energy Supply and Data Centers

59:14 to 1:01:16

Discuss the dynamics of energy supply, data centers, and investment opportunities.

“Because that's, you know, you do the calculation on how many, how many burnt, what's the average revenue per in and out?”

The Investment Landscape in Power Generation

1:01:16 to 1:03:19

Analyze the potential profitability of independent power producers in the current market.

“There's a lot of room for power producers to get outsized returns.”

Understanding the Debt in Data Center Deals

1:03:19 to 1:04:18

Explore the financial structures behind data center capacity and investments.

“Like hyperscalers are paying for transmission upgrades, which people will benefit from, right?”

Circular Financing in the AI Sector

1:04:18 to 1:06:18

Investigate the circular financial relationships among major AI players.

“You know, just having a customer alone spoken for it was enough, right?”

The Future of AI Models and Software Development

1:06:18 to 1:10:00

Discover how AI is transforming software development and modeling practices.

“It's mostly just 99 plus percent of their spend at the company is probably just compute.”

The Impact of AI on Low-Level Knowledge Work

1:10:00 to 1:12:08

Explore how AI tools like Claude Code are reshaping the landscape of entry-level jobs.

“an investment case for clients, as well as like, you know, some other interesting details from someone who's never really coded, just using Claude Code and it like doing this all.”

New Paradigms in Model Development and Productivity

1:12:08 to 1:13:46

Learn about the advancements in AI models and their implications on productivity.

“And so now we're trying to force everyone in my company.”

Praising Sholto: A Remarkable Talent

1:13:46 to 1:14:18

Delve into the impressive qualities and achievements of Sholto, highlighting his skills.

“And speaking of Sholto, we both agreed that he was a perfect specimen.”

Casual Conversations and Life with Roommates

1:14:18 to 1:16:15

Discover the everyday interactions and topics shared among roommates in tech.

“It's like, holy crap, you're a specimen.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Dylan Patel:This is the biggest change in human history, maybe ever. What's about to happen with AI? This is the biggest revolution, bigger than industrial revolution. Jensen is very paranoid about losing. If he just kept making his mainline chip, people crush him on cost and performance. Acquiring Grok is how you get those resources to make more solutions for different parts of the market to stay king. At the end of the day, this is an economic war. If the US and the West win in AI, China will not rise to be the global hegemony. But without AI, China definitely will rise. They're just going to outrun America.

0:29Hi, I'm Matt Turk. Welcome back to the Matt Podcast. Today, I'm joined by the one person Wall Street and Silicon Valley turned to when they need to cut through the hardware hype, Dylan Patel of Semi Analysis. We dove into many of the most important topics of today. Nvidia's massive move to acquire Grok, the truth about the CapEx bubble, whether the US power grid can actually handle the AI boom, and the geopolitical chess match playing out between the US and China. But I have to warn you, this conversation went off the rails in the best possible way. And we ended up going into all sorts of fun tangents like the strange phenomenon of Chinese romance dramas set inside semiconductor factories and what it's really like when three AI famous roommates live together in SF.

1:10Please enjoy this fantastic conversation with Dylan. Hey Dylan, welcome. Hello, how are you? I'm great. I'd love to start with Rock and NVIDIA since it's still fresh. So not so long ago, NVIDIA was saying that one GPU could do it all and now they're doing this acquisition slash non-exclusive deal with Grok. What does that mean from your perspective?

1:32Dylan Patel:It's very clear we're not sure where AI models are headed in terms of, you know, over the next few years, what happens to the architecture. But, you know, the thing that I think everyone has sort of like agreed on is models are pretty autoregressive, right? Next token generation is like the thing. But beyond that, right, attention mechanisms change, how it works, everything changes, right? Could change. And so what's interesting is the reason NVIDIA 1 is because they just took like the widest surface area bet and then people kept developing models on that and that kind of shape worked. But now the workload is so large that there is room for specialization that will give you 10X increases in certain domains, right?

2:07Dylan Patel:In a general purpose workload, crock, crock, didn't work, right? You know, it can't train, it can't, you know, it can't inference really, really large models cost efficiently, right? You can't serve many, many, many users, but what it can do is it can go screamingly fast, right? Same with the Cerebris OpenAI deal. But that's like one workload, right? Very decode focused, right? Doing autoregressive tokens in a single stream super fast. Another direction AI models could head, right? We don't know, are models going to think in one token stream? Or is it actually they're constantly context switching, right?

2:39Dylan Patel:And they're going from, they have this humongous, humongous context and they're generating in multiple parallel streams, right? And so Google and OpenAI have both released mechanisms of this with their pro models where the model actually doesn't just have one single chain of thought for reasoning. It has multiple, right? And then I don't know exactly like, you know, and how they choose which one and what the final answer to you delivers is an area of research. But there is room for that kind of chip, right? Something that works on very parallel, a lot of streams of chain of thought. And maybe the latency requirements are not as crazy, right?

3:13Dylan Patel:Maybe you don't want to go blindingly fast, right? Maybe you're okay with it being, you know, because I can spin up 100 parallel, you know, streams of thought or agents or whatever you want to call them. Maybe I care a lot about cost there. And because it's 100 in parallel, instead of one going super, super fast, it's not as deep, right? The tree search or the depth of the inference is not as deep, but it is much wider. You know, there's other parts of inference. Hey, creating the KV cache. So NVIDIA has a chip for that, right? That's the CPX. So they've made the CPX, they bought Grok for decode, and then they still have their general purpose GPU.

3:43Dylan Patel:So they're kind of trying to cover their bases because unlike the first wave of AI chip companies where they sort of just made chips and then tried to figure out where it would work, right? They had a thesis, Grok and Cerebrus both, as well as SambaNova, right? Which has put a lot of memory on the chip and not necessarily, in the case of Cerebrus and Grok, no memory off chip. And in the case of SambaNova, less memory off chip or slower memory off chip with higher capacity. You know, they sort of all made similar bets in that direction. And it didn't work for a while until it kind of did, right?

4:11Dylan Patel:Because there was a workload that now necessitates it. NVIDIA recognizes they're the leader, they're at the tent pole. Hey, in one respect, they can just run faster than everyone, but it's kind of hard to be 2x better than Google or OpenAI or whoever else's internal chip, right, to justify their, you know, 75 % plus margins, right? And then they have to be 2x to 4x better to justify, 4x better to justify their margins because that's what they're charging above COGS. You know, the question is what architecture will deliver that? Well, yes, keep the programmability of their GPUs is great for training and for a lot of workloads.

4:47Dylan Patel:But, you know, guess what? I think a lot of people will just be downloading an open source model, downloading an inference framework and pressing go, right? A little bit more complicated than that, but that's going to be the consumption method for a lot of enterprises, a lot of startups, a lot of tech companies is they're just going to do that or they're going to rent the GPUs or rent the chips and then download an open source framework and model and go, right? And And NVIDIA recognizes this and hey, there is room for products that aren't general purpose, right? The general purpose GPU will still probably be the main line for training and for a lot of inference and for cost efficient inference, but maybe blindingly fast or workloads that have a ton of pre-fill, i.e.

5:24Dylan Patel:creating the KV cache. Maybe that those workloads could be different chips, right? And the CPX chip they announced, right? They say it's for the context processing, creating KV cache. It's also really useful for video models because video models don't care about memory bandwidth. And so, you know, why pay for the expensive memory that the general purpose chip has? Or why do what Grok is doing, which is tying hundreds or thousands of chips together and not having memory, but keeping the entire model on chip? The tradeoff for that, of course, is you need thousands of chips and you have less compute per chip.

5:50Dylan Patel:And so like NVIDIA is trying to capture the whole surface area because, again, you don't know where models are headed. And it's hard to say where the research is headed. And do you think it's a good thing for the market? It's yet another one of those deals that's structured as a license, but released an acquisition. I certainly think it's not good from an anti-competitive sense, right? I don't think people should just be able to buy companies without any antitrust process at all. Now, in the case of a large company buying a startup, I'm completely fine with it. The flip side is like, hey, we know the deal is happening, right?

6:22Dylan Patel:This happened for a company I was an advisor for. NVIDIA acquired in Fabrica just maybe a few months before they did Grok and similar style of deal, right? If someone wanted to strike it down, that's the biggest limbo, right? We've seen this happen in venture, and you probably know more stories of this, but a company trying to get acquired, they get stuck in limbo for a year, and then it falls apart. Many stories. Yeah, it falls apart, the deal did, because some regulatory BS, and now the company was, and the founders were focused on getting the deal done, instead of making the product better for a year, now they're behind, or they weren't focused on growth as much, right?

6:56Dylan Patel:You only have so much time as a founder. So in that sense, I like the license deals, right? So now is NVIDIA also dominating the inference market? Is there any world where NVIDIA is no longer the king or they seem to be getting stronger? I think the thing about NVIDIA is they take the Andy Grove mentality like more serious than anyone else, right? Like, OK, fine, Google like implemented OKRs because Intel did it. But that's like, you know, management stuff, right? Only the paranoid survive, right? This is like core to the Bay Area, core to NVIDIA. Jensen is very paranoid about losing, right? These specializations, if he just kept making his mainline chip, would mean people could, you know, point solutions for specific parts of the market would crush him on cost and performance, and then he can't justify his margin.

7:40Dylan Patel:That's a threat to NVIDIA's business model as a whole, especially if the best model only changes every three months or the model you want to roll out. Okay, well, then you have three months to figure out how to make a model work on one chip architecture for that point solution. And, you know, it's fine. Software advantage of NVIDIA is not that important then. Jensen's super paranoid about losing. And frankly, it's really hard to hire enough talented chip people. When you look across the market, there is only a few companies who have successfully created a chip architecture, software to run the models accurately, run the models accurately, right?

8:12Dylan Patel:Like, because you can look at random APIs of say an Alibaba, a Quen model, and different people are doing all sorts of tricks like quantizing it, but also many other tricks, which then end up like making the model quality lower, you know, building a rack scale solution, networking thousands of chips together and then deploying an API. And Grok did the whole thing with, frankly, not that many people. So now it's like, OK, well, I'm NVIDIA. I want to make four different chip architectures and actually four different point solutions, maybe the general purpose and then one here, one here, one here.

8:37Dylan Patel:And in addition, my general purpose thing is actually not just like a GPU chip. It's like GPU chips, CPU chips, networking chips, NvSwitch, Nix. Like, you know, there's many, many chips and each of those chips has many chiplets. You don't have enough engineering resources, right? And so like acquiring Grok is like how you get those resources to make more solutions for different parts of the market. And as far as like, are they threatened? Like, I think, I think like, obviously there's some cool startups out there, right. That are raising a lot, right. Currently or have raised such as Etched, Maddox, Positron, these new age of AI companies.

9:08Dylan Patel:There's also the prior age of like Cerebris is out there still, right. You know, TensTorn, et cetera. And there's, so there's a lot of AI chip companies on the startup side, but then there's also, you know, Google's TPU, AMD, GPUs, Amazon Tranium. who are all really credible competitors. And then, you know, Meta's MTIA is somewhat credible. And then, you know, Microsoft's Maya is not credible, but like, you know, maybe it will be one day, right? So you sort of have like a lot of competition. They've got to hold the gates back. And so I think, is there a risk to them being, I mean, there's risk from all of those companies that I mentioned and, you know, effectively California slash Seattle, right?

9:43Dylan Patel:Only two places. There's also chips from other parts of the world, right? Obviously, China has a number of different AI chip companies that are doing cool things. Anyone would have told you Grok was, you know, their business revenue, their revenue was not like stellar, right? In fact, they missed revenue last year significantly and yet they got bought, right? Because the value of the IP was there and the value of the team. Anyone else would have been like, well, why the heck would I buy this, right? Makes no sense. There's definitely a credible threat. Yeah. And do you think CUDA is going to remain that moat?

10:10I guess the combination of CUDA and whatever came out of the Mellanox acquisition, like do those persist as long-lasting advantages?

10:17Dylan Patel:I think they do. I think networking is super important. I think the CUDA software mode is very important, but it's also like changing rapidly, right? It's an incredible amount of the software that NVIDIA GPUs run on is not from NVIDIA. It's the developer ecosystem that's open sourcing it. When you look at, for example, VLLM and SGLang, right? These support AMD GPUs almost as first-class citizens now, And VLM is getting significant support for TPUs, for Tranium, and there will be other chips coming out from startups that also support VLM slash SGLang. Now, like, how difficult is it? You know, the reason why CUDA is so important is like, okay, I can do whatever I need to do, right?

10:56Dylan Patel:Programming a GPU. I think most AI chips will not be consumed by people programming anything for it. They will download an open source inference engine. They will download an open source model and then they will put it on the, and it's really simple to download VLM and like make it work. Like it's not that hard to set up, you know, a server. And NVIDIA is putting out a lot of open source software like Triton Inference Server and Dynamo and all these things to make it easy because that is the consumption model ultimately for the majority of AI, right? Is, and it might be like, oh, it's my own inference engine, but most servers will not run code besides the inference engine in the model.

11:33Dylan Patel:It's not like people are actually like, researchers are like writing code for GPUs to see ideas if they'll work and train models and all these things, or just mess around with them to figure out, you know, performance or whatever it is, but most of it won't be there. And so CUDA as a mode, CUDA language is like, you know, like it's like fine, right? Like, you know, no one actually writes CUDA, right? Like most people write PyTorch and then like Torch compile and then they just run it on the GPU. They don't write CUDA. But a lot of this CUDA mode is like, how does PyTorch translate into high performance GPUs?

12:01Dylan Patel:And that surface area from when people were just writing like hardcore, when people are hardcore writing CUDA kernels to like, hey, they're writing PyTorch and then it's compiling down to GPUs versus, oh, I'm just downloading VLLM. It is a curve of like not a ton of people that can do CUDA kernels. A whole lot more people can do PyTorch, right? Random PhDs and random people. It's very simple, right? A crap load of people can do VLLM, download it, run it on a server. Well, if it now supports other chips, what is the CUDA mode? NVIDIA has recognized this and they've been building software that is not necessarily the CUDA mode.

12:33Dylan Patel:And I can give some examples, right? So the name of the game is fast tokens and lowest cost tokens, right? And lowest cost tokens happens by your chip being fast, but there's also tricks, right? One example, right? Like I mentioned with the CPX versus Grok, right? Is processing your pre-fill contacts, right? Super cheap CPX, right? If I care a lot about speed, then Grok. These are optimizations on the hardware side. There's optimizations on the software side as well, right? And so one example is when I'm doing, for example, if I look at a cloud code or a cursor type application, right? The workload is like, it takes your repo, it takes the relevant parts of your repo, puts it in the context of the LLM, it prompts, it generates, right?

13:14Dylan Patel:And if it's an agent mode, it circulates the context a couple of times, it'll collapse, put things off to the side, access different contexts. But what's, you know, especially when you think about an agent for software, and you can see this in Codex, you know, Codex actually not as good as cloud code, but it can do work on time horizons of like nine, 10 hours. and do like a big refactor better than cloud code can, even though most of the times cloud code is better. And what's interesting about Codex does is it'll like take your repo, it'll identify parts. If you're asking it to refactor it, identify parts, write stuff, you know, make like these notes for itself everywhere, collapse the context, switch from this part of the repo to that part of the repo to this part of the repo.

13:50Dylan Patel:But when you think about it, it's like, oh, if this thing is just generating tokens all the time, plus it's switching what my context is constantly, that's really expensive, right? If you think about like, what's the cost of inference? I want to say it's like, it's$10 per million tokens of output and$3 for decode or 10 for decode and three for pre-fill. And so if you think about, oh, it just worked for nine hours on one task, one refactor, huge value. But if it changed context a ton of times and your context is like 30K usually or 50K or, you know, heading to hundreds of thousands, you know, how long, how big your repository is and how much context switch.

14:28Dylan Patel:Now you're spending all this money on, on pre-fill, right? Not the decode tokens, but actually why am I like regenerating the KV cache? I can actually just like store the KV cache elsewhere. And then when I need it again, I can pull it and plop it into CPU memory or into GPU memory. And so NVIDIA has got this like KV cache manager and they've been working really hard on like making it so they can interface SSDs and stick the KV cache on there and pull it out whenever they want. So for this kind of workload. And then if you do this and you look at like coding as an application and you like look at these coding companies and how much they're paying for prefill versus decode, actually majority of their cost is prefill tokens, not decode tokens because their context is just so large and it's switching all the time, even in agent modes.

15:08Dylan Patel:You know, if you can now not have to do the prefill, your costs go down dramatically. But that's a very complicated thing to do from a software perspective. You know, companies like Anthropic, Google, OpenAI have already done it. But what about the wide world, right? And so NVIDIA is trying to make the open source software for this. And that's like a CUDA emote, but it's like, actually, no, none of this is CUDA, right? Like it's like memory management and like, you know, storage management. And when do you call what and how do you transfer it? And how do you like spread the KV cache across a bunch of different storage nodes?

15:34Dylan Patel:And what happens when you read it and the network congestion, just like all these things. Yeah. It's like NVIDIA's wheelhouse, but it's not CUDA. And I think like the easy way to say it is it is the CUDA mode, right? And so things like this KV cache manager and many other things they're trying to do to reduce the cost of inference, like is how they build the new CUDA mode. Because again, today it's It's, you know, it is quite, I mean, AMD is like not fully there yet. And TPU is being added right now and Tranium is being added soon as well to VLLM. But all of them will have a very good UX for download model, run model on VLLM by the middle of the year, I think.

16:08Dylan Patel:Right. Certainly AMD is already there by the end of this quarter. We have something that like test this, right? It's called inferencemax.a. It's an open source. All the code is and the results are. but we run across, I think,$60 million of GPUs, which are donated to us by companies like NVIDIA, AMD, OpenAI, Microsoft, Amazon, Crusoe, CoreWeave, Together AI. All these companies are sponsoring GPUs for us to run this. We're running VLM and SGLang every night on nine different kinds of GPUs on a variety of different models and different work context lengths and all these things, right? To see the performance and you can see the performance moving every day or pretty often because the software changes all the time.

16:43Dylan Patel:and so like the fact that this exists is the coup de vote right it's not that like amd you can do this on their chips and v you can do this on their chips it's oh when the new model comes out how fast does it get to peak performance because you know it's a moving target or hey can i implement this kv cache management thing how hard is it how many engineers do i need oh just one great like or 10 great if i need 100 people to develop it like google and you know so on and so forth did then that's much harder do you think amd can catch up i think amd will be caught up at times and very behind at other times.

17:14Dylan Patel:Like currently they're super far behind, right? Because Blackwell is just way better than MI355. And then, you know, Rubin comes out and they'll be way, way behind, but then AMD's new chip comes out and AMD will be caught up or even slightly ahead on a hardware perspective, software's behind, right? And you have this like leapfrogging and AMD is a very credible second competitor. I don't think they'll go beyond like, I think they'll stay in single digits market share, single digit percentage market share. But single digit percentage market shares. It's still pretty good. Yeah, I mean, NVIDIA's revenue this year is going to be like, it's a lot.

17:43Dylan Patel:Three gajillion dollars. I think it's actually four gajillion. What about all the startups? You mentioned a few. So there's Cerebrus on the one end of the spectrum and then newer ones, Etched and others. If AMD has a, you know, uphill battle in front of them, like, do you think those guys can take a significant market share? You sort of the whole specialization game, right? You have to specialize because you're never going to be NVIDIA at their own game, right? They're going to have the supply chain on lock. They're going to get to the newest memory technology or process technology or whatever packaging technology, whatever it is, sooner than you.

18:19Dylan Patel:And they're just going to crush you, right? If you play their game, you have to, AMD is trying to play NVIDIA's game, but AMD is like extremely good at engineering silicon, right? Everyone else has to, has to, has to try something weird or different, right? And so when you look at Etched or Maddox or Positron or Cerebris or Tensitorn, you got to look at all these companies, right? There are unique things about what they're doing. And it's not clear if AI models will still be within that realm when that comes out, right? does, oh, now people use like engrams and other sparse attention techniques.

18:54Dylan Patel:Is that like, does that change like some of the specializations people are doing? Or hey, people are now doing like, you know, models are now sparse MOEs instead of being dense models. Does that change things? There's so many optimizations and changes on the model side and you can't predict what's going to happen with the ML research easily, at least. You can't. The thing you're optimizing for today has to be a vision of where AI will be in two years. And NVIDIA is fully accepted. they don't know where that's going to be. That's why they have a portfolio of chips now, not just one GPU line, right?

19:26Dylan Patel:It's not just Hopper, Blackwell, Rubin. Now it's going to be, you know, it's not Ampere Hopper, you know, it's not that line. It's like there's a variety of chips to serve the different markets and different possible scenarios. They think each of them has this vision today, but, oh, it might turn out the general purpose one sucks. And actually AI models have developed in a way where CPX or Grok style chips are the best, right? Well, okay, now we have a solution for that market. And so I think that's the challenge with the startups. With that said, I think they're all taking very interesting bets.

19:52Dylan Patel:I think it's much more exciting than the first wave of AI hardware bets, Graphcore, Cerebris, Samanova, Grok, where they all made the same bet on memory and putting the memory on the chip. They sort of just made a bet and they optimized for a certain kind of model, all similar kinds of model. And it didn't end up working out for a long time, right? They had to pivot and they had to work on a lot of things. It took a long time. I think these companies have like a really clear vision of what they think models will look like, right? Etch does, Maddox does, Positron does. And that's what's really cool about it between the three of them, these new etch.

20:26Dylan Patel:So, I mean, I'm excited for them. I'm very, very skeptical. I don't know what a venture capitalist views as like likely chances of succeeding, but I think all of them are less than 1%, right? But, you know, that's... But the world where they win is a multi-silicon kind of world where any given customer uses a range of different GPUs? It could, it could. Or it could be any given customer has like one workload they care a lot about. Anthropic clearly does not give a crap about video gen, image gen, right? They just don't care. On the flip side, a company like Midjourney cares a lot about image and video gen, right?

21:03Dylan Patel:Image and video gen is very, very, like I mentioned, it's not very memory bandwidth heavy. It loves, loves, loves compute, right? Whereas inference of large language models in the style of like, you know, say for example, coding agents cares a lot about decoding for long streams of time. And that's very memory bandwidth heavy, right? And so that's like a simple example, but there's a lot more nuance there in terms of like, even like the size of like the matrix multiply, you know, the tensor cores that you, you know, the systolic arrays that you use or the ratios of networking and memory and like, what's that memory hierarchy look like?

Read the full transcript

21:36Dylan Patel:And, you know, what are you doing for different kinds of attention? and like, oh, like all these sorts of things, like there's a lot of specialization here. And so some people are betting big on different types of specialization. And I think like you could clearly see a world where companies do care about different stuff, right? Like if, for example, a chip optimized for video and image generation existed today and it was better than NVIDIA or NVIDIA made it, I think MidJourney would absolutely only use that for inference. I think for training, they'd still use the general purpose thing. And as would like Meta and Google would like, they should do that, right?

22:07Dylan Patel:And hey, Meta actually has two lines of AI chips. Their MTIA, there's a line that's focused on recommendation systems. And then there's a line that's focused on Gen AI. The Gen AI one is a new line, but that recommendation systems line is still continuing, right? It's not sexy, no one cares because there's no, and ByteDance also has a recommendation system line of chips. And it's not really focused on Gen AI, which is fine because, you know, this is a$200 billion business or something, which is just deciding what ad to serve me, right? And what order to put my friend's stories and, you know, things like this.

22:36Dylan Patel:So I think like it's perfectly fine for there to be specialized AI chips, given the target market is big enough and you have to have vision to know what that target market is. Unless you're hyperscaler, then you can like just like you can just use general purpose until you've like it's clearly there and then you can make your ASIC. Right. Fascinating. Turning to the geopolitical aspect of all of this, which is always fun, Huawei and NVIDIA in China. Last year, there was like 10 or 12 percent of their overall revenue. And this year they were saying that their market share has basically dropped to not very much.

23:08Is that Huawei chips? Is that restrictions? Is that tariffs? What's happening over there?

23:12Dylan Patel:I think it's a variety of things. Actually, in some quarters last year, it was even north of 20, I think. I don't remember exactly. But anyways, you know, if you look at 2022, China was almost the size of the U.S. in terms of buying server hardware, right? Almost. Not quite, but getting there. And it looked like they were going to be the same size as America in like a year or two after that, right? And if you look at like global data center capacity, global cloud capacity, et cetera, et cetera, et cetera, it's American companies and Chinese companies, right? That dominate the world. American companies are obviously doing a lot better here, but both of those dominate the world.

23:42Dylan Patel:And if you look at like every industry, right? You know, it's very clear that like China wants to insource stuff, right? So in 2015, they made these five year plans for 2020 and 2025, where they set the percentage of semiconductors they wanted domestically produced. And they've missed the goal both times, which is fine, right? They set really aggressive goals and shoot for the moon, even if you miss, you hit the stars, right? And that's sort of what's happened, right? Like, look, China is not caught up on leading edge semiconductors, but microcontrollers from China are almost as good as the microcontrollers are as good and cheaper than the ones from Texas Instruments or STMicro or, you know, et cetera, Right.

24:20Dylan Patel:Or like this power, random power chip is better than or the same as the one from like another company. Right. And so they've really built up a semiconductor industry and started insourcing a lot more. I don't see why China wouldn't be buying, you know, 30, 40 percent of the world's AI chips and the U.S. like 50, 60 percent. And then the rest of the world, like, you know, and when I say U.S., I mean U.S. origin companies. That seems like a more natural state for the world. but there are restrictions and hey this is the biggest change in human history maybe ever knowledge work and you know everything that's going to happen there and and then eventually like robotics and all these things like you know obviously there's there's a lot of geopolitical stuff and so there are restrictions nvidia has been hand capped and handicapped from selling their best chips to china and so that's obviously impacted the sales a lot because like why would you do that and so when you look at who rents the most gpus in the world it's three companies right So one of them is obviously OpenAI.

25:10Dylan Patel:Second one, actually, they were bigger than OpenAI. They are bigger than OpenAI today. Or no, they were bigger than OpenAI and OpenAI eclipsed them recently, is ByteDance. ByteDance rents tons of chips from Oracle and Google and many other cloud companies because they couldn't get the chips they needed in China. They're mostly just serving TikTok, right? Okay, well, they're not allowed to buy them, and that sucks, but they're allowed to rent them. And so, okay, if I'm not allowed to get the best ones, I'm gonna rent externally. And if ByteDance is the second biggest renter of GPUs in the world, that's substituting demand that would have been built in China in many cases.

25:40Dylan Patel:It's instead being built in Malaysia. And Oracle has over a gigawatt of capacity in Malaysia that ByteDance is going to take, right? So things like this are, you know, hundreds of thousands, millions of chips, tens of billions of dollars of capacity that would go to China, but it's not. It's going to Malaysia instead, as an example. Another sort of point around this is China's like, you know, they've had these five year plans. So and you know, the way these initiatives work from China is there is like some top down ordering, but then they just kind of whip the whole like everyone just kind of gets into it.

26:07Dylan Patel:And it's really cool. Like, I don't think it's as top down as many people think. Like, I think the entire country is like semiconductor pilled, right? There are dramas where people fall in love in the fab or dramas where people fall in love and they're photovoltaic like solar cell researchers and engineers. And it's like, it's like, this is just the backdrop. And it's like, actually, this is, it's like super cool for your like significant other to be that semiconductor engineer or to be that photovoltaic, you know, a solar panel researcher. As opposed to an influencer. As opposed to an influencer, right?

26:41Dylan Patel:Like, I'm sorry. Love Island is, I watched like for 10 minutes cause I was forced to, I was like, this is fricking terrible. But you know, like, we are so cooked. No, you know, seriously, we're cooked, we're cooked. I think like when you think about like this happens, it's like it's diffused into drama even. People like like there's multiple dramas like taking place about semiconductor industry. And they're like romance, comedy, like the entire spectrum, right? Drama really is like it's like what the heck is going on? Anyways, you have all these provinces, you have all these local cities setting out ordinances and giving out subsidies and all sorts of stuff.

27:18Dylan Patel:Right. It's truly like crazy. Like there's some national level stuff like, oh, no taxes on this. Oh, we're going to ban a few things. But as far as I understand, the national government has not banned NVIDIA's H20 or H200. But the local ones have, right? A lot of local ones have said, no, you know, you must use China manufactured chips. And it's like, who told you that, you know, you're here to uphold this? It's like, doesn't matter, right? I mean, like, it's cool because then you have this like survival of the fittest. All these provinces and cities are trying to attract different companies with different types of subsidies and grants and industrial parks and like all these different things.

27:55Dylan Patel:And then like the ones who succeed actually develop an industry and they take over. This is how one thinks of China, right? It almost sounds like more like the U.S. or like it was a federal government and states with the provinces of authority over their purchasing. I mean, it's actually like great. there's this one TikTok and Instagram person, and they sing it. They're like, if you want to buy things in China, make sure you go to the right place. And then they just say the most random shit and name the city. And then you look into it and you're like, wow, this city has the entire supply chain for this.

28:25Dylan Patel:It's like lampshades, and then it names the city. It's like, what the fuck? There's a city that specializes in lampshades. It's like microphone arms, like microphones. Literally, there's a city in China that specializes in things. And guitars as well, right? this one city that became the guitar. It's literally everything. Yeah. Literally everything. There's a city and it's not like, hey, specifically for camera arms, for example, there's ball bearings in this and the ball bearings are like there's ball bearings. There's multiple manufacturers of ball bearings for camera arms. And then like most of the camera arms in the world come from that one city.

28:55Dylan Patel:It's like, what the hell is going on? And so like the semiconductor industry, I think people don't realize is absurdly specialized. I'm not answering your question. I'm just going a little bit of a rant because I think people don't understand China semiconductors. It's really sick or semiconductors in general, but like, you know, like in Japan, they like focus on a few different types of chemicals and they're the best at it. And it's like almost a cultural thing, right? Like Japanese people are so precise, like with sushi and like, it's all about the trade and the craft and like, you know, the French food in Japan is better than the French food in France because the Japanese chefs went there and then come back and they perfected it in Japan and like, cause they're so precise.

29:28Dylan Patel:And, and there's so many different like things that like Japan is so good because they're so precise and like dedicated to the craft. And it comes out of like, I don't know, like samurai culture or something. I don't know, right? Like, I don't know exactly know how that culture came up. And so when you look at like, and it's like across the world, there's different places where things like this happen, right? Like, oh, like the Netherlands makes UV tools. Cool, I guess so. And you look across the semiconductor industry, there's a famous economic essay called Eye Pencil or something like that, or talking about how the pencil, like a simple pencil comes from like, oh the rubber comes from like indonesia for the eraser and the graphite comes from this mine here and and the wood comes from these aspen trees in canada and like you actually can't make a pencil without aggregating this entire supply chain semiconductor industry is like way crazier because like i would say there's like 15 or 20 countries that could shut down the entire semiconductor industry right even like austria could right and it's like what it's like well yeah there's two different companies there who have like 90 share in like some random niche stuff.

30:23Dylan Patel:And it's like, okay, cool. I guess Austria can. And oh yeah, those two companies have less than a billion of revenue, but they just happen to have linchpin critical things. And there's linchpin critical things everywhere because the process is so complicated. And so China has been trying to replicate this. Is there one thing they're missing that they don't have yet? I think there's a lot of things. I think if you were to close your eyes and say, or if you were to cut off every country and say, there's no more globalism, China has the most vertical stack in semiconductors today. And they're the best at semiconductors in the world because their fabs could still run somewhat on a lot of things because they have built some of these chemical supply chains, right?

30:54Dylan Patel:Like TSMC for certain kinds of chemicals, a hundred percent share from Japan, right? Or Intel, same thing, right? Or, you know, for certain kinds of tools, a hundred percent share from Netherlands or a hundred percent share from, you know, this American company or that, you know, Austrian company or this or that, right? Like there's just all these, like, you know, this Swiss company, like it was just all these different places have a hundred percent share. It might be one company, it might be three companies, geographically or in the same area. And China's built that up, right? Because they've created this made in China initiatives, which just plowed money into it.

31:22Dylan Patel:And they've got this culture of like the diffused, like, you know, these provinces are like, yeah, I just decided I'm going to fucking focus on, or it might not even be the city, right? It may be the like, you know, someone brought it there and decided, and then people were like, oh, wow, you're doing that? Me too. Like I'm a Patel and I grew up in a motel. And guess what? We like almost all the Patels I know grew up in a motel and it's because some random Patel immigrated to America and like worked at a hotel motel and then bought a motel that like just started happening. Right. Like it's sort of like these things are serendipitous of sorts.

31:52Dylan Patel:And like, I don't know, like it's like I view it as the same kind of specialization. Right. Chinese cities are like starting to do these things. China is missing a lot of things. Right. I would say like if you say minus 10 years tech, China is complete and no one else is complete. Right. Taiwan is not complete. The the fabs would shut down without foreign supply, you know, and you go down or you go across the stack. But if you go to 10 year tech, maybe, maybe more like 20 year tech, you could get a fully vertical supply chain in China, which I do not think any country could do. Like America could not build a fully vertical fab without stuff from elsewhere, even if it's 20 year old tech, probably not even 40 year old tech.

32:24Dylan Patel:And so, so that's interesting. But then when you look, flip side is like, well, like you kind of do need specialization. That's how that chemical gets the purest, best, most engineered, or that slurry of chemicals or that gas or that tool. Because every smart person or a lot of them in that country grew up around that culture and the supply chain is there and everyone kind of knows and it's like a drive away. And this is what makes supply chains work, is that there is this specialization and the best of the best only comes when you have that hyper specialization. So China doesn't have lithography.

32:59Dylan Patel:Their lithography is like 10 years behind. And I think it'll be five years behind in a couple of years, right? They're catching up fast. I don't think they'll be as good as ASML for a long time. You know, maybe, I don't know, maybe they will be, you know, you know, China, you shouldn't ever underestimate China, but like in Chinese engineers or, you know, but like for a while, right. Or like, you know, I don't think they'll be able to make leading edge chemicals like many Japanese companies or many American companies and their tools. And like, you just go across the supply chain. They're not, hey, forefront on really anything in the manufacturing supply chain.

33:31Dylan Patel:On the design supply chain, there's some things that they're starting to be similar par, but like cheaper or like a year or two behind, but cheaper. And that's like fine for a lot of stuff. An example of that is Huawei, right? Huawei in mobile phones was on par with Apple, like entirely. And they had become Apple TSMC's biggest customer when they were designing the best thing. And they are number one in telecom. And their tech is just literally better. and so when you think what happens is you know is is is china missing anything well it's like they don't they don't they don't have the best of much that you know today in the ai supply chain they have a complete package and a couple years behind and they'll figure out how to make it cheaper slash do more slash catch up and and create a robust industry but there's a reason like i don't think that like jensen is scared of amd really he's paranoid i mentioned he's paranoid i'm sure he's a little bit scared of them right like i think some of the things that they've done are reactions and competitive dynamics with AMD or Google's TPUs or whatever, right?

34:26Dylan Patel:There's a CoreWeave deal today. And I think that's directly the result of what Google's been doing. Yeah, the 2 billion pipe that NVIDIA announced. Yeah, NVIDIA invested 2 billion in CoreWeave. But what's more important is that that's like sort of just like the sticker. What's really relevant is NVIDIA is going to work with CoreWeave to acquire and backstop and all these things, the land, the power, the energy, the transmission, help build the data center, all this capital side stuff because NVIDIA has so much money, they can backstop CoreWeave doing it because CoreWeave then can be the one who generates demand.

34:59Dylan Patel:Anyways, there's like, because Google was doing that and they did that with like a couple of companies such as FluidStack and Terrowoof and Cypher. These are some public deals that have been announced. And so Google is doing that with TPUs and NVIDIA reacted, right? And so in the same way, I think NVIDIA has reacted to AMD and in the same way, I think the thing is, Nvidia is like deathly terrified of Huawei because Huawei has caught up to Apple and actually surpassed them as TSMC's biggest customer before they got banned. Right. They did just crush Nokia, Sony, Sony Ericsson, et cetera. Right.

35:27Dylan Patel:Like the entire telecom supply chain, they just like completely destroyed them. And there's so many other areas like they straight up made a folding phone. Right. You know, I have a Samsung folding phone. They have a folding phone that's better than Samsung's folding phone. Yeah. And it's like, bro, what? Like, you know, you know, Huawei is really, really cracked. And so, of course, they're terrified of, and Huawei is the most vertical company in the world. No company is more verticalized than Huawei, which then leads to huge innovations. It's something that we don't fully appreciate in the US, but when you travel in Europe, you see everybody who's like honors, honor phones.

35:59And it's like the footprint of Huawei is huge in phones in a way that people -

36:04Dylan Patel:But not just phones, you know, security cameras. Actually, I think they have like, you know - There's a lot of training on the - a captive group of testers. Exactly, exactly. I think Huawei is terrifying, right? And so, like, yes, their chips are not as good today. And is that already happening? I mean, obviously, the U.S. and China are the two biggest markets, but, like, for other markets, I don't know, UAE, Middle East, Europe, are NVIDIA and Huawei already head-to-head in deals? Huawei shipped a little bit, but, like, mostly just, like, sticker capacity. Like, there's nothing, like, no, like, I would say, like, a little bit as in like a few servers, not like a billion dollars worth of stuff.

36:41Dylan Patel:Right. The thing is, China's supply chain has to ramp up. Right. China, China's express goal is to have all internalized. But then like a company like Alibaba is like, I don't want to use Huawei. Right. Like I want to make, I want to use NVIDIA and just make the best freaking models. Right. Because that's my business. My business is not, you know, using a Huawei thing, but it's like, okay, it's being pushed upon me. There's other companies too, like CameraCon and so on and so forth. And so there's sort of like supply chain, you know, companies in China don't want to use Huawei, They're kind of encouraged, obviously, and pushed, you know, you must some local provincial government be like, well, you're doing this much business here.

37:14Dylan Patel:You got to do this, right? Like there's all sorts of like crazy stuff that, you know, pushing of companies to use Huawei. The challenge is Huawei can't manufacture enough, right? We've like done a lot of work on this and we've just put it for free, you know, instead of like to our customers, because it's like something that's like national security, which is how was Huawei actually building chips? Well, actually, they were using shell companies to get chips from TSMC and using different methods of sneaking HBM, which is memory, from Korea through Taiwan to China. All sorts of crazy stuff we've reported on.

37:45Dylan Patel:And people, it's like a whack-a-mole, right? They shut it down. Or tools that get shipped to China and they shouldn't be for making leading-edge chips, but they actually are. And all these sorts of things are happening because they can't make everything. And if they want to make the leading-edge stuff, they do need to rely on the foreign supply chain quite a bit in terms of the upstream supply chain, right? memory, logic chips, tools for fabs, chemicals for fabs, et cetera. Huawei cannot satisfy the market because there's not enough advanced leading edge capacity in memory, logic, you know, and all these other things domestically in China.

38:16Dylan Patel:And they're trying to build it as fast as they can, but that means there's just not enough to satisfy the market. So NVIDIA has a market. I think they'll figure out how to sell chips to China. And Jensen's in China, I think like right now, or was yesterday. And so like, he's clearly like a wheeling and dealing to try and get his chips into China because, you know, I think NVIDIA's argument is if we sell them chips, then they won't, you know, there won't be enough of as much of a domestic market. The feedback loop for software and everything else won't be there. That was sort of like really challenging, right?

38:42Dylan Patel:Like most of the open source software for AI has a lot of Chinese contributors, right? BLM and PyTorch, SGLang and like all of these other like libraries and things that are just like, you know, and it goes to low level software, especially, right? A lot of the best open source stuff is actually just from a Chinese company who decided to open source it. And same with models, right? And so it's like, okay, well, if they can't use NVIDIA chips anymore, then this open source stuff won't be designed for NVIDIA chips. It'll be designed for Huawei chips. And now does that weaken the CUDA mode? And now not only is China domestic, now they have a feedback loop internally, and then they can externalize across the rest of the world.

39:17Dylan Patel:So this is the argument NVIDIA makes. I'm not sure if I am like, I think my AI timelines are so fast. I'm not that fast. not in terms of AGI, but like, hey, AI is$100 billion of revenue across the industry. I think the industry could hit$100 billion ARR by the end of this year. Like$45.50 for OpenAI, like$35.40 for Anthropic, and then Vertex, DeepMinds models at Google, Gemini, right? And then Vertex API for Anthropic models and Bedrock APIs and Azure Foundry APIs. I think$100 billion is end of this year. That's a lot. And then what's the economic value of that$100 billion? Now, how much of that is in China, right?

39:59Dylan Patel:Like China's number is probably 10x lower, right? Because they just haven't been able to pervasively push AI, right? ChadGPT has a billion users, roughly. And, you know, then you add on Gemini and Meta claims they have 500 million users. I don't know. I think people just accidentally click like generative sticker or something. But like, anyways, like there's like, there's like a lot of usage of AI in the West already and it's going to climb, it's going to keep climbing. And like, you kind of have to get used to it. And so like, the question is like, do you, you know, what's, what's the economic benefit to the world?

40:29Dylan Patel:Right. And at the end of the day, this is an economic war, right? If the U S and the West win in AI and control, you know, more powerful AI systems that have this feedback loop that improved economic growth and weapon systems and whatever else, right? Engineering of grids and cyber attacks and all these sorts of things, they have this like advantage over China, then China will not rise to be the global hegemony. But without AI, China definitely will rise to be the global hegemony. They're just going to outrun America. And so the question is like, you know, that's, I think like the other view, right?

40:58Dylan Patel:And how fast are super powerful AI systems versus, you know, China building a domestic ecosystem for chips and models and everything that is a few years behind. Like what's actually the value, right? Like that's It's sort of like around restrictions and regulations. Where do the U.S. onshoring efforts fall in that category? What do you make of them from the Chips Act to like all the thing that is being built? Everything looks like it's massively delayed, by the way, which perhaps is not surprising. I think TSMC is manufacturing wafers and they're like building real wafers and there's real fabs.

41:30Dylan Patel:And like, you know, there's some other fabs that have been announced and like they're doing well. And there's like a bunch of like different kinds of plants, like a Korean company making a random gas plant in Texas for, you know, their chips, right? Like for chips and all these sort of things are happening. I think the chip stack did really well with its$50 billion. It's just, I don't think people understand the scale of the semiconductor industry. It is the most complicated supply chain in the world, right? It's much bigger than, you know, say manufacturing airplanes. It's much bigger than like, you know, really anything else, right?

41:58Dylan Patel:If you look at the top 10 companies like of the world, I think eight of them designed semiconductors, right? Now, obviously like Google designed semiconductors, but it's like, oh wait, no, but their cost of search would be like 10x higher if they didn't have TPUs. And TPUs were super optimized for search, right? Or like, you know, you go down the list, right? Like Meta serves recommendation systems with their chips, right? Like you go down the list, it's everyone is making their own chips. Apple devices would be materially worse if they didn't have their own chips, right? And you just go down the list.

42:26Dylan Patel:It's like, it's the most complicated supply chain. And they're spending something on the order of like$150 billion roughly in subsidies a year to the chip industry. We are doing 50 over like a decade. There's a difference in scale here, right? The collective total amount of like CapEx that has been spent in Taiwan is like 500 billion plus, right? Across the industry, across all the companies that are making semiconductors in Taiwan. And Taiwan doesn't have a domestic industry. How is$50 billion of subsidies going to change America's needle, right? It does move it a little bit, right? I want to be clear.

42:59Dylan Patel:Like the Chips Act is awesome. I don't understand why like EVs or like solar was given this massive, massive like trillion dollar package. Semiconductors were only given 50. Like semiconductors need a lot bigger package to actually incentivize the on-shoring. I think what's happened so far has proven that it's working well. TSMC is literally making chips for NVIDIA and Apple and AMD and others in Arizona today. Right. And I think that's really great. Is your sense that the broad American government is just aware of all of this? Well, you know, the chipset only passed because the automotive prices went up because car manufacturers are like the worst because they do just-in-time inventory, right?

43:39Dylan Patel:Or worse, but it's just like a thing, right? Just-in-time inventory systems. COVID happened, sales plummet, fabs that were making, you know, random power ICs or random microcontrollers for engines got repurposed to the boom from COVID, which was data centers and PCs and smartphones. so that stuff was booming and then when people were like oh wait actually like you know i have some money i stayed at home i didn't go out i didn't drink i have a lot of i have some cash right let me buy a car they went on about cars and cars started skyrocketing in prices oh let's restart and let's let's oh yeah can you sell me that microcontroller for the engine again it's like no i i'm making a slightly different microcontroller that works for you know uh let's say a keyboard or a mouse right or whatever and it's like and and they actually didn't just leave me flat footed and they were like a partner through COVID, right?

44:23Dylan Patel:You know, versus you just left me screw you Ford or whoever, Toyota, um, or automotive OEM, you know, you know, that supply chain. And so chips act did not get passed only got passed because that happened. And people were like, Oh my God, the semiconductors are why cars can't be made. If that didn't happen, we wouldn't even have the chips act. It's like, it's like silly. So like, I don't know, like, I think, you know, whereas like, and even though that's what was pitched to all the senators, like I know people who are running around Capitol Hill, just pushing that narrative and story and that's why it finally got passed in reality it was all for advanced leading-edge chips right nothing that goes in a car right and so it's like this like funny thing so in other words do you think my words my words not yours but is it is it hopeless that the u.s is going to i'm very optimistic okay i mean do you think that's a world where the u.s just decides to invest in semiconductor at the scale that you know i thought we just needed a bigger chips act but look, Trump's kind of gotten TSMC to promise to invest a fuckload more and they're moving on it.

45:19Dylan Patel:Right. They're like actually like just building it. It's like, I'm going to tear off the shit out of you unless you build a fab. And it's like, we'll build a fab and they're building it right now. The timelines for fabs just takes forever. Cause again, it's the most complicated thing in the world. The cleanest space in the place in the world is not like a hospital or a biotech lab or whatever. It's a semiconductor fab. And the most expensive tools in the world are not, you know, any of these medical tools or whatever it's semiconductor tools or it's not a rocket it's a semiconductor tool right like everything you know i describe it as um i remember when i was a kid i was like i want to be a rocket scientist and then i was like oh i want to be a surgeon and i'm like wait chips are like rocket surgery uh but even cooler right like i think um anyways like sort of like there are there are fabs being built in america they won't take america to self-sufficiency i don't think that's a relevant i don't think that's a goal relevant like that's relevant right like globalism is generally just good hot take uh like in terms of economics turn this into a short a youtube short globalism is good dude you're gonna get me like canceled i tweeted about ice it was a complete joke but so many people got mad at me because i can't be you know i'm too i'm too much of a joker you know these are serious things yeah no i know the i know the feeling yes anyways um i think i think you know i think we are building fabs and i think it's like gonna move and now even elon's talking about building fabs now because he sees the shortages in the world right uh there's a lot of semiconductor related shortages for building out ai and so i don't think it's hopeless i think i'm like very optimistic that we're going to do more and more and more and maybe this administration threatens tariffs and they get the deals and the next administration comes back with the carrot if it is the democrats whatever happens i don't know So I was at a comedy club on Sunday night and like he's like, oh, I use Chad GPT.

46:58Dylan Patel:And then like there are a couple of people who booed. He's like, yeah, I'm one of those guys. I know. And like it's like, wow, people hate AI. And that has not even started, right? Like the actual impact of AI. Or like New Jersey power prices are up, right? Is it because of a data center? Well, New Jersey, the governor's election, like I think literally like there's like an election that changed recently in New Jersey because power prices were up. and people blamed a Microsoft Nebius data center in New Jersey for that reason. But in reality, that data center has nothing to do with power prices going up.

47:30Dylan Patel:It's super storm standee, like five years ago, knocking or whatever, how many years ago, knocking down the state's electrical infrastructure and then improving all these improvements. And then those improvements have to be paid by someone. And it turns out the consumer has to pay for them with higher power prices. Right. And so like, you know, like there's like, there's a lot like going on in that regard. right um that kind of is uh sad um and and people hate ai and they're blaming ai on it and artists hate ai and like you know you see all this deep fake stuff and like i think i think it'll be the hottest button issue especially as like we're really getting into like i think last year google spent three billion dollars on waymo and we're waiting for their guide for this year three billion dollars on waymo taxis but their their waymos went from like 300k to like 100k or 90k the new waymo car and they're gonna spend more than three because they've just launched in like four cities now, right?

48:18Dylan Patel:Or five cities. And they're testing in a lot. And the same with Robotaxi. Like, people are going to hate AI for that reason. People are going to hate AI because the slop on the internet. People are going to hate AI because, you know, the perceived job replacement. People are going to hate AI for all these reasons. And so, yeah, it's going to be a hot button political issue, don't you think? Yeah. Talking about that. So, CapEx, is there a CapEx bubble? Are we investing too much or actually are we investing not enough given what what you were saying earlier about the rate of revenue increase and therefore implied demand that you expect for this year?

48:52Dylan Patel:I'm obviously a maxi. I think we're going to need a lot of infra. And I think I'm literally paid to like analyze the supply chain and do consulting. Like that's what my company does. So like, obviously I'm very biased. I think we're pretty good at calling when things go down though, right? Before like part of the supply chain or whatever. But anyways, you know, again, going back to the economics of it, It's north of$100 billion of revenue exiting this year for AI from a base of sub$1 billion, gen AI, because ads and stuff is already a multi-hundred billion dollar AI industry. Go back to 2023, it was less than a billion.

49:26Dylan Patel:In 2024, I don't know exactly what number, maybe let's call it 10. And 25 was maybe like 30, 40. It'll be north of 100 easily. If you're talking about$100 billion of revenue, let's say at a 50 % gross margin, So that's$50 billion of gross profit and$50 billion of COGS. That$50 billion of COGS needs to run on infra, which costs roughly, if you're talking about five-year depreciation, call it$250 billion, right, of infra for$100 billion of revenue. Okay, what is the actual spend on AI and for this year? It's going to be like, I mean, it depends on what layer. If you're talking about energy, those are longer-lived assets and all these other things, right?

50:05Dylan Patel:Data centers are longer-lived assets. The chips are not as much. People are putting CapEx down. And the hyperscalers CapEx is going to be like$500 billion this year or something like this. And then besides them, there's also a lot more CapEx elsewhere. And so, you know, is it a bubble? I mean, theoretically, like, you know, it's twice as much as it should be, but it's also like, well, no, there's an R &D component to this. And the excess spent that wasn't revenue generating last year is what led to models being so good this year and led to like everyone who can using cloud code and like that changing their life.

50:39Dylan Patel:this is like, it's not a bubble, right? I don't think it's a bubble yet. I think if AI model progress stops, and that's the main thing, right? The moment model progress stops, all the spending is for naught. But so far, we've had consistent improvement. As you put in more compute, you get more performance and better models. Yeah. Model performance being a lagging indicator of hardware progress or data center capacity. Or yeah, of CapEx, right? Ultimately, the CapEx that Microsoft spent in 2024 for OpenAI is what results in 2025 for OpenAI or CoreWeaver or whoever is what results in their models being so good this year.

51:12Dylan Patel:Same with Anthropic and Amazon Google and their models now being so good now is that CapEx. And actually, they still haven't paid for those chips yet because those chips still have a useful life for another few years, right? I think model progress is very clear. The moment that stops happening, right? If we hit a wall, there's no new research directions, then it's cooked, right? And that assumes that a better model leads to more demand, which is a reasonable assumption. Yeah, for sure. I mean, there's still the adoption curve, regardless of how good the model is in the enterprise. Yeah, but like 2 % of GitHub commits today are CloudCode.

51:48Dylan Patel:Yeah. As in committed by CloudCode. You can disable that where it's not automatically committed. But 2 % of GitHub commits today are CloudCode. $2 trillion of software wages paid in the world. If it was 2%, then you're like, wait a second. this is an insane amount AI is under earning the value that it's producing in the world right? By a significant margin already today. Boris Cherny from Cloud Code who we had on the pod was saying that he's written all of Cloud, what is it called, Co-Work the new product entirely we're very much in that world one of my roommates I was asking him because he's always been a really low level good programmer and he started, you know, I was like, he's like, he had this, um, holiday obsession, right?

52:35Dylan Patel:I mean, he was using cloud code for work already, right? Like whatever. Um, but he had this holiday obsession. We got into playing age of empires to myself, you know, my roommate, a handful of people from like open eye GDM anthropic. We just would do land parties of AOE to over the holidays a bit. Not, not like Christmas, but like a little bit before a little bit after, you know, cause most of us went home for Christmas. Um, but like we do these lands, my roommate got so obsessed with like the game that during a christmas week because he didn't go home he just stayed in san francisco um he just worked on an rts game and he built an entire rts game and i think i kid you not i think he he used like ten thousand dollars of clod in one week and built an entire rts from scratch uh about age like but instead of like being a standard rts where it's like oh age of empires where you advance through ages or starcraft it is it is an rts where it's China versus the US and you're in the AI race and you go from the start of the information age all the way through to, you know, AGI and like robots and humanoids and like space-faring civil...

53:36Dylan Patel:Like, it's crazy. He built it in a week and he didn't type a single line of code, right? He only dictated to the model. And he told me, yeah, like we have an indicator internally at Anthropic where you see how many people actually write code now. There's only a few holdouts left. But I guess the question to the bubble is really a question of... timing as well, right? It's whether the build, which is the supply side and the demand side, are going to land sort of at the same time. Is that fair? Yeah. But also the economics of like, say, you spend, let's say you spend, you build a gigawatt, you put down roughly$50 billion across, you know, the data center, the chips, the networking, blah, blah, blah, blah, blah, right?

54:15Dylan Patel:Let's say it has a five-year useful life. So it's$10 billion a year. Is it a bubble if the first year you didn't make any money, it's zero? The second year, it's zero. And then third, fourth, fifth year, you're at 50 % gross margins. And so you make 20, 20, 20. Now you've made$60 billion off of this$50 billion investment. It's not the best return on invested capital, but it did pay for itself. Is that a bubble? Well, that's what's happening today is that people are spending all this money on infra and there's no return for a lot of it, right? A lot of it is just doing research and trying to get adoption and it's free users.

54:46Dylan Patel:And what does that mean? Yeah. Depends a bit on the timing. That's the timing though. But that$50 billion CapEx was spent in year one. What about energy? In the data center world, you had this fun post about the gas replacement for energy. So is AI basically destroying the grid? What if the utilities were willing to let it? But I think the utilities are so slow and dumb that they don't want to. Not destroy, but like expanding the grid. I think the U.S. could have a way better grid, but we just don't want to. Like no one's made the effort or initiative. You know, there's not enough power. America's not built power for 50 years, really, right?

55:23Dylan Patel:It's like converted from coal to gas and like things like this, but like really just have not built wholesale new power on a large scale. And there've been a lot of times where the industry blew up, right? Independent power producers, IPPs, have blown up multiple times in the 2010s when Korean and Japanese investors like flooded the market with, because they saw such a good return there. Or before in the early 2000s, power was growing a little bit for a little bit. And so people overbuilt on power. So power industry has been burned a couple of times. So no one really builds power. And then you've got data centers now all of a sudden coming online and going from 2 % to 10 % of the U.S.

55:55Dylan Patel:grid in just a handful of years. And so you've got this humongous, humongous change in the industry. We don't have the labor, right? I think ultimately that's the biggest problem is the equipment and the labor. And equipment is basically, you know, again, labor and time takes time to build a factory so you can build the things. I think the equipment side of things will be solved like more reasonably. And one example is like gas, right? People initially thought, oh, you can only use like the two vendors, right? Siemens or GE Vernova for gas turbines. They have the best ones, the most efficient ones.

56:22Dylan Patel:It's like, okay, well, like, okay, also Mitsubishi exists, and they're ramping up production fast. Oh, Doosan and Korea exist, and they're ramping up production fast. Oh, actually, I can just take Cummins engines, right? Like, you know, if you've ever, like, ridden a pickup truck or, like, you know, like diesel trucks, like, everyone loves Cummins, right? You know, you see the Ram on the street and it has the Cummins, like, badge. It's like, that's like an aura symbol for a certain kind of redneck from South Georgia, which I have a little bit of. Anyways, I don't have a truck. I have, though. But anyways, there's all these engines.

56:50Dylan Patel:People are figuring out how to make the equipment. Solar sucks. It's too intermittent. Wind sucks. It's too intermittent. Nuclear sucks. It takes forever to build. Coal sucks. It's way too dirty. How do you make power for data centers besides gas? Okay, the grid's not willing. Just put the gas on your site. That's what Elon did. Now everyone's doing it, right? This other cool poster just last week or two weeks ago that was about the water consumption. Do you want to talk to that? Yeah, yeah. So there's this annoying thing where everyone's like, oh, AI is using all the water. Oh, wow. AI and data centers are going to use up all the water and now we don't have any water.

57:24Dylan Patel:And it's like, that's so silly. Water is a distribution problem, not a we don't have enough problem, right? You look at California. It's like California has shitloads of water. But people decide to make oat milk, which consumes like 1 ,000x the water of anything else, like regular milk even. And cows obviously consume a lot of water. But anyways, like, you know, data centers consume very little water actually, right? So the U.S. grid will get to like 10 % of power by like 28, 27 is data centers. For water consumption, it's not even gonna crack 1 % by the end of the decade. And what was the metric?

57:58Dylan Patel:And so the comparison we made is because like, you know, it's a bit of a shit post, but it was like serious research. Basically like we were doing serious research because we keep getting this like question and debunking it and we would do it seriously. But then I was like, no, no, no, this is like too like complicated. Like let's make it very simple. So I was like, guys, why don't we just compare it to like hamburgers? Right. Cause, cause you know, I've heard that argument from some like vegetarian people before or some Hindus or like, I'm Hindu myself. Although, you know, and I do eat beef sometimes, you know, like I'm Hindu, but like, you know, so we made this comparison to hamburgers, right?

58:32Dylan Patel:Hamburgers require a shitload of water because cows, you know, for them, they require a ton of water. And when a cow is taking a lot of water, it's not the cow itself. It's all the feed you're feeding them. Right. Because no one grass feeds their cows and just lets the rain take care of the grass. They either rain the grass or most likely they do mass industrial farming of corn, soybean, alfalfa, et cetera, which uses shitloads of water, right? Or almond milk uses tons and tons of water. Produce is the main user of water. I think the metric was the entirety of Elon Musk's Colossus Data Center, right?

59:10Dylan Patel:uses as much water as two and a half in and outs. Because that's, you know, you do the calculation on how many, how many burnt, what's the average revenue per in and out? And how many hamburgers does that translate to, right? If everyone's ordering like a combo, right? Okay, let's ignore the drink. Let's ignore the fries. Let's just talk about the hamburger. Let's ignore the bread, which does have grain. Let's just do the meat and the cheese. And all of a sudden, all this water is, there's so much water, right? Like a single query, like all of your AI usage from chat GPT of the average user is like a hamburger, right?

59:43Dylan Patel:Like it's like, okay, this is nothing, right? You know, because these things are, the data centers actually are like, they're mostly closed loops and like, sure, they evaporate some water for like cooling reasons, but like by doing evaporative cooling, they're using less power, right? And that's actually better for the environment than not using evaporative cooling. There's all these reasons why. This myth or hoax of AI using all the water is just nonsense, right? Like Meta's data center in Louisiana is getting protested because the water, it's going to be the largest data center in the world.

1:00:11Dylan Patel:It's going to be like four or five gigawatts, at least announced so far. We're tracking some other ones that may be as big or bigger. But Meta is getting protested because the local population around that area is like, oh, the water's dirty. It's because of this Meta data center. And like there's these trucks on these big trucks on these back roads that used to be empty completely. They're just like mad and annoyed about that, right? But at the end of the day, what actually made the water dirty is that that's an area where you go fracking. Like fracking is absurdly worse. And almost all of that gas is being shipped to an LNG terminal and being shipped to Asia, like, you know, like Japan or Taiwan or China or Korea and some Europe as well.

1:00:48Dylan Patel:Right. Like like actually all of this water is dirty because of regulation slash fracking. Like I support fracking, by the way, but that's that's an insane take to maybe. But like water usage is like not a relevant argument. Are you bullish on the sort of energy companies? I'm thinking Constellation for Nuclear or Vistra, I guess, as an independent power producer. I think IPPs will do well. I think IPPs can secure contracts at premiums to what they've previously been able to for new power plants that are either dedicated or grid connected, but come with a pairing of a grid load. right for example utilities won't let you just do data centers now but if you come with a a pair right you're like hey i'm going to build this massive data center but we're also going to have this massive uh power generating asset right say you know whatever it is right some ipp they're going to partner with and they'll build the load and the uh consumption even if it's connected through the grid for better stability and more reliability um or it's not it's behind the meter i.e not connected to the grid at all um like some part some data centers like partially like Colossus from Elon, the original one, or part of Abilene's Texas OpenAI, right?

1:01:59Dylan Patel:Like Crusoe. There's a lot of room for power producers to get outsized returns. I'm not necessarily bullish nuclear. Existing nuclear, fine. Yeah, it can find a higher buyer, higher priced buyer. But majority of it will be gas. But like you can do like renewables backed by gas and then just turn off the gas and like it's cost more, but whatever, right? Or you can do wind backed by gas. And why not nuclear? Takes too long. Takes too long. No one can build nuclear fast. Mm-hmm. even China takes like five years to build nuclear, right? Like it's complicated. It's unsafe, right? You know, and I love nuclear.

1:02:30Dylan Patel:I wish it would work. It's just not relevant in the timescale that like AI's power is going crazy. But yeah, there's a lot of interesting stuff. Like have clients would like, had a client buy a coal plant and we were advising them on the transaction based on, they just like showed up and they're like, yeah, we want to buy power assets. We believe in this power story. It's like, okay, great. So yeah, so here's all of the like power plants that we know of. Like you can get some from EIA, blah, blah, blah. which are these like, and then we like work through the economics and we looked at new data centers being built in the region and all this.

1:02:58Dylan Patel:And then they decided to buy a coal plant and they restarted it. And they're like making tons of money now because now someone, a certain hyperscaler wants to buy the entire pipeline of power and put a load near it, right? Instead of just being a grid connected asset. So it's like a super awesome investment. So like, you know, power is, power is going to do great. Yeah. I was going to talk about piece dividends of the AI boom. Generally. Yes. Right. Like hyperscalers are paying for transmission upgrades, which people will benefit from, right? Or like, you know, investors are obviously going to benefit.

1:03:28Dylan Patel:People who work in the industry, electricians' wages are skyrocketing, you know, et cetera, right? Like plumbers' wages are skyrocketing. So there's like a lot of trades that are doing really well too. I think that's definitely also part of it. Yeah. I wanted to come back quickly to that NVIDIA and CoreWeave deal that you mentioned as we sort of closed the discussion on CapEx and a bubble. It seems like there is circular deals, but also a lot of debt kind of like flushing around. I don't know the specifics of that deal, but I did hear variations of this where effectively you have a large player guaranteeing the debt being the last recourse for a lot of infrastructure build.

1:04:08It is sort of this plus the whole like Oracle commitment. There is a fragility into this whole thing that can be a little unnerving. What do you make of it?

1:04:19Dylan Patel:I think it's like completely fine and I think like people are like freaking out and making narratives where there really is shouldn't be one it's like well okay Google doesn't have enough data center capacity they need people to build data centers but no one can build a data center because they don't have the capital like don't have you know in many cases capital is not the you know they don't have capital right or like no one will give them a loan because they don't trust some random fucking company and it's like but then Google's like well no we've due diligence then we think they can build it here we'll like even guarantee we'll buy the thing or start using it once they build it.

1:04:47Dylan Patel:You know, just having a customer alone spoken for it was enough, right? In the case of CoreWeave, they were actually able to, no backstop, right? They were able to just say, hey, look, here's our Microsoft contract for this many GPUs. I want to put in that data center, that data center, that data center. Here's the contract for renting those GPUs. I want to hire these people. I want to do this. No one will like, they don't have any money, but then they were able to like have it work out because they were able to get people to lend to them. I think like CoreWeave did that and there was no circular financing, but that was when there was like the scale of investment was like single digit billions or less than a billion, right?

1:05:15Dylan Patel:Now the scale of investment is hundreds of billions. And so the question is like, oh, well, if I want data center capacity, how do I get data center capacity? I just go to everyone who's going to build it, looks smart, is smart enough to do it, but can't afford to do it and tell them I'll take it. And in fact, I won't just take it. I'll go to your debtor and be like, I'll guarantee you. Because, you know, obviously you're a new company. I vetted you, but the debtor hasn't. And so, you know, like, you know, they don't want me to just be able to walk away. Because like in the Microsoft CoreWeave deals, Microsoft could have walked away if CoreWeave fucked it up, right?

1:05:45Yeah.

1:05:45Dylan Patel:I mean, yeah, there's always like sort of like cancellation or whatever possibilities. And so there's just a further form of guarantee as far as on like a lot of these backstops, as far as on like Oracle getting the money and then OpenAI getting money and NVIDIA, you know, paying and it's a whole circular. It's kind of nonsense because it's like NVIDIA is getting equity in OpenAI. They're basically saying, hey, every gigawatt you buy will also buy some equity. Yeah. Right. OK, well, cool. Now NVIDIA owns an asset which they think is valuable, OpenAI, right? Right. OpenAI is turning around and is like trying to rent those, use the equity they buy.

1:06:17Dylan Patel:What are they, what are they, what was their use of equity? People's cash pay isn't that great. Right. It's mostly just 99 plus percent of their spend at the company is probably just compute. Yeah. So, so sort of like, it's like, okay, well then I raise this money. I'm going to do the, the whole thing where I explained earlier, right? Year one and two, I lose money. Year three, four, five, I hope to make money on it. Right. And OpenAI has been doing that. Right. So I'm going to, okay, I'm going to go out there. I've raised$50 billion. I've raised$10 billion. I'm for five years for$65 billion. And I've rented that contract.

1:06:48Dylan Patel:And now I only have enough to pay for the first year, to be clear. But I think, you know, you trust me, Oracle, you think I'm going to grow and you think I'll be able to pay for it. Oracle's like, yeah. Or if you're not, I think I'll be able to sell it to someone else. So like, okay, cool. I'm going to spend$50 billion this year to build that data center. And this is like for a gigawatt. And so is it like circular that OpenAI is every amount of GPS they consume, maybe it's an investment, that investment is turned around to pay for the first year of the rent of the cluster or second year. And then first two years go, you know, it's sort of like, it's fine.

1:07:17Dylan Patel:Yeah. Yeah. Like, it's like, it's like, it is a little bit funky, but like, I don't think it's a big deal. Yeah. Love it. Contrary intake. Maybe let's finish with models and the software side of things. We talked extensively about hardware and supply chain and all the things. I get a sense that you are super, super bullish on what's happening next in AI. Your roommate, Sholto, I assume was the roommate that you were talking about earlier on this pod, effectively making the point that we're just starting to scratch the surface and there was so much low-hanging fruit around RL and all the things.

1:07:50You're in Silicon Valley circles. Is that your sense as well? And what are you tracking on the model side?

1:07:56Dylan Patel:One thing is like, you know, simple stuff like GitHub commits. Other things are like, what's the amount of usage? How much are people using? Like all these sorts of things. I think there's so many different alternative data sources for tracking AI model progress. Area tokenomics, token economics, tokenomics. And so that's like an entire practice for us. Are you rebranding the term from crypto? Yeah, I don't believe in crypto people. Like I've always hated them. So now you're taking the term. Yeah, yeah. And Jensen's used it now. So I've like, I've convinced him to use the word. He's used it as sovereigns.

1:08:27Dylan Patel:And so I think, I think we won. That's awesome. Congratulations. I've said it to him. We've written it in articles. It's an entire practice of consulting that I just started. I started in like 23. was token economics. And we've been trying to build out these like, you know, but basically I think the main things are like, people who don't code can use cloud code now, right? I think people don't understand that. Like, even if you don't code, you've never had any training in software development. You've never had a job as a software developer. You can code. Let's take an example of what's one of the analysts that my company did, right?

1:08:53Dylan Patel:Comes from a engineering background, but on like semiconductor systems, right? Like worked on mechanical systems, worked on these sorts of things. And they coded this thing, which was they want to do an analysis of area of clean rooms, right? Clean rooms are the building that the fab has all the tools and the most complicated kind of building in the world, has all sorts of chemical systems and all this. Area of that, a company who builds these systems and revenue of that company, right? And so it was like, okay, we have this fab data set, pointed at it as like, hey, here's this fab data set. What's the square footage of all of them?

1:09:28Dylan Patel:And we have this thing that we built, which just pulls with Cloud Code separately, which for data centers and fabs and everything else just calculates the area of something from a satellite image, right? Very simple. So we have the square footage of all these things. What's that? Here's the company name. Okay, go find the filing. So it dug through all these filings. It pulled the data, right? Okay, great. Now tell it to compare these two. Make a chart. Great. Oh, wait, there's this like weird inflection. Oh, that's because they bought a company five years ago. Can you do a pro forma of this analysis without those financials of that company they acquired?

1:09:58Dylan Patel:Okay, great. And then like we were able to like figure out an investment case for clients, as well as like, you know, some other interesting details from someone who's never really coded, just using Claude Code and it like doing this all. And this is like not even there. And it wrote the note and they just like, they didn't even like work on this full time for like three hours. They just told the model and would go work on other things and told the model and worked on other things. They just did this. People don't understand that like the skill sets that like, I think like if you go talk to an analyst, right, a very junior analyst at any company, right?

1:10:29Dylan Patel:Whether it's venture or especially growth venture or public markets or private equity, their job is like finding data, cleaning it, making charts. It's like, this is Claude Code now. You don't need junior analysts. Just like a lot of companies have stopped hiring L4 engineers because it's useless. Why would I hire an L4 engineer? I just tell Claude to do it. You sort of like have this has happened. And this is a really big like shift, I guess, like is that like low level knowledge work just doesn't matter, right? Why would I use Excel when I can just tell Claude to manipulate CSVs? Why would I use Word when Claude will just generate the markdown and I can copy and paste the markdown directly into our WordPress and then that WordPress is fully formatted now and it's like, oh my God, what's the point of Word, right?

1:11:12Dylan Patel:And what's the point of doing all sorts of stuff? I think when we look at model progress, that's just Opus 4.5. OpenAI's new model I think will be better than Opus 4.5. It's coming somewhat soon in March-ish timeframe, maybe February, March-ish, but yeah. because OpenAI has a better RL stack than Anthropic today. It's just their pre-trained models suck compared to Anthropic's pre-training, right? And so like if they catch up a lot on pre-training and keep their better RL stack, they would actually have a model that's much better, right? Flip side, Google has a better pre-trained model than Anthropic or OpenAI, but their RL stack sucks.

1:11:43Dylan Patel:So if they catch up on RL, like these models are gonna get ridiculously, and then Anthropic is obviously advancing as well, right? And then you look across the ecosystem, everyone's advancing, really fast progress. these moments are happening, right? You know, chat GPT was a moment. Ghibli was a moment. Those are more consumer. Those are less like, I mean, chat GPT is everyone using it for work too. But like, I think Cloud Code is like a new moment, right? Opus 4.5 on Cloud Code is a new moment where the way you work has forever changed. And so now we're trying to force everyone in my company.

1:12:10Dylan Patel:There's 54 people here. I think like half of them have coded. The other half, we're trying to force them to use like Cloud Code. And it could be like, oh, well, actually you come from a semiconductor consulting background. Oh, you come from like a semiconductor, like engineering of like packages. oh, you worked in a fab, right? Like these kinds of people, they're using Cloud Code now, right? And their productivity is being boosted. And it's like, you know, workspace, Cloud Workspace is new. It sucks compared to Cloud Code, but it'll get there, right? He said he coded it entirely in Cloud Code.

1:12:37Dylan Patel:You know that, right? Or it was on your pod, right? Yeah, yeah. So like, you know, I've heard that. And I think maybe that might've been from your pod, original disclosure. My pod was before that, but yes. Oh, okay, okay. That's the guy on my pod, but he subsequently said that. Ah, okay, okay. I think it's like a brand new age and like there's so much low-hanging fruit. As Shilto said on the episode when he was here, there's so much low-hanging fruit. Yeah, I mean, for the models progressing and then I think model progress will translate to revenue. Adoption is difficult, but like actually the UX of Cloud Code sucks, but like give it six months, the models will be good enough that the UX can be like talking to it.

1:13:10Dylan Patel:And you don't even have to have like, you know, CLI integration, right? It's something even easier. Or like Cloud for Excel was released recently and it's like not bad. You know, building models and like all these sorts of things are just going to be like tell someone, right? Like why tell a junior analyst, right? When you can just do it yourself. I think it's a whole new world and it's a$2 trillion of software work, but also of wages, but it's also, we have more north of 2%, 2 % is Claude and then, you know, there's Kodaks and Cursor and all these other guys. So probably like 5 % of code committed today is AI generated, if not higher, marked as AI generated.

1:13:40Dylan Patel:What's going to happen when normal workers who do spreadsheets and office processing start automating their workflows? I think it's a whole new world. And speaking of Sholto, we both agreed that he was a perfect specimen. dude i've i've been i'm straight but i've been accused of being uh homosexual which is perfectly fine for for how much i like praise this man because like think about it right he's like six foot four he's like really good looking he's like australian accent sounds amazing like you've heard his episode i have like a annoying voice probably his voice sounds amazing he's absurdly good at coding he was an olympian level fencer like like he picks up any sport he's really good at it right because he's athletic.

1:14:21Dylan Patel:It's like, holy crap, you're a specimen. Yeah, yeah. So this is going to be eclipsed incentive for sure. Yeah. It must be, you know, I guess maybe some people don't follow the play-by-play on Twitter and haven't heard of the fact that all of you guys are roommates. You're a roommate with Sholto and then with Dworkish. And Dworkish is like the podcaster's podcaster. So it must be absolutely— What's a podcaster's podcaster mean? the podcaster that other podcasters aspire to become or learn from. Yeah, yeah. When he's preparing, you know, it's like he's so locked in and he prepares so hard for interviews.

1:15:02Dylan Patel:It's great. No, he's just incredible. And then he might only say like 100 words on the episode. Yeah. But he's prepared so hard. And then like I think people just realized, oh, wow, he's not just like, you know, it's like, oh, he just has good guests. No, no, no. Like he's preparing really hard. but you can't tell if you're not like realizing that and then once he started writing more and he started writing more people like oh wow he's actually really really smart it's like yeah because he's studying like crazy like it's like oh i'm interviewing an ai researcher who worked on this i'm gonna try and train a freaking model yeah right it's like that's the level of like commitment he goes to when he records this stuff what do you guys talk about when you bump into each other is that is that ai non-stop or you talk about everything but ai with sholto it's like the age of empires game you know because we got super into it for a bit we talked only about that and his RTS that he made with Dwarakash.

1:15:46Dylan Patel:I mean, it's all sorts. It's like normal roommate stuff. It's like, how's your dating life? Oh, okay. You want to know how to date? It wasn't well. It didn't go well. Okay, well, okay, yeah. You know, like, oh, you know, like, that's me. That's me. You know, my dates don't go well. No, just kidding. Or like, it's like, oh, you want to like have dinner? We can invite a few friends. Like, yeah, great. Or like, you know, it's like all sorts of like normal stuff too. Obviously, we also do talk about a lot about tech, right? Like we are like, this is our lives. and tech is the most fun thing. Awesome.

1:16:13Well, great. Great San Francisco lore. Dylan, thank you so much. That was absolutely fabulous. We enjoyed it. Learned a lot. So we appreciate your coming on the pod. Thank you so much. Hi, it's Matt Turk again. Thanks for listening to this episode of the Mad Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests. Thanks and see you at the next episode.

From the publisher

Dylan Patel (SemiAnalysis) joins Matt Turck for a deep dive into the AI chip wars — why NVIDIA is shifting from a “one chip can do it all” worldview to a portfolio strategy, how inference is getting specialized, and what that means for CUDA, AMD, and the next wave of specialized silicon startups.


Then we take the fun tangents: why China is effectively “semiconductor pilled,” how provinces push domestic chips, what Huawei means as a long-term threat vector, and why so much “AI is killing the grid / AI is drinking all the water” discourse misses the point.


We also tackle the big macro question: capex bubble or inevitable buildout? Dylan’s view is that the entire answer hinges on one variable—continued model progress—and we unpack the second-order effects across data centers, power, and the circular-looking financings (CoreWeave/Oracle/backstops).


Dylan Patel

LinkedIn - https://www.linkedin.com/in/dylanpatelsa/

X/Twitter - https://x.com/dylan522p


SemiAnalysis

Website - https://semianalysis.com

X/Twitter - https://x.com/SemiAnalysis_


Matt Turck (Managing Director)

Blog - https://mattturck.com

LinkedIn - https://www.linkedin.com/in/turck/

X/Twitter - https://twitter.com/mattturck


FirstMark

Website - https://firstmark.com

X/Twitter - https://twitter.com/FirstMarkCap


(00:00) - Intro

(01:16) - Nvidia acquires Groq: A pivot to specialization

(07:09) - Why AI models might need "wide" compute, not just fast

(10:06) - Is the CUDA moat dead? (Open source vs. Nvidia)

(17:49) - The startup landscape: Etched, Cerebras, and 1% odds

(22:51) - Geopolitics: China's "semiconductor-pilled" culture

(35:46) - Huawei's vertical integration is terrifying

(39:28) - The $100B AI revenue reality check

(41:12) - US Onshoring: Why total self-sufficiency is a fantasy

(44:55) - Can the US actually build fabs? (The delay problem)

(48:33) - The CapEx Bubble: Is $500B spending irrational?

(54:53) - Energy Crisis: Why gas turbines will power AI, not nuclear

(57:06) - The "AI uses all the water" myth (Hamburger comparison)

(1:03:40) - Circular Debt? Debunking the Nvidia-CoreWeave risk

(1:07:24) - Claude Code & the software singularity

(1:10:23) - The death of the Junior Analyst role

(1:11:14) - Model predictions: Opus 4.5 and the RL gap

(1:14:37) - San Francisco Lore: Roommates (Dwarkesh Patel & Sholto Douglas)


More from The MAD Podcast with Matt Turck

All 44 episodes
Dylan Patel: NVIDIA's New Moat & Why China is "Semiconductor Pilled”The MAD Podcast with Matt Turck · 1 h 17 min
Listen in VO