Dylan Patel on the AI Chip Race - NVIDIA, Intel & the US Government vs. China

22 Sep 2025 · 1 h 40 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

a16z Podcast Episode Notes

Episode Title

Dylan Patel on the AI Chip Race - NVIDIA, Intel & the US Government vs. China

Episode Summary In this episode, host Erik Torenberg is joined by Dylan Patel (Chief Analyst at SemiAnalysis), Sarah Wang (General Partner at A16Z), and Guido Appenzeller (A16Z Partner and former CTO of Intel's Data Center and AI Business Unit) to discuss NVIDIA's surprising $5 billion investment in Intel. This partnership signals a significant shift in the semiconductor landscape, potentially reshaping the AI and cloud computing sectors amidst the ongoing US-China tech tensions.

---

Key Discussions

  1. NVIDIA and Intel's Strategic Partnership
  2. Surprise Investment: NVIDIA's $5 billion investment in Intel is unprecedented, given their historical rivalry.
  3. Market Implications:
  4. For NVIDIA: Seen as a "Buffett Effect"—showing confidence in Intel's future, driving stock prices up.
  5. For Intel: A much-needed lifeline amidst its struggles in the semiconductor market.
  6. For Competitors (AMD, ARM, Huawei): An alarming development, indicating a major strategic shift in the competitive landscape.
  1. Industry Perspectives
  2. Dylan Patel's Analysis:
  3. Sees the partnership as a poetic turnaround, with Intel collaborating with a former rival.
  4. Notes a shift in the semiconductor market dynamics, with potential impacts on product offerings.
  5. Expresses cautious optimism regarding Intel's ability to capitalize on this partnership to regain market position.
  • Guido Appenzeller's View:
  • Emphasizes customer benefits in the short-term due to enhanced collaboration.
  • Raises concerns about Intel’s internal products and competition.
  1. US-China Tech Tensions
  2. Huawei's AI Roadmap:
  3. Discussion on Huawei's recent unveiling of their AI capabilities and tech advancements.
  4. Impact of US-China tech bans on Huawei's chip production and competitive positioning.
  5. Potential Shifts in Global Chip Production:
  6. The episode illustrates the geopolitical tensions affecting technological development and supply chains.
  1. The Future of AI Chips
  2. NVIDIA's Moat:
  3. Discussion on NVIDIA's historical developments that built its competitive advantage.
  4. The importance of rapid execution and innovations in chip design.
  • Emerging Hardware Technologies:
  • Introduction of new chip architectures and their implications for AI processing.
  • Insight into the upcoming trends in GPU capabilities, including the latest developments in pre-filling and decoding processes in AI tasks.
  1. Market Forecasts
  2. Growth Projections:
  3. Predictions on increasing demand for AI infrastructure and GPU utilization.
  4. The evolving role of hyperscalers and potential shifts in market dynamics as new players emerge.

---

Key Takeaways

  • The collaboration between NVIDIA and Intel marks a significant turning point in the semiconductor industry, potentially reshaping the competitive landscape.
  • The geopolitical climate, particularly the relationship between the US and China, plays a crucial role in determining the future of tech advancements and AI infrastructure.
  • Emerging technologies, such as new chip architectures (GB200, CPX), are expected to drive further innovations in AI processing and efficiency.
  • The conversation indicates a shifting paradigm in how companies approach AI chip investment and infrastructure development, especially amidst increasing demand.

---

Resources and Links

  • Follow Dylan Patel on X: [@dylan522p](https://x.com/dylan522p)
  • Follow Sarah Wang on X: [@sarahdingwang](https://x.com/sarahdingwang)
  • Follow Guido Appenzeller on X: [@appenz](https://x.com/appenz)
  • Learn more about SemiAnalysis: [SemiAnalysis](https://semianalysis.com/dylan-patel/)
  • Follow the a16z Podcast on Spotify: [Spotify Link](https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX?si=70ca2d87cf9342d9)
  • Follow the a16z Podcast on Apple Podcasts: [Apple Podcasts Link](https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711)

---

Disclaimer The content here is for informational purposes only and should not be considered as legal, business, tax, or investment advice.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00How you buy GPUs is like buying cocaine. You call up a couple people, you text a couple people, you ask, yo, how much you got? What's the price? If your two arch nemesis suddenly team up, it's the worst possible news you can have. I did not see this coming. I think it's an amazing development. Like a Warren Buffett coming into a stock. Jensen is like the Buffett effect for the semiconductor world. It's kind of poetic that everything's gone full circle and Intel's sort of crawling to NVIDIA. Today, we're talking about one of the biggest surprises in semiconductors in years. NVIDIA just put $5 billion into Intel.

0:34Two long -term rivals now teaming up on custom data centers and PC products. A deal nobody saw coming. For NVIDIA, it's the Buffett effect. For Intel, it's a lifeline. And for AMD, ARM, and the global chip race, the fallout could be massive. To break it all down, I'm joined by Dylan Patel, Chief Analyst at SemiAnalysis, Sarah Wang, general partner at A16Z, and Guido Appenzeller, partner at A16Z and former CTO of Intel's data center and AI business unit. Let's get into it. Dylan, welcome back to the podcast. Thanks for having me, yeah. It just so happens that there's some big news, just as we're having you, NVIDIA announcing a $5 billion investment in Intel and them teaming up to jointly develop custom data centers and busy products.

1:22What do you think about the collaboration? I think it's hilarious that NVIDIA can invest, it gets announced, and their investment's already up 30%. $5 billion investment, $2 billion profit already. I think it's fun because they need their customers to really have big buy -in. So when potential customers buy in and commit to certain types of products, it makes a lot of sense. And it's kind of funny in a way because in the past, there was this whole thing around how Intel was sued for being anti -competitive with their chipsets, and NVIDIA actually got a settlement from Intel way back when the graphics were separate from the GPU, and the graphics were really put on the chipset, which had all this other IO, like USB and all this stuff.

2:14So it's kind of a funny turn of events that now Intel is going to make a chiplet and package it alongside a chiplet from NVIDIA and then that's like a PC product. It's kind of poetic that everything's gone full circle and Intel's sort of crawling to NVIDIA, but actually it might just be the best device. I don't want an ARM laptop because it can't do a lot of things. An x86 laptop with NVIDIA graphics fully integrated it would be probably the best product in the market. So are you optimistic? How do you think this will go? I mean, sure. I mean, I hope. I hope, right? I'm a perpetual optimist on Intel because I have to be.

2:58I was thinking that the structure of the deal that at least a lot of the government folks and Intel were sort of trying to go for was big customers and the biggest suppliers directly give capital to Intel. But this is sort of the other way around where they're buying some of the stock, having some ownership, but they're not really diluting the other shareholders. And then the other shareholders will get diluted, slash everyone will get diluted when Intel finally does raise the capital from the capital markets. But because they've announced these deals, and they're pretty small, right? 5 billion NVIDIA, 2 billion SoftBank.

3:35U .S. government was 10. These are still relatively small. Pretty small, yeah. Yeah, on the nature of things, right? I mean, last time I think I said Intel needs like $50 billion, right? Now when they go to the capital markets, it's better. And hopefully they get another couple of these announcements. Maybe there's all sorts of speculation that Trump is involved in getting these companies to invest. NVIDIA, and now the government as well, of course. And now is Apple going to come invest, right? And also do something with Intel, or who else will come in? and that'll really boost investor confidence and they can dilute slash go get debt.

4:16Like a Warren Buffett coming into a stock. The Jensen is like the Buffett effect for the semiconductor world. Guido, you were the CTO of the Intel Data Center and AIBU. What are your thoughts? I think it's really good for customers and consumers in the short term. Having both Intel and, like specifically the laptop market, having the two collaborators is amazing. I wonder what's going to happen with any of the internal graphics or AI products at Intel. They might just push a reset and give up on that for now. They currently don't have anything competitive. There was the Gaudi effort that's more or less done.

4:52There was the internal graphics chips, which never competed really at the high end. From that perspective, it makes a lot of sense for both sides. Look, I think for Intel, they needed a breath of fresh air. They were sort of desperate, so I think it's a very good thing. I think AMD is fucked. If your two arch nemesis suddenly team up, that's the worst possible news you can have. They were already struggling. Their cards are good. Their software stack is not. They were getting very limited traction. They now have a bigger problem outside. I think ARM is a little bit screwed as well because their biggest selling point was sort of like, look, we can partner with everybody that doesn't want to partner with Intel.

5:34In a sense, they're number one. And, you know, like, NVIDIA is probably the most dangerous of the future CPU competitors, right? And so they now suddenly have access to Intel technologies and might get in that direction. It remixes the card, right? I did not see this coming. I think it's an amazing development. Yeah, it will be very interesting to see this play out. To Eric's point, PAC News Week. The other thing that we wanted to pick your brain on, since we have you here, Dylan, is the other news dropping on Huawei unveiling their kind of AI roadmap. And, you know, obviously they're hyping up the capabilities.

6:08I think you guys have been sort of ahead of the curve of trying to gauge, hey, what can the 950 supercluster actually do? But would love your thoughts on everything that's going on from the China front, right? And this is kind of coupled with DeepSeq saying their next models are going to be on domestically produced Chinese chips. The Chinese government kind of banning companies from buying the produced specifically for China NVIDIA chips. So there's just sort of a lot of dominoes falling right now in the semi -market in China, but would love your take overall and drill into some detail. Yeah, I think when you sort of zoom out to even like, let's walk from 2020, because I think it's really important to recognize how cracked Huawei is, or even just historically, they've always been really good.

6:54Sure, initially they stole Cisco source code and firmware and all this stuff, but then they rapidly passed them up as well as every other telecom company. In 2020, they released an Ascend chip and submitted it to impartial public benchmarks. And they were the first to bring seven nanometer AI chips to market. They were the first to have that, right? Now, you could still say NVIDIA was ahead, but the gap was like nothing, right? And this is when they could access the full foreign supply chain. This was when they just passed Apple to be TSMC's largest customer. They were clearly ahead of everyone on a manufacturing supply chain sort of design standpoint in a total basis.

7:42Now, of course, NVIDIA still had higher market share, but it was so nascent then, they could have really taken over the market. Quality got banned by the Trump 1 administration from accessing, and then it went into effect in 2020, the full ban. And so they were only able to make a small volume of these chips, but they had trained significant models on these chips that they made then. And then over the next couple years, NVIDIA continued to accelerate. Huawei, because they were banned from TSMC, had to go and try and figure out how to manufacture at SMIC, the domestic TSMC. And then they were also, in parallel, trying to go through shell companies to manufacture at TSMC and acquire memory from Korea and so on and so forth.

8:26so by the end of 24 this had gotten in full swing and it was caught and they finally shut it down but they were able to acquire 3 million chips 2 .9 million chips from TSMC through these other entities roughly $500 million worth of orders which ends up being a billion dollar fine that the US government gave TSMC if I recall correctly at least there was a Reuters article I don't know if they actually issued it which is important and interesting to gauge because the number of ascends floating out there has not consumed this entire capacity yet. So now we get to 2025, the H20 got banned in the beginning of the year, NVIDIA had to write off huge amounts of money, our revenue estimate for NVIDIA in China for just H20 was north of $20 billion because that's what they were booking in capacity slash had to write off.

9:23and then it got banned, they cut the supply chain, they just said, no, we're not doing this anymore. They had their inventory, it gets re -approved, they resell the inventory, but now they're like, do we even restart production? Is NVIDIA's question? And now you have China saying, hey, we don't need NVIDIA, we have domestic alternatives, whether it be Huawei or CameraCon. These companies have capacity, but most of this capacity is still foreign -produced, whether it be wafers from TSMC, memory from Korea, Samsung and SK Hynix. So the question is how much can they do domestically? And there's sort of two fronts there.

10:05There's the logic, i .e. replacing TSMC, and there's the memory, i .e. replacing Hynix, Samsung, Micron. And on the logic side, they're behind, but they're really ramping there, and I think they can get to the production capacity estimates needed, and the US is still allowing them to import all the equipment necessary pretty much. The bans are really for beyond the current generation of technology, beyond 7 nanometer. The bans are really for 5 nanometer and below. Even though the government says they're for 14 nanometer, the actual equipment that's banned is only for below 7 nanometer. And so they'll be able to make a lot of 7 nanometer AI chips and maybe even get to 5 with using existing equipment for 5 nanometer rather than taking the new techniques.

10:52And so there's the logic side and then there's the memory side. And the aspect of Huawei's announcement that was surprising was that they're doing custom memory, right? That's the part that is sort of like, hey, this is really exciting. They announced two different types of chips for next year. One that's focused on recommendation systems and pre -fill and then one that's focused on decode. There's a twin these days. Yeah, so in NVIDIA, the same thing. They just announced a pre -fill specific chip recently. There's numerous AI hardware startups that are really focusing on pre -fill versus decode.

11:27And so the sort of split of inference up to two workloads, Huawei's doing the same thing for their next year chip. And what's interesting is the decode one has custom HBM. What does that mean? What is the manufacturing supply chain? Because that's the one that's tricky, right? How much can they manufacture of that custom HBM? and NVIDIA and others are also adopting custom HBM only starting next year. So it's not like, yes, the manufacturing capacity is not there, maybe it is going to consume a bit more power, it's going to be slightly lower bandwidth, but the fact that they're able to do some of the same things that NVIDIA plans to do, AMD plans to do, in their memory is evidence that they're catching up.

12:08But then the main question that remains is production capacity. So as far as like, hey, NVIDIA is banned in China, like they're saying don't buy NVIDIA chips I think for a period of time that's fine for China from a perspective of hey I'm China that's fine because you have all this capacity that you shipped in in 2024 they haven't turned into AI chips now you're turning them into AI chips you're running all that stockpile down what about the transition from running that stockpile down to ramping your new stuff and that transition is the one that's really tricky China's either shooting itself in the foot by not purchasing NVIDIA chips during that time period or China's able to ramp.

12:49I think they'll be able to ramp. I think it'll take a little bit longer and there will be a gap in between where China probably backtracks and says it's fine. ByteDance is begging for NVIDIA chips. They don't want to use... They use some CameraCon, they use some Huawei, but they really want to use NVIDIA because it's way better. They don't care about the domestic supply chain. They want to make the best models. They want to deploy their AI as efficiently as possible. So the government can mandate them to not do it. So it's not that NVIDIA is not competitive. It's that the government is trying to instigate it.

13:30I guess the last sort of thing is there's always the argument of hey, if banning NVIDIA chips to China is so good for China, why didn't China do it for itself? And they're finally doing it for themselves. Again, it'll be interesting to see. Smuggling is still happening, re -exportation of chips from other countries to China. That is still happening at some volume, low volume. Lower medium volume, right? But then the direct shipments of NVIDIA chips that are legally allowed to China are not necessarily happening today, but may have to restart at some point, because China won't have the production capacity.

14:12They would just have so many fewer AI chips being deployed domestically versus the US. And at some point you kind of have to pick, am I all about the internal supply chain or am I all about chasing super powerful AI? Is there an angle here about the negotiation angle as well? Because currently there's still discussions ongoing what exactly are the boundaries, what can be exported to China. So these are sort of well -timed announcements. if you want to make a point that the US should allow more exports. Do you think that's a factor or not? In the report we did a few weeks ago about the production capacity of Huawei and the supply chain, there was a bit in there that we wrote about how, honestly, if you were China and you do want NVIDIA chips, actually, how do you play this?

15:02It's by hyping up your domestic supply chain. And it's like, yes, we can do everything. It's Huawei announced the most crazy shit possible, announced three years of roadmaps. I think they knew. They were already banned, and then say, we're banning NVIDIA. Then the government official is going to think, alongside lobbying from domestic players, of course we want to ship them better AI chips. We're losing this market. We can't lose this market. And it's sort of like, it is 10 ,000 IQ, right? And we're here playing checkers while they're playing chess. Well, so I guess negotiating chip aside, in that report you talked about HBM or high bandwidth memory being a bottleneck to Huawei.

15:50To your point on one of the surprising aspects of the announcement, do you think it's credible that it's no longer a bottleneck based on what they're saying? Or is it just hype? I think production capacity -wise, it is still absolutely a bottleneck. Like they, certain types of equipment required for making HBM need to be imported. They're working on domestic solutions, but as far as we know, they have not imported enough equipment for this. Although if you look at Chinese import data for different types of equipment, right? There's sort of like fabs spend, you know, roughly, it depends on the process technology, but fabs spend roughly different amounts of money on lithography, etch, deposition, metrology, right?

16:28Like these different steps. and historically lithography has hovered around you know 17 -18 % with EUV it grew to 25 % right but China because they they sort of like wanted to stockpile lithography and they were worried about the incoming ban they were importing lithography at a much higher rate than that right like 30 -40 % of their equipment imports were lithography and they were just stockpiling lithography equipment this is sort of like reversed now in that like hey If you look at the monthly import -export data both into provinces in China but also out of countries, you can see that Etch specifically is skyrocketing.

17:09The main thing about stacking HBM is that when you have each wafer, you have to Etch create this thing called a through -silicon via so it can connect from the top to bottom and then you stack them on top of each other, 12 high, 16 high for HBM. That's how you make super high bandwidth memory. Their import for Etch is skyrocketing now. So it's like, they don't have the production capacity yet. How fast can they ramp it is a function of how much equipment can they get, A, and B, the yields, right? Improving yields is really hard on manufacturing. Intel and Samsung are really good, and TSMC is just amazing.

17:43Not that those companies suck, I think is a better way to put it. And so, you know, it's those two things. I think yield, they haven't even started production of high speed, of HBM3, right? They've only done some sampling of HBM2. HBM 3 came out a few years ago. So there's still quite a bit of ways on going up the learning curve. Obviously, I expect them to catch up faster than it took the technology to be developed because it exists in the world. We know how to do it. It's just a matter of actually doing it versus inventing it. And then the other one is the production capacity. A couple months of import -export data is not enough to set up for years worth of supply chain buildup, which is what we have today in Korea for the Korean companies.

18:30Now, Hynex is also investing in the US, in Illinois, and then Microns primarily in Japan. The American memory company is primarily in Japan and Taiwan, but they're also expanding in Singapore and the US now. There's so much capital that's been invested. It would take some time for China to build up that production capacity to actually match the West, and when I say the West, I mean East Asia in production, non -China East Asia in production capacity. So it'll take some time to get there. I think it's like, hey, we can design this. It's always a question of can we manufacture? And then the thing that Jensen would say is you're betting on China not being able to manufacture?

19:10It's a matter of when, not if. And that's the whole calculus that I think the US government has to be aware of when they're like, hey, what level of AI chips do we sell? Do we sell everything? Probably not because AI is far more powerful and the end market of AI is going to be way larger than the end market of semiconductors and equipment. Do we sell, you know, what level do we sell at? Well, how much can China make at each specific, you know, sort of performance tier and then, you know, analyze that and what's the volume and then figure out, like, what is okay, which is, like, maybe a little bit above or around the same level.

19:42Yeah. So if you, to your point on, like, playing chess versus checkers, if you're Jensen, what would your next move be given the situation at hand? It's both partially true that he's afraid of Huawei more than he is an AMD. Right. He called them formidable. Yeah. I mean, Huawei's beat Apple, right? They passed Apple up in TSMC orders. They passed Apple up in phone market share. Not in the US, but in many parts of the world before the bans came down. And then even now they're growing back again in market share without Western supply chains. they've done this to numerous other industries. I would say Apple's a formidable competitor.

20:26They've beaten a lot of industries. It's reasonable that he's afraid of them. He's not afraid of A and B. I think the best thing is

20:41what Huawei announced is reality rather than their hope target. and so away all doubt on manufacturing capacity, which I think is not fair. I think manufacturing capacity is a real bottleneck for them. And then the yield learning is a real bottleneck. Temporary, maybe. We'll see how long and we'll see how fast the rest of the NVIDIA technology advances past what Huawei is capable of and how fast Huawei is able to close the gap. But I think his main sort of pitch would be Huawei is real. They're a formidable competitor. They're going to take over not just the Chinese market, but also foreign markets, right?

21:25Whether it be the Middle East or Southeast Asia or South Asia or Europe or LATAM, right? Everywhere besides America. And there's sort of, there's a, I think Noah Smith has this analogy, right? This whole idea is that you should Galapagos China, make them have their own domestic industry that is so different from the rest of the world. Kind of what happened with Japan in the 70s and 80s and 90s, their PCs were so specific and hyper -optimized to the Japanese market with the weird, I don't know if you've seen the weird scroll wheel on these Japanese PCs. You literally, it's like you go like this and it scrolls.

22:07And then the touchpad is a circle and then that's around it. It's like, things like that are so weird. And the rest of the world doesn't care, but Japan market likes it, right? And his whole idea is like, let's Galapagos them, i .e. keep their technology within China and then that's like deadweight loss and they never expand outside versus that we serve the whole world. But the whole risk is that the opposite can also happen, right? Our technology is hyper -optimized to running, you know, language models at this scale and RL and you keep like hardware software co -design can take you down a path of the tree that is a dead end.

22:43And then China, because they're not allowed to access this tree, they're like, oh, okay. And then they end up in the optimal spot. We had a local minima, they had a local maxima, they had a global maxima. That's sort of technological Galapagosing is sort of what Noah Smith's analogy is. I like it a lot. I don't know if it's accurate, but it's an interesting one. Yeah, I love that. Well, actually, maybe just taking a step back from current events, even though there's so much to talk about right now. Last time you appeared with us, NVIDIA came up, obviously. And you talked about a couple of the potential paths forward for NVIDIA.

23:19Give us maybe the bull case, the bear case. Yeah, fair enough. There's a lot embedded in their numbers now. but what's interesting is consensus for the banks is like for across the hyperscalers so Microsoft, CoreWeave, Amazon, Google, and Oracle right meta so it's the six hyperscalers the consensus for the banks is $360 billion of spend next year across all of them and my number is closer to it's like $450 ,000, $500 ,000. And that's based on all the research we do on data centers and tracking each individual data center in the supply chains. So this is just NVIDIA spend. This is CapEx for the hyperscalers.

24:09And then CapEx gets split up across different companies, but the vast, vast majority still goes to NVIDIA. And NVIDIA is in a position not where they can't take share. They grow with the market slash defend share. And so the question is, how fast is the growth rate of CapEx for hyperscalers and other users? And the reason I included Oracle and CoreWeb is hyperscalers, even though they're traditionally not called hyperscalers, is because they are opening eyes hyperscaler. So when you look at the Oracle announcement, first of all, the Oracle announcement, I don't understand why people don't think this is crazier.

24:48They did the most unprecedented thing in the history of stocks and companies ever. They gave a four -year guidance and it made Larry the richest man in the world. The question is, how fast does revenue grow? Do you think OpenAI, which signed a $300 billion plus deal with Oracle, will actually be able to pay $300 billion across raising capital and revenue? and it gets to a rate of over $90 billion a year in just a handful of years. Do you believe the market will grow that fast? It's very possible, yes. It's very possible for OpenAI, what is their revenue going to be exiting next year? Some people think $35 billion, some people think $40 billion, some people think $45 billion.

25:44ARR by the end of the year next year. This year they hit 20. If that growth rate is maintained, then all of that cost goes to compute plus all the capital they continue to raise. Again, their financials that they gave to investors for their last round was like, hey, we're going to burn $15 billion next year. It's probably more likely going to be $20 billion. You stack this on and they're not going to be profitable until 2029. so you sort of have like they're going to continue to burn 15, 20, 25 billion dollars of cash each year plus revenue growth, that's their compute spend and you do this for Anthropik, you do this for OpenAO you do this for all the labs it's very possible that the pie does get to you know more than 500, not 360 billion next year, 500 billion next year for total capex and the pie continues to grow for hyperscalers Nvidia says actually it's going to be multiple trillions a year on AI infrastructure and he's going to capture a huge portion of it.

26:43That's his bull case, right? That's the bull case is AI is actually so transformative and the world just gets covered in data centers and the majority of your interactions are with AI, whether it's business productivity and telling an agent to do some code or you're just talking to your AI girlfriend, Annie, it doesn't matter. All of this is running on NVIDIA for the most part. The bear case is, even if it does grow a lot... Save the bull case for a second. Fundamentally, the value creation, I think, personally, is there. Trillions of dollars of value with AI, I can totally see this happen. So assume it's true, where will NVIDIA top out?

27:22I guess, how much do you believe in takeoffs, right? Yes. If there is a takeoff scenario where powerful AI builds more powerful AI builds more powerful AI, or that creates more and more, each level of intelligence enables more for the economy. How many monkeys can you employ in your business versus how many humans? Or how many dogs? What is the value creation of a human versus a dog? It's the same with AI. In this case, the value creation could be hundreds of trillions, if not the negative. If we take every white -collar worker and make them twice as productive with AI, that's in the hundreds of trillions isn't it yeah but like what is twice you know like I mean like if you talk to people at the labs right like twice as productive what does that even mean it's replaced them right and it's been 10 times better than death like I mean like I don't know how soon that happens white color work is essentially useless without a constant stream of LLM tokens right that make them productive right at that point you basically can tax every single knowledge worker in the world right which is most workers in the world long term yeah I don't know.

Read the full transcript

28:36What's your guess? Give us a number. What's the cap? Cap? I mean, why aren't we making a Matrioshka brain? I don't know. I mean, at some point, the machine says humans don't need to live and I need even more compute. What's that before that, man? Are we colonizing Mars yet? TBD. I don't know, man. I find it completely impossible to predict anything beyond five years given how much stuff is changing. Five is a large number. I'll leave it to Economist. Supply chain stuff is three, four years out and that's it. And then fifth year is sort of like yellow. I just try and ground myself in the supply chain stuff.

29:24Supply chain and then what is the adoption of AI? What's the value creation? What's the usage? And you can see that in a short horizon. beyond that like I don't know like are we all going to be connected to computers like BCIs and stuff like I don't know dude are humanoid robots are they going to be you know I mean you saw Elon's thing right like he's like yeah humanoid robots are why Tesla's worth more than 10 trillion so go ahead great what is all that being trained on great NVIDIA okay awesome so that's worth also 10 trillion right like I don't know like it's too out there for me I don't like the out there discussions read some sci -fi books So just pulling out the thread where you talked about, I mean, this is kind of a throwaway comment, but how market share can't really grow just because it's such a dominant market share.

30:12And we talked about or you guys talked about the moat of Nvidia last time. And obviously this moat is tied to maintaining that very high market share that they currently have. And I love this sort of historic journey you took us through with Huawei just earlier. can you kind of walk through what Nvidia did throughout history to build their moat? It's super awesome because they failed multiple times in the beginning and they bet the whole company multiple times Jens is just crazy enough to bet the whole company whether it was certain chips ordering volume before he knew it even worked and it was all the money he had left or ordering volumes for projects he had not won yet.

30:58I heard a rumor that or not a rumor, but a story from someone who's a gray beard in the industry and I think would know. NVIDIA ordered the volume for the Xbox before Microsoft gave them the order. They were just like, fuck it, YOLO. I don't know. I don't know how true this. I'm sure there's more nuance there, like verbal indication or whatever, but the order was placed before he got the order. is what he said. There's cases like with the crypto bubbles. There was a couple of them, but NVIDIA did their damn best to convince everyone on the supply chain that it wasn't crypto and that it was gaming, that it was durable, real demand, and it was gaming and data center and professional visualization.

31:46Therefore, you guys should ramp your production. They all ramped production and spent all this capex on increasing production and building out new lines for them. They pay per item and then they bought them and sold them and made shitloads of money. And then when it all fell apart, they just had to write down a quarter's worth of inventory. Whatever. Everyone else was like, well, crap, I have all these empty production lines. And so it's like, but what did AMD do then? Their chips were actually better for crypto mining on an amount of silicon cost versus how much you hash. AMD was like, we're going to not really raise production.

32:23as a reasonable thing. It's a sort of like strike while the iron's hot. The same has happened with NVIDIA. In recent times, they've ordered capacity that no one believes multiple times. They see that in demand, obviously, but in many cases, their number for Microsoft was higher than Microsoft's internal planning. and then Microsoft's internal planning went up but their number for Microsoft was way higher and it's like, we just don't think Microsoft's going to need this much even though they tell us this it's like, who the heck is like, no, no, no, customer you're going to buy more and orders, right?

33:07and then when the orders come through the supply chain it's like, I have to put pay NCNR, non -cancelled, non -returnable

33:15I asked a question in Taiwan once it was Colette, which is the CFO, and Jensen CEO. They were both there. It was a room full of mostly finance bros and they were asking stupid finance questions three days before earnings. So obviously they just could not answer anything because it's like SEC regulations. But then my question to them was like, look, Jensen, you're so vibes driven and very gut feel and very visionary. And then Colette's CFO, she's amazing in her own right, but those personalities clash, how do you work together? He's like, I hate spreadsheets. I don't look at them. I just know. It's his response.

33:56And it's like, of course the best innovators in the world have really good gut instinct. Right? And so the gut instinct to order with non -cancelable what you don't know, and they've had to write down over their history multiple times, many, many billions of dollars in accumulative orders. accumulate in total orders. Whether it be the H20, which is more regulatory, but other cases they've ordered and had to cancel. Is that many billions? It's many billions. Peanuts. Well, it depends, right? The crypto write -down was like multiple billion when their stock was less than 100 billion. It's peanuts compared to the upside.

34:36I think everything you did was right. I think everything AMD did was wrong in that scenario. but it is crazy to, especially in a cyclical industry like semiconductors where companies go bankrupt all the time, which is why we have all this consolidation, is every down cycle, companies go bankrupt. If you look from a risk -return perspective, these bets were totally worth taking. If you look at it from, I'm a CEO, I want to have predictable quarters for Wall Street, it's a very different story. I think that's where part of detention is from. though. Yeah, so we I don't know if you've seen these Lee Kuan Yew edits where they're like him saying some fiery speech and then it's some cool music at the end and it's showing different pictures of him.

35:21And so we made one of Jensen recently and put it on social media, right? On Instagram, TikTok, XHS, Redbook, Twitter, of course, all the different social media. And I really liked it because he's like the goal of playing is to win. And the reason you win is so you can play again. And you compared it to pinball where actually you just play all day and you keep getting more rounds. And it's like, his whole thing is like, I want to win so I can play the next game. And it's only about the next generation. It's only about now, next generation. It's not about 15 years from now because it's a whole new playing field every time or five years from now.

36:02You're right, the risk -reward is correct. There's few people take these kind of risks. It's the only semiconductor company that's worth, I think, even north of $10 billion that was founded as late as it was. MediaTek was in the early 90s and then NVIDIA and everyone else is from the 70s mostly. From the big ones. Yeah. Yeah, I think you raised this great point on the bet the farm and he's actually been wrong a couple times to your point. Mobile, right? What the hell happened with mobile? Exactly. And he still takes them. And I think Mark actually had this great conversation with Eric where he talked about being founder run, where you have this memory of the risks you took to get to where you are today.

36:46And so in a lot of cases, if you're a CEO brought on later on, you're sort of like, okay, continue to steer the ship as is. But in this case, he remembers all the times they almost went belly up. And he's like, I've got to keep making bets like that. how do you think he's changed I mean he's been one of the longest running CEOs he's kind of right up there with Larry Ellison now how do you think he's changed over the last 30 years or so I mean obviously I'm 29 I don't know what he was like I've watched a lot of old interviews I won't say he wasn't he's been CEO longer than you've been alive yeah exactly Nvidia was founded before I was born I'm 96 right Yeah, maybe anything over the last couple of years?

37:35No, no, it's probably better. I think even like watching old interviews, right? Like I watched a lot of old interviews, a lot of old like presentations he's given. One thing is that he's just like sauced up and dripped up like way, like the charisma he's gotten has only gotten stronger, right? Yeah. Which is an interesting point. I don't know if it's quite relevant. I don't agree with that, yeah. But like the man like has learned to be a rock star more, even though he was always charismatic. It was like, he's a complete rock star now. And he was a rock star a decade ago too. It's just people maybe didn't recognize it.

38:10I think the first live presentation that I watched, it was extreme, was like, it was CES, like 2014 or 2015 or whatever.

38:24It's a consumer electronics show. I'm like moderating gaming hardware subreddits at the time I'm a teenager and the dude is like talking only about AI, he's telling all these gamers about AlexNet and self -driving cars it's like know your audience first of all but also like it has nothing to do with consumer electronics at gaming at the time I was also like I was half like holy crap this is amazing but also half like, I want you to announce new gaming GPU. But I know on the forums, quickly everyone was like, screw this, I want to hear about the gaming GPUs, NVIDIA's price gouging. Of course, NVIDIA's always had to price the value plus a little bit because we were just smart enough to know.

39:18I'm guessing Jensen just has the gut feel of how to price things. He'll change the price, at least on gaming watches, he'll change the price up until like right before the presentation so like it really is like a gut feel thing probably um and anyway so so he he had that charisma to know what was right but i think people a lot of people were like oh no whatever jensen's wrong he doesn't know what he's talking about but now like he he talks people like oh very very uh you know so it might just be that he's been right enough yeah there's a post on x recently that said he had moved up into god mode with a select group of CEOs, but that this was, like, it's exactly.

39:57Who's the other gods? It was Zuck. Who are the other gods? Elon. Elon. Elon, Zuck, and Jensen. Nice, nice. Good crew to be in. So we prayed a Silicon Valley. It's sort of the cult now, is it? Basically. Just one last thing on people. You mentioned Colette, his CFO, and, you know, there's sort of a famously loyal crew at NVIDIA, even though all of the OGs could retire at this point. Is there anyone akin to a Gwynne Shotwell at SpaceX or previously a Tim Cook to Steve Jobs at Apple that is at NVIDIA today? I mean, he had two co -founders, right? Like, that's, you know, let's not overlook that. One of them's, like, you know, not involved and hasn't been for a long time, but the other one was involved up until just a few years ago, right?

40:48So it's not just Jensen running the show, right? Totally. Although he was running the show. There's quite a few people on the hardware side. I've always, there's someone at NVIDIA that's like mythical to me. Like when you talk to the engineering teams, he leads a lot of engineering teams. He's a private person, so I don't want to say his name actually. Fair enough. But, you know, he's like effectively like chief engineering officer is like his role. and people within his org will know who he is. I think there are people like that, but he's intensely loyal and there's a number of these types of people.

41:31There's another fella who's like, there's all these innovative ideas at NVIDIA and he's the guy who literally is like, we need to get this silicon out now, we're cutting features. And that's what he's famously known for and all the technologists at NVIDIA hate him. This is like a second guy, this is a second guy. Also intensely loyal to NVIDIA, has been around for a long time, but it's sort of like when you have such a visionary company and forward, one problem is that you get lost in the sauce. Oh, I want to make this. It's got to be perfect, amazing. You got to have that sort of like... Obviously, they're close to Jensen for a reason because Jensen also believes these things.

42:12Have the visionary future looking, but also screw it, cut it, we'll put in the next one, ship. Ship now, ship faster. In a space like Silicon, which is really hard to do so.

42:28The thing about NVIDIA that's always been super impressive, and it's from the beginning days where he's talked about this before, is their first successful chip, they were going to run out of money. He had to go get money from other people to even finish the development. Even then, he just had enough money because he'd already had a failed chip before this. was the chip came back and it had to work. Otherwise, it would not. And so they were like, because they could only pay for, it's called a mask set, right? Basically, you put these, I'll call them stencils, into the lithography tool. And then it says where the patterns are.

43:02And you put the stencil in, you deposit stuff, you etch stuff. You deposit materials on the wafer, etch it away. And you put the stencil in and you tell it where to put stuff. And then the deposition and etch keeps happening in those spots. and you stack dozens of layers on top of each other and you make it with chip. These stencils are custom to each chip, right? And they cost today in the orders of tens and tens of billions of dollars. But even back then, it was still a lot of money. It wasn't that much then, of course. You know, it sort of, they could only pay for one set. But the typical thing with semiconductor manufacturing is, you know, as good as you can simulate, as good as you can do all the verification, you'll send a design in and you have to change it.

43:46There's going to be something. It's so hard to simulate everything perfectly. And the thing about NVIDIA is they tend to just get it, right, the first time. Even great executing companies like AMD or Broadcom or whoever, they often have to ship, they're denoted in A and then a number or B and then a number, so it's like two different parts of the masks. So NVIDIA always ships A0, almost always. They sometimes ship A1. And a lot of times, even if they'll start production of the, you know, the A is basically the transistor layer, then the number is like the wiring that connects all the transistors together.

44:23So NVIDIA will start production of the A and ramp it really high and then just hold it right before you transition to the metal just in case they do need to change the metal layers. And so like the moment they're ready and they've confirmed that it works, they can just, you know, blast through a lot of production. Whereas everyone else is like, oh, let's get the chip back. Oh, okay, A0 doesn't work. we've got to make this tweak, make this tweak, and then shift back. It's called a stepping, right? At Intel, we were very jealous of NVIDIA at that time. They consistently delivered in the first one.

44:50We did not.

44:54The data center CPU group, there was one product where I said A1, A0, A1, or you go to B if you have to change the transistor layer as well. So it's like B. Intel got to like E2 once. E2, that's like a 15 revision. this is the peak of AMD's when they went skyrocketing on market share versus Intel was when Intel was at E2. 15 steppings. It's catastrophic for a go -to -market. Yeah, each time is a quarter of delay or something. It's absurd. I think that's the other thing about NVIDIA is screw it, let's ship it, let's get the volume, ASAP,

45:39let's things that... And so anyways, they have some of the best simulation, verification, etc. that lets them go from design, from idea, to shipment as fast as possible, cutting out any necessary features that could delay it, making sure they don't have to do revisions so that they can get... They can respond to the market ASAP. There's a story about how Volta, which was the first NVIDIA chip with Tensor Cores, they saw all the AI stuff on the prior generation P100 Pascal and they decided we should go all in on AI and they added the tensor cores to Volta only a handful of months before they sent it to the fab.

46:26They said, screw it, let's change it. And it's like if they hadn't done that, maybe someone else would have taken the AI chip market. So there's all these times and those are major changes but there's often minor things that you have to tweak. number four maths or some architectural detail. NVIDIA is just so fast. The other crazy thing is they have a software division that can keep up with that. If you come out with a chip and basically no stepping required, it's immediately in the market, then being ready with drivers and all the infrastructure on top, that's just super impressive. I love that point because you think of NVIDIA benefiting from tailwind after tailwind, but I think both of you are saying you have to move fast enough and execute well enough So take advantage of those tailwinds.

47:11And if you think about, and by the way, I loved your CES story. I'm just envisioning him more than 10 years ago talking about self -driving cars. But, you know, if you think about nailing the video game tailwind, VR, Bitcoin mining, obviously AI now. You know, one thing that, or one of the things that Jensen talks about today is robotics, AI factories. Maybe my last question on NVIDIA, what do you think about the next 10 to 15 years? I know calling beyond five is hard. But what does NVIDIA's business look like?

47:47It's really a question of, and this is like, I think every time I've talked to some executives at NVIDIA have asked this question because I really want to know and they won't answer it, obviously. But it's like, what are you going to do with your balance sheet? You are the most high cash flow company and you have so much cash flow. Now the hyperscalers are all taking their cash flow way down because they're spending on GPUs. What are you going to do with all this cash flow? Even before this whole takeoff, he wasn't allowed to buy ARM. So what can he do with all this capital and all this cash? Even this $5 billion investment in Intel, there's regulatory scrutiny there.

48:38It's in the announcement, yeah, this is subject to review. I imagine that'll get passed, but he can't buy anything big. He's going to have hundreds of billions of dollars of cash on his balance sheet. What do you do? Is it start to build AI infrastructure in data centers? Maybe. But why would you do that if you can just get other people to do it and just take the cash? Well, he's investing those, right? Investing peanuts. He gave recently a core weave a backstop because today it's really hard to find a large number of GPUs for burst capacity. Like, hey, I want to train a model for three months.

49:18I have my base capacity where I don't know my experiments, but I want to train a big model for three months. We know from our portfolio. So NVIDIA sees this issue. They think it's a real problem with startups. It's why the labs have such an advantage. But what if I could? And right now, most companies in the Valley spend, what, 75 % of their round on GPUs? At least, yeah. What if you could do 75 % in three months on one model run, right? And really scale and have some sort of competitive product and then you have the model, then you raise more capital and start deploying. What do you do with it?

49:53Is it start buying a crap load of humanoid robots and deploying them, but they don't really make good software. They don't make really that amazing software for them in terms of the models. The layer below is great. Where they deploy their capitals is a question. He has been investing up and down the supply chain a little bit though, right? Investing in the NeoClouds, investing in some of the model training companies. Yeah, but again, small fries. He could have just done the entire Anthropic round if he wanted to. Of course he didn't, right? And then really got them to use GPUs. Or he could have done the entire OpenAI round.

50:28Are you going to have done the entire XAI round? Do you think these are things he should be doing? Yeah, good question. I don't know, right? We'll quote you for the next round. But anyways.

50:43He could make venture a dead industry. Take all of the best rounds. That is a lot of business, yeah. You can do the scenes and then have Jensen mark you up. That's why I'm going to work. No, I don't think. I don't think it. I think picking winners is obviously really tough for him because he has customers all across this ecosystem. If he starts picking winners, then his customers will be even more anxious to leave and give even more effort to whether it's AMD or some startup or their internal efforts, et cetera, et cetera, buying TPUs, whatever it is. He can't just invest in these. He can do a little bit.

51:21A few hundred million in an open AI round is fine or a few hundred million in the next AI round is fine. core weave right like yeah everyone's like throwing a fuss about it but it's like he invested a couple hundred million plus you know early on plus you know rented a cluster from them for internal development purposes instead of renting it from a hyperscaler which is cheaper for nvidia to do right it's better for them to do it from them than the hyperscalers that's like did he really like is he really backstopping core weave that much right or you know any of the other customers are neoclouds.

51:55There's some investment, but it's more like, this is a good cloud, we'll throw 5 % or 10 % of the round. It's not he's taking 50 % plus of the round. Is he also reshaping his market? I mean, look, a couple of years ago, there were four big purchases of these cards. You just listed six. To what extent is that... There's a long list there. Is that a strategy? It is. I think it absolutely is. But he didn't have to put much capital down to do this. Just chip one earlier than the other? I don't know. No, but if you look at the grand amount of capital that he spent investing in the NeoClouds, it's a few billion dollars.

52:38But he has a lot of other levers if he wants to. Right, right. Allocations, as you mentioned. What's nice is historically, you gave volume discounts to hyperscalers, but because he can use the argument of antitrust, he's like, everyone gets the same price. So fair. It's very fair. What should he do? What should guide his... I think there's the argument he should invest in data centers and only the data center layer, not what goes in the data center so that more people build data centers and then if the market demand continues to grow up, data centers in power are not the issue. Invest in data centers in power.

53:17I've said that to them, they should invest in data centers in power, not in the cloud layer because the cloud layer is quite commoditized, but it's commoditized or complement, right, is the whole phrase. And I won't say being a cloud is commoditized, but it's certainly like you have a lot of competitors who are decent now. And you've educated the commercial real estate and other infrastructure investment firms into going into AI infra as well. So I don't think it's the cloud layer that you invest in, right? Do you invest in data centers and energy? Yeah. do you invest it because that's the bottleneck for your growth really is A how much people want to spend and can spend and B the ability to actually put them in data centers and then like robotics I think there's areas he could invest in but nothing requires $300 billion capital so what do you do with the capital?

54:09I really don't know and I feel like Jensen has to have some idea there's some visionary plan here because that's what shapes the company, right? I mean, they could sell cash. They could just continue to, you know, I mentioned $200 billion of free cash flow, $250 billion of free cash flow a year. What do they do with it? Like, do they just buy back stock forever? Like, do they go Apple route? And the reason why Apple hasn't done anything interesting in like, you know, nearly a decade is, you know, they've got a not visionary at the head. Tim Cook's greatest supply chain. And they're just plowing the money into buybacks.

54:43They're not really, you know, automotive, the self -driving car thing failed. we'll see what happens with AR, VR. We'll see what happens with wearables, but meta and OpenAI might be even better than them. We'll see in others. What does he invest in? I have no clue. What requires so much capital is the tough question. It actually gets a return. Because the easy thing is my cost of equity. I just buy back. It doesn't completely change the company culture. I think that's another thing. There are probably areas you could invest it in, but you suddenly end up with the company doing two completely different things which are very difficult to keep on.

55:18But they do like 10 completely different things, right? I mean, one way to look at it is we build AI infrastructure. And in the guise of we build AI infrastructure, robots, humanoids around the world are AI infrastructure. Or data centers and energy is AI infrastructure, right? So the humanoids would totally work, right? If you're suddenly pouring concrete and building power plants, it has completely different culture, completely different set of people. There's different ways to do it, like invest in the various companies or backstop the building of power plants. There's no one who wants to build power plants because they're 30 -year underwriting things.

55:55There's all these different areas where could it use capital to allow something to happen, not necessarily owning it himself. Look and bear in mind how it's Intel. One of the biggest problems we had was that our customer base sucked. We were selling to, most of the chips went to the large hyperscalers, which they're way too concentrated, and they build their own chips, and so you can push down your prices. So honestly, spending it on diversifying the cloud. Well, the problem was in 2014, you guys should have just charged so much that your margins were 80%. What would the world have done? Nothing.

56:33The margins were pretty good back then. That wasn't the problem. That was the primary problem. They were 60, 65. They were 80. They were 80. Still, yeah. Boy. That was Jensen. It's Jensen on a server podcast. PTSD is kicking in here. Well, wait. I think Guido's comment is actually a really good segue into something else we wanted to talk to you about, which is the hyperscalers. And one of the reasons that I love reading semi -analysis is you guys make these out -of -consensus calls that you're often right about. And one of them recently was calling. Only often? You have a Jensen hit rate. It's very high.

57:14Where's my billion dollar, you know, PV positive bet? But the one that caught my eye was Amazon's AI resurgence. So I wanted to talk to you a little bit about that just because, you know, I think we found it pretty interesting being on the ground, helping our portfolio companies pick who their partners are. And so we have some micro data on this, but you sort of walk through why they're behind. Yeah. So in Q1 2023, I wrote an article called Amazon's Cloud Crisis. And it was about all these neoclouds are going to commoditize Amazon. It was about how Amazon's entire infrastructure was really good for the last era of computing.

58:00What they do with their elastic fabric, ENA and EFA, their NICs, the whole protocol and everything behind them, what they do for custom CPUs, etc. It was really good for the last era of scale -out computing and not this era of scale -up AI infra. And how neoclows were going to commoditize them and how their silicon teams were focused on cost optimization, whereas the name of the game today is max performance per cost. But that often means you just drive up performance like crazy. Even if cost doubles, you drive up performance more, triples, because then the cost per performance falls still. That's sort of the name of the game today with NVIDIA's hardware.

58:43And it ended up being a really good call. Everyone was calling us out like, no, you're wrong. And this was when Amazon was the best stock, and Microsoft really hadn't started taking off yet, and nor had all these other, Oracle and so on and so forth. And since then, Amazon has been the worst -performing hyperscaler. And the call here is that they still have structural issues. They still use Elastic Fabric, although that's getting better, still behind NVIDIA's networking, still behind Broadcom's slash Arista, like type networking, Nix. Their internal AI chip is okay, but the main thing is that they're now waking up and being able to actually capture business.

59:30So the main call here is that since that report, AWS has been decelerating revenue. Year -on -year revenue has been falling consistently. And our big call is that it's actually going to start re -accelerating. And that's because of Anthropic, it's because of all the work we do on data centers, tracking every single data center, when that goes online and what's in there, the flow -through on cost, how much the chips cost, the networking cost, the power cost. you know how much generally margins are for these things and you can start estimating revenue. So when we build all that up, it's very clear to us that they trough on AWS revenue growth this point.

1:00:10This is the lowest AWS revenue growth will be on a year -to -year basis for at least the next year. And it's re -accelerating to north of 20 % again because of all these massive data centers they have online with Tranium and GPUs, it depends on which one, it depends on which customer. The experience is not as good as, say, a CoreWeave or whatever, but the name of the game is capacity today. CoreWeave can only deploy so much, they only can get so much data center capacity, and they're really fast at building. But the company with the most data center capacity in the world, that and still today, although they may get passed up in the next two years, is Amazon.

1:00:54Actually, they will get passed up based on what we see is Amazon, but incrementally, Amazon still has the most spare data center capacity that's going to ramp into AI revenue over the next year. Let me ask a question. Is that the right type of data center capacity? Like for the high -density AI build -outs today, you need massively more cooling, you need to have enough water close by, you need to have enough power close by. Is it the right place, or is that the wrong type of thing? Data center capacity, in this sense, I mean, all the way from power is secured to substations built to transformers to you can provide the power whips to the racks.

1:01:29Now, obviously, the data center capacity will differ, right? Historically, actually, Amazon's had the highest density data centers in the world, right? They went to like 40 kilowatt racks when everyone was still at 12. And if you've ever stepped foot inside of most data centers, they're like pretty cool and dry -ish. if you step inside of an Amazon data center, they feel like a swamp. It feels like where I grew up. It's like humid and hot because they're optimizing every percentage. Your point here is that Amazon's data centers aren't equipped for the new type of infrastructure, but when you compare them to the cost of the GPU, having a complex cooling arrangement is fine.

1:02:14We made a call on Acera Labs a few months ago, a couple months ago when they were at 90 and it's gone to 250 the month after because of what orders Amazon is placing with them. But there's certain things with Amazon's infrastructure, I won't get too much into it, but their rack infrastructure requires them using a lot more of a Sterilabs connectivity products. And the same applies to Coolie. It's on the networking and cooling side. They just have to use a lot more of this stuff. But again, this stuff is inconsequential on cost compared to the GPU. you can build my question was more like look I may need a major river close by for cooling at this point in many areas I just can't get enough water it's probably power in the same region there's 2 gigawatt scale sites that they have power all secured wet chillers and dry chillers all secured everything's fine it's just not as efficient but that's fine they're going to ramp the revenue they're going to add the revenue not that I necessarily think Amazon's internal models are going to be great or hey their internal ship is better than NVIDIA's or competitive with TPU or their hardware architecture is the best.

1:03:26I don't necessarily think that's the case but they can build a lot of data centers and they can fill them up with stuff that will be rented out. It's a pretty simple thesis. How important has Anthropic been to the co -design for Tranium? Because I remember we had a portfolio company. This was summer 2023. They invited them to AWS. They spent, man, I think eight hours with them over the course of a week trying to figure out Tranium back then. It was just impossible to work through. Is that, you know, obviously that portfolio company hasn't gone back and tried it now, but like how different is it now based on what you're hearing?

1:04:08in. Oh, it's still bad. It's tough to use.

1:04:17This is sort of the argument that every inference company offers, including the AI hardware startups, is because I'm only running three different models at most, I can just hand -optimize everything and write kernels for everything and even go down to an assembly level. It is pretty hard. It is pretty hard. But you tend to do this for production inference anyways. You aren't using QDNN, which is NVIDIA's library that's super easy to generate kernels and stuff. Or not generate kernels, but anyways. You're not using these ease -of -use libraries. When you're running inference, you're either using Cutlass or stamping out your own PTX or in some cases people are even going down to the SaaS level.

1:05:07When you look at say an OpenAI or an Anthropic, when they run inference on GPUs, they're doing this. And the ecosystem is not that amazing once you get all the way down to that level. It's not like using NVIDIA GPUs is easy now. You have an intuitive understanding of the hardware architecture because you work on it so much and everyone's worked on it and you need to talk to other people. but at the end of the day, it's not easy. Whereas Anthropic, Traneum, or TPUs, actually the hardware architecture is a little bit more simple than a GPU. Larger, more simple cores rather than having all this functionality.

1:05:46Less general. So it's a little bit easier to code on. There's tweets from Anthropic people saying when they are doing that low level, actually they prefer working on Traneum and TPU because of the simplicity. to be clear Tranium and TPU I mean Tranium especially is very hard to use not for the faint of heart it's very difficult but you can do it if you're just running like if I'm Anthropic and I must only run Claude 4 .1 Opus for Sonnet and screw it I won't even run Haiku I'll just run Haiku on GPUs I'm just going to run two models and actually screw it I'm just going to run Opus on GPUs too and true TPUs.

1:06:30Sonnet is the majority of my traffic anyways. I could spend the time. And how often am I changing that architecture every four or six months? It's not even changing that much, honestly. I mean, from three to four definitely did change, right? Yeah, I mean, define architectural change. At a high level, the primitives are more or less the same across the last couple of generations. I don't know enough about anthropics model architecture, to be honest, but I think from what I've seen at other places, there have been enough changes that it takes time to program this. The main thing is, if I'm anthropic and I have 7 billion ARR now or whatever, by the end of next year, north of 20, ARR is maybe even 30.

1:07:20My margins are 50%, 70%. That's $15 billion of training that I need. They can run on Sonnet. And most of that's going to be Sonnet 3, 5, or sorry, 4, 5, whatever it is. It's going to be one model serving most of the use cases. So I could spend the time, and it'll work on this hardware. Yeah, totally. Maybe on the topic of non -consensus calls you've made, and maybe I'll move to another cloud. In June, you guys said that Oracle is winning the AI compute market. And then in this pod, we've already referenced the big jump, obviously, that Oracle had. I think it was the single largest gain that a company with over $500 billion in market cap has ever had.

1:08:06Was the 2023 Q1 NVIDIA not bigger? It might have been smaller. I think it was maybe close. We'll fact check ourselves. That's amazing. But obviously this is the massive commitment that was announced. Can you walk us through why you made that call then and just sort of why Oracle is poised to do so well in such a competitive space. Yeah, so Oracle, they're the largest balance sheet in the industry that is not dogmatic to any type of hardware, right? They're not dogmatic to any type of networking. They will deploy Ethernet with Arista. They'll deploy Ethernet through their own white boxes. They'll deploy NVIDIA networking, InfiniBand or Spectrum X.

1:08:54And they have really good network engineers. They have really great software across the board, right again, like ClusterMax. They were ClusterMax Gold because their software is great. There's a couple things that they needed to add that would take them higher, and they're adding those, right? To Platinum, right? Which was where CoreWeave was. And so a couple of two things, right? OpenAI's got insane compute demand. Microsoft is quite pansy. They're not willing to invest in, they don't believe OpenAI can actually pay the amount of money. I mentioned earlier, the $300 billion deal, OpenAI, you don't have $300 billion.

1:09:32And Oracle's willing to take the bet. Now, of course, the bet is a bit like, there's a bit more security in the bet in that Oracle really only needs to secure the data center capacity. So this is sort of like how we came across the bet. and we've been telling our institutional clients especially in a super detailed way whether it be the hyperscalers or AI labs or semi -electric companies or investors in our data center model because we're tracking every single data center in the world Oracle doesn't build their own data centers either they get them from other companies they co -engineer but they don't physically build them themselves and so they're quite nimble in terms of being able to assess new data centers, engineer them so we saw all these different data centers Oracle is snatching up in deep discussions, snatching up, signing, et cetera.

1:10:17And so we have, hey, GigaWatt here, GigaWatt there, GigaWatt there, right? Abilene, two GigaWatt, right? You have all these different sites that they're signing up and discussions with, and we're noting them. And then we have the timeline because we're tracking the entire supply chain, we're tracking all the permits, regulatory filings through language models, using satellite photos constantly, and then supply chain of chillers, transformer equipment, generators, et cetera. We're able to make a pretty strong estimate of quarter by quarter in our data center, or quarter by quarter, how much power there is for each of these sites.

1:10:53Some of these sites that we know of aren't even ramping until 2027, but we know that Oracle signed it. And we have this sort of ramp path. So then it's this question of like, okay, let's say you have a megalot, for simplicity's sake, which is a ton of power, but now it doesn't feel like much. We're in the gigawatt era. If you're talking about a megawatt, you fill it up with GPUs. How much do the GPUs for a megawatt cost? Actually, it's even simpler to do the math. If I'm talking about a GB200, each individual GPU is 1 ,200 watts. But when you talk about the CPU, the whole system, it's roughly 2 ,000 watts.

1:11:35At the same time, all in, everything, simplicity's sake, $50 ,000 per GPU. The GPU doesn't cost them. There's all the peripheries. So $50 ,000 CapEx for 2 ,000 watts. So $25 ,000 for 1 ,000 watts. And then what's the rental price for GPU? If you're on a really long -term deal, volume 270, 260 in that range, then you end up with, oh, it costs like $12 million per megawatt to rent a megawatt. And then each chip is different. So we track each chip, what the capex is, what the networking is. So you know what each chip is. You can predict what chips they're putting in which data centers, when those data centers go online, how many megawatts by quarter.

1:12:20And then you end up with, oh, well, Stargate goes online in this time period. They're going to start renting it this time. It's this many chips. Each Stargate site, right? And so therefore, this is how much OpenAI would have to spend to rent it. And then you prick that out, and we were able to predict Oracle's revenue with pretty high certainty, and we matched pretty dead on what they announced for 25, 26, 27, and we were pretty close on 28. The surprise for us was that they announced some stuff that 28, 29, data centers that we haven't found yet, but we'll find them, of course. This methodology lets you see what data centers are you getting, how much power, what are they signing, how much incremental revenue that is, when that comes online.

1:13:06And so that's sort of the basis of our Oracle bet. Obviously, in the newsletter, we included a lot less detail. But it was that thesis, right? That like, hey, they have all this capacity. They're going to sign these deals. In our newsletter, we talked about two main things. We talked about the OpenAI business, and then we talked about the ByteDance business. And presumably, tomorrow, on Friday, there's going to be an announcement about TikTok and all this. but the ByteDance business, huge amounts of data center capacity that Oracle is also going to lease out to ByteDance. So we did the same methodology there.

1:13:43With ByteDance, it's pretty certain they'll pay because they're a profitable company. With OpenAI, it's not. So there's got to be some error bars as you go further out in terms of will OpenAI exist in 28, 29, 30 and will they be able to pay the 80 plus billion dollars a year that they've signed up to Oracle with? That's the only risk here. and if that happens, then Oracle's downside is also somewhat protected because they only sign the data center, which is a minority of the cost. The GPUs are everything, and the GPUs they purchase one to two quarters before they start renting them. The downside risk is pretty low for them in terms of if they don't get the deal, well, they don't get the revenue, but it's not like they're stuck with a bunch of assets they bought that are worthless.

1:14:24Is there another angle here? OpenAI and Microsoft were off BFFs and now they're filed to voice papers and they just want to diversify and then that's pushing them away towards other providers? Yeah, so Microsoft was exclusive compute provider. It got reorg to write a first refusal. And then Microsoft... Is it now your last choice or something like that? No, it's still write a first refusal, but it's like Microsoft... Those two are not mutually exclusive. Well, if OpenAI is like, we're going to sign a $80 billion contract or a $300 billion contract for the next five years, do you guys want it? And they're like, no.

1:15:01what? Okay, cool. And then they go to Oracle. And OpenAI is sort of like OpenAI needs someone with a balance sheet to actually be able to pay for it. And then they'll make tons of money off of OpenAI on the margins, on the compute, and the infra, and all these things. But someone's got to have a balance sheet. And OpenAI doesn't have a balance sheet. Oracle does. although given the scale of what they signed, we had also had another source of information which was that they were talking to debt markets because Oracle actually just needs to raise debt to pay for this many GPUs over time. Now they won't do it immediately.

1:15:44They can pay for everything this year and next year from their own cash. But in 27, 28, 29, they'll start to have to use debt to pay for these GPUs, which is what CoreWeave has done and many of the neoclouds, most of it's debt financed. Even Meta went and got debt for their Louisiana mega data center. It's literally better on a financial basis to do buybacks with your cash and get debt because the debt is cheaper than the return on your stock. It's like a financial engineering thing.

1:16:15Who's out there, right? It could be Amazon, it could be Google, it could be Microsoft, or it could be Oracle or Meta. Meta's obviously not. Microsoft's chickened out. Amazon, Google, and Oracle. That's all that's left. Google would be an awkward fit. Yeah, Google would be an awkward fit. Amazon would be a fine fit. Exactly, right? It's a very drop -back. Well, I guess maybe on the topic of these giant data center build -outs, you guys just released a piece on XAI and Colossus 2. Are you getting less impressed by these feats of building something this massive in six months? Or is it still very impressive to you guys?

1:17:01You know, this is the thing I've said about AI researchers is that they're the first class of humans to think about things on an order of magnitude scale. Whereas people have always thought about things in terms of percentage growth ever since industrialization. And before that, it was just absolute numbers, right? You know, sort of like humanity is evolving in terms of how we think because things are changing faster. Everything is an ox scale.

1:17:28It was really impressive when GPT -2 was trained on so many chips and then GPT -3 was trained on 20k 100s. Sorry, GPT -4, 20k 100s. It's like, holy crap. Then it was like, oh, the error of 100k GPUs clusters. We did some reports around 100k GPU clusters. But now there's like 10 100K GPU clusters in the world. I was like, okay, that's kind of boring. But it's like 100K GPUs is like over 100 megawatts. Now it's like literally in our Slack and some of these channels, it's like, oh, we found another 200 megawatt data center. There's someone who puts the yawning emoji every time. and I'm like, dude, what?

1:18:18Like, now it's only exciting if you do gigawatt scale. Like, we're in gigawatt era, yeah. Yeah, yeah. And I'm sure, like, you know, I'm not sure. Maybe we'll start yawning to that, too. But, like, you know, the log scale of this is, like, the capital numbers are crazy, right? Like, you know, it's crazy enough that OpenAI did, like, $100 billion trading run. You know, or, you know, like, then they did a billion dollar trading run. Now we're talking about $10 billion trading runs, right? It's crazy that we think in log scale. But yes, things are only impressive. What Elon's doing in Tennessee, in Memphis, first time was crazy.

1:18:58100K GPUs in six months. He bought a factory in February of 24 and had models training within six months. And he did liquid cooling, first large -scale data center. at this scale for AI, doing liquid cooling. All these sorts of crazy firsts. Putting generators outside, like cat turbines, all these different things to get the power, mobile substations, all these different crazy things. Tapping the natural gas line that's running alongside the factory. So he does this, and it's like, holy crap. And he did it for 100K GPUs, 200, 300 megawatts. Now he's doing it for a gigawatt scale. and he's doing it just as fast, right?

1:19:46And so you would think this is obviously way more impressive than he did it again. But maybe I'm desensitized, but you've given the child too much candy, right? And now the child doesn't like apples, right? I don't know. So yeah, a gigawatt data center. there was all these protests around his Memphis facility. People were like, oh, you're destroying the air. And it's like, have you booked around that area of Memphis? There is a gigawatt gas turbine plant that's just powering generally that area. There's a sewage plant that's servicing the entire city of Minnesota. Or sorry, city of Memphis. And there's open air pits of the open air mining.

1:20:34There's all sorts of disgusting shit around there which is needed, right? We need that stuff to have a country run, right? To be clear.

1:20:44It's like people are complaining about a couple hundred megawatts in air of generation. So he got protests from all sorts of people. You got super into the politics side of things. NAACP even protested him. He really got some local municipalities to be like, oh, I don't like this. And so he couldn't do as much as he wanted to in Memphis. but he still needed the data center to be close because he wanted to connect these data centers super high bandwidth, super close. And he already had a lot of infrastructure set up there. So he bought another distribution center at this time. And it's still in Memphis, but the cool thing about Memphis is it's right across the border from Mississippi.

1:21:25So now, it's like 10 miles away from his original one, but his facility is like a mile away from Mississippi. And he bought a power plant in Mississippi and he's putting turbines there. is the regulation is completely different. If the question is really galvanize resources and build it really fast, maybe Elon is ahead of everyone. He hasn't made the best model yet or he doesn't have the best model, at least today, I think. You could argue Grok 4 was the best for a little period of time, but it's truly amazing how fast he's able to build these things. For first principles, most people are like, fuck.

1:22:06We can't build the power. We can't do power here anymore. I guess we have to find a new site. It's like, no, just go across the border. Go to Mississippi. My favorite thing is Arkansas is right there. So Mississippi gets mad.

1:22:22The regulation, the future data centers built in places where multiple states meet. Four quarters, yeah. The optimal regulation. There we go. Is there a point in the U .S. with five? I know there's a point with four. Four states intersect there. Maybe that's according to a data center. I'm going to buy real estate in that area. I've read it. Well, I guess on the topic of just maybe new hardware, you had this piece analyzing TCO for GB200s. And I'm kind of going to ask this question on behalf of our portfolio companies, which it sounds like you're helping them already. But one of the findings that I thought was really interesting was TCO was sort of 1 .6x H100s for GB200s.

1:23:10And so obviously, you know, there's this point on, okay, that's sort of the benchmark for the performance boost that you're going to need to at least make the sort of performance cost ratio benefit from switching over. Maybe just talk about what you've seen from a performance standpoint and what do you recommend to portfolio companies, maybe in a smaller scale than XAI, who are thinking about new hardware, try to get it. There's capacity constraints, obviously. Yeah, I mean, that's a challenge, right? With each generation of GPU, it gets so much faster that you end up like you want the new one.

1:23:47And in some metrics, you could say GB200 is three times faster than, or two times faster than the prior generation. Other metrics, you can say it's way more than that, right? So if you're doing pre -training versus inference, right? They can run everything four -bit, right? Yeah, if you can run it four -bit or just inference and take advantage of the huge NVLink, NVL72, you know, there's ways you can squint and say GB200 is only 2x faster than H100, in which case 1 .6x TCL. It's worthwhile, right? It's worth going to the next gen. But more marginal. It's more marginal. It's not a big deal. Then there's other cases where it's like, well, if you're running deep seek inference, the performance difference per GPU is north of 6, 7x, and it continues to optimize for deep seek inference.

1:24:40And so then it's like, well, I'm only paying 60 % more for 6x. It's a 4x or 3x performance per dollar gain. like absolutely right and if you're like in running inference of deep seek that can also include rl right um and so the question is sort of and then and then the other question is like well the gpu is new you know there's also b200 there's gb200 there's b200 b200 is much more simple from a hardware perspective it's just eight gpus in a box so then it's not as much of a performance gain especially in inference but you have um you have all the stability right it's an eight gpu box it's not going to be unreliable.

1:25:16The GB200s are still having some reliability challenges. Those are being worked through. It's getting better and better by the day, but it's still a challenge. When you have a H100 box, or H200, eight GPUs, one of them fails. You take the entire server offline and you have to fix it. If your cloud's good, they'll swap it in. but if it's GB200, what do you now do with 72 GPUs? If one fails, do you break the whole thing and get a new 72? The blast radius of a failure, GPU failure rates at best are the same and likely worse, gen on gen because everything's getting hotter, faster, et cetera. So at best, the failure rates are the same.

1:26:00Even if you model the failure rates as the exact same because you go from one out of eight to one out of 72, it's a huge problem. So now what a lot of people are doing is they run a high -priority workload on 64 of them, and then the other eight, you run low -priority workloads. Which is then like, okay, this is this whole infrastructure challenge. I have to have high -priority workloads, I have to have low -priority workloads. When a high -priority workload has a failure, instead of taking the whole rack offline, you just take some of the GPUs from the low -priority one, put it in the high -priority one, and then you just let the dead GPU sit there until you service the rack at a later date.

1:26:32And it's like, there's all these complicated infrastructure things that make it so, oh wait, actually that 3x or 2x performance increase in pre -training is lower because the downtime is higher. Slash, I'm not using all the GPUs always. Slash, I'm not smart enough or I don't have the infra to have low priority and high priority workloads. It's not impossible. The labs are doing it. If I'm running a cloud, it's actually really hard because I probably have to rent the spares out of spot instance or something. No, no, no, no, because it's a coherent domain. It's NVLink. You don't want anyone touching that.

1:27:09So it has to be the end customer. You don't have to leave them because it's empty, spare. So it's even worse. The end customer usually would just be like, I want them and I will. And the SLAs and the pricing, everything is accounting for that, right? So generally when you have a cloud, you have an SLA, right? That is, hey, it's going to be uptime, it's going to be 99%, blah, blah, blah, right? For this period. with GB200 it's 99 % for 64 GPUs not 72 and then it's like 95 % for 72 now it differs across every cloud every cloud is a different SLA but like they've adjusted for this because they're like look this hardware is just finicky do you still want it we will credit you in that 64 of them will always work right not 72 and so there's this whole finicky nature and the end customer has to be capable of dealing with the unreliability.

1:28:00And the end customer can just continue to use V200, performance games not as much. The whole reason you want this 72 domain is so you can have some of these gains. But you have to be smart enough to be able to do it. And that's challenging for small companies. Totally. And we just announced the Rubin pre -fill cards, like CTX, CPX, there we go. What's your take on that? Does it cannibalize? Dude, by the way, I don't know if this is like brain rot or like, I don't know, but like, I can't remember what I had for lunch yesterday. But I know the model number of every fucking chip, like. Hots you in your dreams.

1:28:38We're broken, we're broken. Living the dream. No, no, no, no, no. You know. Why do you pre -announce a product that's 5x faster for certain use cases? Is that that much? I think it's a potential, right? Historically, AI chips were AI chips, right? And then we started getting a lot of people saying, this is a training chip, this is an inference chip. Actually, training and inference are switching so fast into what they require that now it's still one chip. Actually, there are still workload -level dynamics that differ, but the main workload is inference even in training, right? Because of RL, most of that is generating stuff in an environment and trying to achieve a reward.

1:29:23So it's inference still. Training is now becoming mostly dominated by inference as well. But inference has two main operations. There is calculating the KV cache for pre -fill. Here's all these documents. Do the attention between all of them, between all the tokens, whatever type of attention you use. And then there's decode, which is auto -aggressively generate each token. These are very, very different workloads. Initially, the ideas or infrastructure techniques, the ML systems techniques were, oh, okay, I will just make the batch size every single forward pass this big. Let's call it, I'll make it 1 ,000 big.

1:30:05Maybe I'll run 32 users concurrently. That way, now I still have 900 something left, 960 left. That 960 is actually doing the prefill for if a request comes in, it chunks it. It's called trunk pre -fill. You pre -fill chunks of it now, you get really good utilization on GPUs, but then that ends up impacting the decode workers. The people who are auto -aggressively generating each token that having slower TPS. And tokens per second is really important for user experience and all these other things. So then the idea is like, okay, these two workloads are so different, and they are literally different.

1:30:42You pre -fill and then you decode. It's not like you're interleaving them. So why don't we split them entirely? And this is done on the same type of chip, right? OpenAI, Anthropik, Google. Pretty much everybody does that. Everyone good. They're big guys. Together, fireworks. All these guys do pre -fill decode, disaggregated pre -fill decode. So they run pre -fill on a set of GPUs, decode on a set of GPUs. Why is this beneficial? Because you can auto -scale them, right? You can, hey, all of a sudden, I have a lot more long context workers. I allocate more resources to pre -fill. Oh, all of a sudden, not all of a sudden, but over time my traffic mix is not long input, short output.

1:31:20It's short input, long output. I have more decode workers. This way I can guarantee, and so now I can auto -scale the resources differently and I can also guarantee that my pre -fill time is, what's really important in search is how fast you get the page to start loading, not when does the resource happen. What do people do in games? The loading screen often has some sort of interactive environment or it blends in over time or whatever it is. It has tips and tricks, ways to distract, you. The same thing is, there's studies and papers out there that users prefer a faster time to first token, right?

1:31:54First token gets streamed to me sooner, even if the total time to get all my tokens is a little bit longer. I can't read that fast anyways, right? I mean, I like to skiv. I mean, most models return about speed reading. But you need that, right? The idea is that you want to guarantee time to first token as a certain level for user experience reasons. Otherwise, people are like, screw this, not using AI. The decode speed matters a lot too, but not as much as time to first token. And so by having separate pre -filled decode, you do this, right? But now you've already, and this is all in the same infrastructure, you've already done this.

1:32:31So now it's like, what's the next logical step? These workloads are so different. Decode, you have to load all the parameters in and the KV caches to generate a single token. You batch a couple users together, but very quickly you run out of memory capacity or memory bandwidth because everyone's KV cache is different. The attention of all the tokens, right? Whereas on Prefill, I could even just serve one or two users at a time because if they send me a 64 ,000 context request, that is a lot of flops, right? 64 ,000 context requests. I'll use Llama70B because it's simple to do math on like 70 billion parameters.

1:33:08That's 140 gigaflops per token times 64 ,000. and that's many, many teraflops. You can use the entire GPU for a second, potentially, depending on the GPU, to just do the pre -fill. And that's just one forward pass. So I don't necessarily care about loading all the tokens or all the parameters in KVCache and FAST. All I care about is all the flops. And so that leads us to, I think it was a long -winded explanation because it's hard for people to understand what CPX is. I've had a lot of, even my own clients, we set multiple notes explaining and they're like, I still don't understand. I'm like, shit, okay.

1:33:46Sentence is all you need, paper. You can't expect. Think about a networking person. They're like, I don't need to know about this. Attention is all you need. Or think about an investor. Data center and a parade are like, oh, there's two chips, why? Should I build my data center differently? I got to explain everything. Or just like, no. You don't have to build differently. But anyways, you get to know... In Stanford, there's 25 % of all students, not CS students, of all students read their paper. Read what paper? Attention is all you need. That's low. They do gym majors and you don't like the philosophy guys.

1:34:24I find this amazing. Anyway, sorry. The Middle East, I can't remember what country it is, has AI education starting at age eight and in high school they have to read Attention is all you need. Wow. Someone told me that their hand had to read Attention is all you need. Which is, I don't know. Look, look, top -down mandates for education. Maybe they work, maybe they don't. Maybe people like homeschooling their kids. I don't know, I went to public school. Back to your readers. Just on the topic of hardware cycles, I wanted to maybe... Sorry, I didn't actually explain what CPX is. So CPX is a very compute -optimized chip, whereas for pre -fill and then decode, just to specifically say it, is the normal chips with HBM.

1:35:08HBM is more than half the cost of the GPU. if you strip that out, you end up having a much cheaper chip passed on to the customer. Or if NVIDIA takes the same margin, then the cost of this pre -fill chip is much, much lower, and now the whole process is way cheaper and more efficient. Now long context can be adopted. I love that we're actually going into all this detail because I had a more 10 ,000 -foot view question for you, which is, I haven't been following the semi -market as closely as you have. I probably started with the A100. And I remember helping GNOME at Character, this is summer of June 2023, chase down GPUs.

1:35:51And the only thing that mattered at that time was delivery date because there was a huge capacity crunch. And then to see that over the last two years evolve where, you know, let's say six to 12 months ago, people were doing these RFPs to 20 NeoClouds, right? and the only thing that mattered to some degree was price. People actually do RFPs for GPUs? Yes. So just to be clear, my opinion on how you buy GPUs is that it's like buying cocaine or any other drug. This is described to me, not me. I don't buy cocaine. Someone tells me this. I'm like, holy shit, that's right. You call up a couple people, you text a couple people, you ask, yo, how much you got?

1:36:31What's the price? It's like... Exactly. This is fucking like buying drugs. Sorry, sorry. It's the same way. We have Slack connects with 30 NeoClouds as well as some of the major ones. And we just send them a message like, hey, customer wants this much. This is what they're looking for. And then they send quotes. I know this guy. I know a guy. Well, so I think that's actually a very accurate description. And I've sent countless port codes, your ClusterMax original posts, because I thought it did a really good job breaking them down. but maybe one question to end on for me is just what era are we in now with Blackwell's coming online?

1:37:11Are we sort of back to the summer 2023 era and that's kind of the cycle that we've just entered or what's sort of your view on where we are? That's a very good question. For one of your portcos, we were like, you know, after their difficulties with Amazon, we were like, okay, let's actually get you to fuse. The original deals we got you were gone but here's some other deals. It turned out that multiple major neoclouds had sold out of hopper capacity. Their Blackwell capacity comes online in a few months. It's a bit of a challenge. Due to inference? Inference demand has been skyrocketing this year.

1:37:53Reasoning models. These reasoning models are revenue. It's been skyrocketing this year. Also, there's a bit of the Blackwell comes online but it's hard to deploy so it takes a little, you know, there's a learning curve to deploying it. So whereas you got down to like, you buy the hopper, you install the data center, it's running within like a month or two, right? For Blackwell, it's a longer time frame because of reliability challenges, it's a new GPU. I mean, it's just learning pain, right? Learning growing pains. So there was this gap of how many GPUs are coming onto the market right as revenue is starting to inflect.

1:38:27And so a lot of capacity got sucked up, right? and actually prices for Hopper bottomed like three or four months ago or like five or six months ago. And actually they've like crept up a little bit now. They're still like, you know, not, so I don't think we're quite 2023, 2024 era of GPUs are tight. But certainly if you want to, if you want like just a few GPUs, it's easy. But if you want a lot, it's hard. Like you can't get capacity that instantly. Yeah. Wow. What a time. shall we wrap on that Dylan this was another instant classic thank you so much for coming on the podcast it's like two hours bro we couldn't stop thanks so much it was great thanks for listening to the A16Z podcast if you enjoyed the episode let us know by leaving a review at ratethispodcast .com we've got more great conversations coming your way see you next time as a reminder the content here is for informational purposes only, should not be taken as legal business, tax, or investment advice, or be used to evaluate any investment or security, and is not directed at any investors or potential investors in any A16Z fund.

1:39:41Please note that A16Z and its affiliates may also maintain investments in the companies discussed in this podcast. For more details, including a link to our investments, please see a16z .com forward slash disclosures.

1:40:00Thank you.

From the publisher

Nvidia’s $5 billion investment in Intel is one of the biggest surprises in semiconductors in years. Two longtime rivals are now teaming up, and the ripple effects could reshape AI, cloud, and the global chip race.

To make sense of it all, Erik Torenberg is joined by Dylan Patel, chief analyst at SemiAnalysis, joins Sarah Wang, general partner at a16z, and Guido Appenzeller, a16z partner and former CTO of Intel’s Data Center and AI business unit. Together, they dig into what the deal means for Nvidia, Intel, AMD, ARM, and Huawei; the state of US-China tech bans; Nvidia’s moat and Jensen Huang’s leadership; and the future of GPUs, mega data centers, and AI infrastructure.

 

Resources: 

Find Dylan on X: https://x.com/dylan522p

Find Sarah on X: https://x.com/sarahdingwang

Find Guido on X: https://x.com/appenz

Learn more about SemiAnalysis: https://semianalysis.com/dylan-patel/

 

Stay Updated: 

If you enjoyed this episode, be sure to like, subscribe, and share with your friends!

Find a16z on X: https://x.com/a16z

Find a16z on LinkedIn: https://www.linkedin.com/company/a16z

Listen to the a16z Podcast on Spotify: https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX

Listen to the a16z Podcast on Apple Podcasts: https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711

Follow our host: https://x.com/eriktorenberg

Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.

Stay Updated:

Find a16z on X

Find a16z on LinkedIn

Listen to the a16z Podcast on Spotify

Listen to the a16z Podcast on Apple Podcasts

Follow our host: https://twitter.com/eriktorenberg

 

Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.


Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

More from The a16z Show

All 489 episodes
Dylan Patel on the AI Chip Race - NVIDIA, Intel & the US Government vs. ChinaThe a16z Show · 1 h 40 min
Listen in VO