Why Hardware-Software Co-Design Is AI's Real 100x: Dylan Patel of SemiAnalysis

30 Jun 2026 · 1 h 10 min · 31 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Hardware-software co-design as the main driver of AI efficiency (the “real 100x”), with a focus on inference benchmarking, cost/throughput tradeoffs, and where bottlenecks will be next (memory bandwidth, power density, energy).

Key claims

point-in-time inference benchmarks go stale; inference efficiency improves ~60x cost for equal quality and ~40x intelligence-per-watt, largely via co-optimization across model, software, and hardware; the throughput vs interactivity Pareto curve is the most important curve for downstream infrastructure decisions; space data centers are negligible in 3–5 years but could dominate incremental compute by ~2040; future gains require co-design rather than isolated kernel or hardware tweaks.

Notable examples

SemiAnalysis’s “InferenceX” runs daily automated benchmarks on donated hardware (CoreWeave, Crusoe, Nebius, Oracle, Microsoft, Amazon, Google, OpenAI) with ~50M hardware donated (potentially >$100M) across many chip types and models; co-design examples include DeepSeek/DeepSeek-V3 optimized for Hopper shapes and later models for Blackwell/Huawei; TPUs can be great for some model families but “suck” for DeepSeek.

Guests

Dylan Patel, founder of SemiAnalysis; background includes teenage forum moderation across Android/Apple/Google and hardware, early hardware tinkering (Xbox “red ring” repair), quant risk-firm experience, and building SemiAnalysis from 2020 blog traction after being doxxed.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Dylan's Journey into Technology

0:45 to 3:18

Dylan shares his background growing up in a family business and early interests in technology and hardware.

“We're here in the semi-analysis office with Dylan Patel.”

The Impact of Early Experiences

3:18 to 6:00

Dylan discusses formative experiences, including fixing hardware issues and engaging in online forums.

“My birthday is in May, and it was April when the Xbox 360 was announced.”

Transitioning to Semi-Analysis

6:00 to 8:10

Dylan explains his transition from working in finance to founding Semi-Analysis, driven by personal events and market trends.

“Like Spanish, I got like not the greatest grades.”

Journey of Building Semi-Analysis

8:10 to 10:28

Dylan describes the early days of Semi-Analysis, including personal struggles and his strategy in the semiconductor space.

“And so I was like posting even more than normal.”

Networking at Conferences

10:28 to 12:36

Dylan shares insights from attending various conferences and the value of networking in the semiconductor industry.

“I go to 40 plus conferences a year, no matter where in the supply chain it is.”

Understanding Supply Chain Dynamics

12:36 to 14:00

Dylan discusses the complexities of the semiconductor supply chain and shares interesting anecdotes about industry challenges.

“Whereas you go to NeurIPS, a couple of times you can understand, okay, what's neuro-symbolic reasoning?”

The Future of Inference

14:00 to 14:43

Discussion on the potential of inference as a major market.

“Inference, going to be the biggest market on earth, biggest market beyond earth.”

Challenges in Inference Benchmarking

14:43 to 16:42

Exploring the challenges of point-in-time benchmarking in AI inference.

“Like semi-analysis, we do a lot of stuff that's like, A lot of it is research for institutional clients and our subscription-first products, but a lot of it is also like, hey, this would just be cool to figure out.”

Creating Optimal Inference Configurations

16:42 to 18:15

Insights on how to create optimal configurations for AI inference.

“We've got over$50 million of hardware donated to us.”

The Importance of Cost and Efficiency

18:15 to 21:07

Analyzing the relationship between cost, efficiency, and inference compute.

“The throughput interactivity curve is the most important one?”
Show all 31 chapters

The Future of Inference Compute in Space

21:07 to 23:02

Discussing the potential for inference compute to move into space.

“It's the cost of building power on terrestrial land and how much power are you going to be able to do on terrestrial land?”

Intelligence Per Watt: Trends and Predictions

23:02 to 26:57

Examining trends in intelligence per watt and its implications.

“As far as where we are from the human brain, we're many orders of magnitude away.”

The Co-Design of Hardware and Software

26:57 to 28:01

Understanding the significance of co-design in hardware and AI models.

“But there's been innovations on every layer, and really the biggest gain and the beauty of the best labs is when they co-optimize all three.”

The Importance of Hardware-Software Co-Design

28:01 to 29:11

Explore the significance of co-optimizing hardware and software for breakthrough innovations.

“There will always be bottlenecks somewhere in that optimization that they're lagging behind and then need to get pulled forward.”

Tracking Bottlenecks in Technology

29:11 to 30:26

Discuss the major bottlenecks in technology, particularly in memory and power consumption.

“If you had to predict at any level of the stack, it can be literally anywhere, what are some of the bottlenecks you're tracking most acutely the next year?”

Innovative Approaches to Energy Solutions

30:26 to 32:41

Learn about innovative approaches to solving energy bottlenecks in technology.

“And so if a chip is 100 millimeters squared, generally the power consumption's around 100 or a little bit less.”

Comparing NVIDIA and TPU Technologies

32:41 to 34:28

Examine the competitive landscape between NVIDIA GPUs and Google TPUs.

“Sean, you've been media trained so well.”

The Role of Model Architecture in Hardware Choices

34:28 to 37:13

Understand how model architecture impacts the choice between GPUs and TPUs.

“And the way that Anthropic and Google's models are headed, it's actually a terrible decision, potentially, for them to train with GPUs.”

Impact of Open Source on Hardware Decisions

37:13 to 38:47

Consider how the trend towards open source affects hardware and software decisions in AI.

“ecosystem or open source models themselves.”

Fast Mode and Economic Considerations

38:47 to 41:30

Explore the advantages of fast mode and its economic implications in technology.

“I think Cerberus is a really innovative company.”

The Reality of AI Progress and ROI

41:30 to 42:00

Debate the misconceptions about AI's ROI and ongoing progress in model development.

“I think it's really fun inside of semi-analysis because we have 90 people and a big chunk of them are technologists, engineers across the whole supply chain.”

Understanding Semiconductor Misconceptions

42:00 to 44:20

Explore common misconceptions in the semiconductor industry and their implications.

“You don't wrestle with a pig because a pig enjoys it.”

Exciting Long-Term Innovations in Space and Silicon

44:20 to 45:40

Discover long-term innovations in the semiconductor industry and space technology.

“Are there longer-term things that you're really excited about?”

The Future of AI Chip Development

45:40 to 50:40

Discuss the future landscape of AI hardware and the potential for bespoke chip architecture.

“And he's so ahead of his time with Mosaic.”

Evaluating the Compute Crunch and Future Capacity

50:40 to 56:00

Assess the ongoing compute crunch and factors influencing future capacity and demand.

“have their niche and actually make money, even if the majority of the pie goes to NVIDIA and TPU and training them.”

Larry Page's Strategic Investment

56:00 to 56:42

Discussion on Larry Page's significant investment and its implications for the company.

“Larry Page invested a billion dollars at a$10 billion valuation, got 10 % of the company.”

Capital Raising and Market Dynamics

56:42 to 58:09

Exploration of current capital raising trends and their impact on major tech companies.

“But right now, every GPU that Amazon adds, they're making higher revenue or every TP or Tranium, whoever anyone adds, is making gross profit.”

Data Center Value Disparities

58:09 to 1:00:03

Analyzing the differences in value and rental rates across various data centers and technologies.

“of the people that are not as good at it kind of getting hit a little.”

Optimizing Data Center Utilization

1:00:03 to 1:02:24

Insights into how companies maximize data center utility and the implications for profitability.

“And I've seen stuff go as low as 100 still, or in India go as low as 80 because the grid's not reliable, the internet connection's not great, and it's a pretty mid data center, but at least it's a data center.”

AI Cloud and Market Competitiveness

1:02:24 to 1:04:28

Examining the competitive landscape of AI cloud services and the challenges hyperscalers face.

“And sometimes there are levers where you're selling more gigawatts, where each gigawatt is selling at a different price.”

The Future of NeoClouds and AI Labs

1:04:28 to 1:09:20

Discussion on the emergence of NeoClouds and their role in shaping the future of AI computing.

“So in 2023, I wrote a report that had Amazon really hate me.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Dylan Patel:I think it's really fun inside of Semi-Analysis because we have 90 people and like a big chunk of them are technologists, engineers across the whole supply chain. And then a big chunk is people who are formerly at hedge funds. And you see these arguments like people are like, oh, well, that doesn't matter. And it's like that someone's like, well, but cost. And then the engineer's like, no, no, no, but this technology is the closest. You see this organically like fight it out. And we're pretty informal. And, you know, given the fact that I was a forum moderator, you can imagine what this internal smack looks like.

0:29Dylan Patel:enjoying it. You don't wrestle with a pig because a pig enjoys it.

0:50We're here in the semi-analysis office with Dylan Patel. I'm Sean from Sequoia. I'm a partner, Sonia Huang. It's pretty insane what you've done. Semis five years ago were not very sexy in the West. They were sexy in the East, but people here in the West had kind of forgotten about them. You did not forget about them, though. You went very long. You created probably the premier research company in the space that's been educating the world and the state of the art from very technical details to supply chain to the bigger picture. There's rumors that Semi-Analysis recently passed 100 million of revenue.

1:27I don't know how accurate those are, but whatever the numbers are, you guys are crushing. It's as accurate as the information is, you know? You never know. There's also rumors that you might start a venture fund. I hear all the time in the ecosystem, people wanting affiliation with Semi-Analysis. You've built this trusted brand. And so whatever you do, it's working. It's clearly just the beginning of the journey for you. Congratulations on all of that. But how did this happen? like how did you first question is like what is the background how did you kind of get to where

1:58Dylan Patel:you are now well when i was a young boy and you know coming out of the womb so so okay so i grew up in like a small business my parents had a motel we lived in the motel we laid our gas station so you know uh i was selling you know i joke a lot of times the first neural network i trained was uh racially and and visually profiling people based on when they enter the gas station which cigarette to pick. Basically, the cigarettes were all extrude across the top, and I was too short to actually reach them, and technically it wasn't legal to sell cigarettes at that age, but whatever. I had to move the stepstool over to the right area.

2:34Dylan Patel:I started working my first office before it was legal, too. It was a good experience. Well, I didn't get paid, right? It's a family business. We had a motel, and then across the street was our gas station. Sometimes, someone would walk in, and so if an old white lady with curly hair walked in, I'd move the ladder or the stepstool over to where the camels are. And if, you know, and, you know, different, different age, demographic, uh, profession, you know, race, et cetera, I would move the stepstool over. And I joke, this is the first neural network I trained. Cause if I waited for them to tell me, I'd have to like move it over and then I'd step up versus like just being ready.

3:05Dylan Patel:Um, so, you know, menthols versus, you know, a hundred slims and all these things, you know, I think that's the first neural network I trained, but I grew up in family businesses, um, lived in a motel and, um, it all really goes back to when I was like, you know, was my eighth birthday. My birthday is in May, and it was April when the Xbox 360 was announced. For my birthday, I didn't ask for the Xbox, or I didn't ask for a birthday gift. My parents asked what I wanted. I asked for it for Christmas. We celebrated Christmas, but there was no way, at least at the time, I thought there was no way they would ask, would give me the Xbox 360 for Christmas, and so I got it for, asked for my birthday for tab for Christmas.

3:40Dylan Patel:Anyways, Christmas comes around, I get it. Fast forward a couple months, my cousin who lives in Alabama, they also lived in a motel, was going to come over for spring break, for his spring break, and we're going to hang out at my house. And he's in between me and my older brother and age. Brother's a bit more jockey, so he didn't really care too much about the Xbox. He played sometimes, but didn't really care. But my cousin, I wanted him to think I was cool, right? So I bragged many times on the phone. I was like, yeah, I got an Xbox. And then the Xbox broke. There was something, there's a hardware defect called the red ring of death.

4:11Dylan Patel:But long story short, I had to open it up and short the temperature sensor and it fixed it. But there was many other tricks I tried first and then other worked. And so that's sort of how I got into hardware is like open-pear doors box. By the time I was 12, I was on these forums a lot, reading, posting a lot. And this was around the time when Reddit ate all other forums. And so I became a moderator of Android and Apple and Google, as well as hardware and was looking at Intel, NVIDIA, and AMD and all these other forums. So I was built a PC, all these forums. I was watching, reading, posting a lot.

4:42Dylan Patel:but some of them I was moderating a lot. And so, you know, smartphones, watching smartphones develop from like very simple to speed racing to being technologically more advanced than PCs in many ways, architecturally, and same with like, you know, all the GPUs, like just tracking and watching at that, reading every comment, always having the economic tinge because I grew up in a small business. So I was always looking at the economics, right? There was a time where all the like, let's say neckbeards on the internet loved AMD GPUs. And like, I personally had bought an AMD GPU too because price performance.

5:15Dylan Patel:But then when it came down to like, what's technically better, I'd always be like, no, no, no, Nvidia's better because they use a smaller chip to get better performance and better power efficiencies and their margins better. And so like, I would always like talk about how Nvidia's margins were better than AMD's in the GPU landscape. And so it was like very fun. And you were 12 at the time? I started moderating when I was 12, but this is all through my teenage, tween age and high school years, right? Do you have any other weird hobbies or was it just semis? I played a ton of Starcraft at one point.

5:44Dylan Patel:I was Grandmaster on the North American ladder. Oh wow, I'm serious. Okay, so you've gotten just obsessively good at multiple things. Yeah, I mean, obsession is good. How were your grades? They were decent. I would say like, I had mostly A's, but they're classes that I thought were really boring, or I just didn't enjoy. Like Spanish, I got like not the greatest grades. you know but but it was like I speak fluent Spanish by the way so it's really dumb but like it's just sort of maybe that's why you didn't get a good grade I didn't learn Spanish to be fair but yeah so it's sort of my grades are fine right like I mean like they were fine enough for Asian parents I was better than most of school but you know it wasn't like you know try hard maxing for like you know all A's okay so you're very much a student of the internet then and this is how you develop this expertise.

6:37At what point did you decide to start Semi-Analysis, and what's been the biggest surprise since starting the company?

6:42Dylan Patel:Yeah, so I went to school. I got a few degrees in stuff that wasn't related to semiconductors. Was a quant for two years at a small quant risk firm. And then basically, there was a culmination of events that happened, right? One was that my, you know, sort of like I got screwed out of a bonus. I'd made my company many millions of revenue, of risk-free revenue, because I exploited a risk thing in the market, I think well over 10 million. And then someone else took credit for my work and all this sort of stuff. But eventually I did get right size, but I lost the social contract with the company I was working with.

7:16Dylan Patel:My grandparents grew up in my house with us or in the motel with us. They lived with us. And so very close to them. And my grandmother got dementia and she forgot who I was. And she fell down some stairs and had like a tragic accident and passed away. So all of that happened in early 2020. Additionally, there were some like, you know, girl things. And so, you know, there's a few things that happened that made me like kind of very sad. And so all of those things sort of culminated. Then COVID happened. And my brother's like, dude, just come stay with me. He lived in Nashville. So I came and stayed with him in Nashville.

7:48Dylan Patel:We were like, oh, lockdowns will be a few weeks. You can stay with me while they happen. And then you can go back home and, you know, whatever. Famous last words, lockdowns lasted much longer. But, you know, living with my brother for a few months, you know, it was like sort of like, okay, I didn't know what I was doing. I was now at my brother's home. Everything was his rules, you know, sort of like, you know, him and his fiance at the time, now wife, you know, were like there. And so like, I basically had to tiptoe around, but I didn't care about my job. And so I was like posting even more than normal.

8:13Dylan Patel:I'd always been posting a lot on the internet. I'd always been trading stocks a lot, but like I made a lot of money shorting COVID and long and COVID and like all this stuff. Semiconductor shortages happened around then too. And anyways, I was very much obsessed with posting and things like that. Eventually, around that time, I got into an argument with someone on the internet and they doxed me. They publicly revealed my identity for my anonymous account. At the time, I was like, oh no, I was scared. I stopped posting for three weeks. I was like, what am I doing? Why do I care? Then I just started posting under...

8:42Dylan Patel:I had blogs and stuff as well. I made a real blog, semi-analysis, and on my 24th birthday, I posted two blogs. And then from there, it was not a newsletter, but I got so much traction. Because now, instead of posting on anonymous, name is a real name. And I put a lot more effort into those two posts than I usually did. Instead of shit posting on the internet, it was real effort into the blog. You can actually go back and read those if you want. They're not that great, but they were good for the time. They were the best stuff you could find on the internet about SEBIES. And I just kept posting, posting, posting.

9:14Dylan Patel:I started getting a lot of consulting business. You know, 2020, I also sort of, I was, again, crashing out, didn't know what I wanted to do. So I packed everything up or sort of I took my truck. I bought a tent that fits on the back of the tent truck, bought an air mattress, whatever, and would like drove around all these national parks all around America. And so like two or three or four days of the week, I'd stay in a random motel where I negotiated the price to be like$30 a night for a room. And I would work on something else's stuff. And then the weekends I'd read books and oftentimes read textbooks while in some random national park or hiking and listen to audio books about semiconductors, about AI, about all the things that I cared a lot about and got way more educated over these six months where I'm just like going to every national park.

9:56Dylan Patel:And then the whole time I was posting. I was alone the whole time. I was posting blogs. Everyone was like, Dylan, what the are you doing? Pre-Star Link or the very early days of Star Link? Pre-Star Link. Pre-Star Link. Yeah. So it was like very much like, what are you doing? I travel around LATAM again for a year, initially with my friend and then with my ex for about a year. And then I go, then 22, 23, 24, end of 21, 22, 23, and 24. I'm still completely homeless since mid-2020, right? But I'm traveling around to every conference in the world. I go to 40 plus conferences a year, no matter where in the supply chain it is.

10:33Dylan Patel:I'm like, oh, that looks interesting. I guess I'll go to that. And I'm like, I went to one conference, like, wow, this is amazing. You get to talk to the experts. And they just like, they're going to talk to you because, and then you're so excited. And in the case of semiconductors, everyone's a boomer. So it's like, it's great to like, you know, they're like, they don't see young people who are like excited about it. So they're really happy to tell stuff. And so you just have to ask on this. Was there like a part of the supply chain or one of these conferences that, you know, particularly changed your view of the semi world or that you felt then or feel now is particularly underrated?

11:05Dylan Patel:I think the trade shows like rate and conferences range really widely. Obviously, some of the ones I have the most fun at include NeurIPS. Why is that? Because it's 20 ,000 AI researchers and they're generally in my distribution of age range, so it's a lot of fun. But they're also leading AI researchers and it's a lot of fun and you learn a lot. There's also a lot of parties. Then it ranges all the way to like, there's random chemical conference in Japan where it's 300 Japanese dudes, it's like 20 guys from ASML, 20 guys from TSMC, 20 guys from Intel, and those are the only people who speak English.

11:35Dylan Patel:everyone else speaks only Japanese and you're like, I guess it's still pretty interesting and fun. I think one thing that I have a skill set of is I'm able to bond with anyone regardless of their background and who they are. I'm able to talk to them, find something interesting to talk about. Oftentimes, it's the tech stuff. I think the most interesting conferences are oftentimes the really big ones because that's where the biggest stuff is happening. I think the niches that are really, really exciting is like SPIE. So there's IEEE, which is International Electrical Engineering, something. And then there's SPIE, which is another ecosystem.

12:11Dylan Patel:SPIE conferences are super, super deep in details. Every single one that I went to, especially like SPIE Advanced Lithography or SPIE Photo Mask, I went to them the first time. I didn't even understand 90 % of what I heard. And then I read, read, read, read. I'd made some contacts, of course. And then next time I went, I understood like half of what I went to. Third time I went, I understood 75 % of what I went to. Even now I went and I was like, I still don't understand everything that's going on. Whereas you go to NeurIPS, a couple of times you can understand, okay, what's neuro-symbolic reasoning?

12:41Dylan Patel:Okay, what's this? What's that? You can get a mapping of what everything is pretty quickly, but some parts of the supply chain are so arcane and so deep and so technical, it takes a lot of times for you to even understand what's happening on everything, for every research paper. It doesn't necessarily mean you didn't. You go to a conference for a few reasons. right? You understand the research, you understand, but like, it's all the research that's being published. But what you really care about is understanding how does that research intersection, intersect with technology. Also, how does that research differ from what's there today?

13:11Dylan Patel:And none of these research papers tell you what's happening today. But then you just ask people and you build contacts and you learn and then you like learn about the supply chain and oh, this company supplies this company, even though it's not publicly stated anywhere. Like, you know, you learn that the, the, this chemical is like cost about this much and a tool uses about this much. And you hear the horror stories of like this chemical had a shortage and it totally threw off this part of the supply chain. And then it turns out there's only three companies in the world that make that chemical.

13:40Dylan Patel:And it's like my favorite one is I learned a Japanese guy at that specific Japanese conference that I went to where almost no one spoke English in very broken English. He told me about how his father worked in this industry in the 1980s that the only factory in the world that built this chemical burned down, and that caused memory prices to double or triple. I was like, wow, not too different from today. Not at all. That's crazy. Inference, going to be the biggest market on earth, biggest market beyond earth. Agree or disagree? Obviously, use of tokens is going to be the biggest market, and the value that's created from tokens is going to be the biggest market.

14:18Dylan Patel:But I think tokenomics, the use of tokens, adoption of AI is the most important thing that's happening. Inference, whether it's open models or closed models will be like one of the biggest markets in the world. Much bigger than oil, I think much bigger than like, you know, many other parts, like inference of AI will be, you know, many percentage points of the GDP. Yeah. Right. What you've done with inference X, I think is, you know, industry standard, maybe say a word on why you started it, what it does. And, you know, what do people misunderstand about performance benchmarking on inference? Yeah.

14:48Dylan Patel:So, so to zoom back, right. Like semi-analysis, we do a lot of stuff that's like, A lot of it is research for institutional clients and our subscription-first products, but a lot of it is also like, hey, this would just be cool to figure out. Let's figure out how to figure it out and just post it publicly, and that gets more and more scale. We've done this with a lot of GPU benchmarking and testing and training performance and inference performance, but ultimately we saw inference benchmarking was point in time. You test it, and you take some time, and you release it, and it's slow and arcane and outdated because models change all the time.

15:22Dylan Patel:I feel like every week there's a new model, whether it's a Chinese model or today, Mythos 5, Fable dropped. New models are coming out all the time. On the software layer, PyTorch, VLM, SGLang, new drivers, new something drops. In fact, the update cycle for most of these libraries is twice a week. So you basically have the software updating all the time and therefore performance changing. New inference optimizations are coming out and those get updated. And so I feel like it's a relentless breakthrough after breakthrough after breakthrough that keeps driving efficiency and cost down, which is why we've seen model costs drop for equivalent quality by 60x a year.

15:58Dylan Patel:It's incredible. But to stay on top of that, you can't have point-in-time benchmarking. You need to have benchmarks be living and breathing, i.e., constantly running on the latest hardware, on the latest models. And so we embarked on a project, and we got a lot of buy-in from the ecosystem. and this was only possible because we had enough aura with some of the ecosystem where we were able to get CoreWeave and Crusoe and Nebius and Oracle and Microsoft and Amazon and Google and OpenAI to contribute to us compute. And then we were able to work with SGLang and VLM and now Radix, Arc, and Infraact, which are the private companies who are sort of leading those efforts, the open source efforts, to collaborate with us.

16:36Dylan Patel:We were able to get NVIDIA and AMD and Google and Amazon now because we're adding TPUs and Tranium to collaborate. Now we've got all these people collaborating. We've got over$50 million of hardware donated to us. Once we launch TPUs and Trainium, it'll actually be over$100 million of hardware. Maybe about 15 different chip types, all running these benchmarks every single day on all the latest model. The best model for Moonshot, the best model for Alibaba, the best model from... There's about five different Chinese models, the best open source models, the best Chinese labs there. We run benchmarks on their models every day and then also the best US open source models, GPT-LSS, NemoTron, et cetera.

17:12Dylan Patel:We're running these benchmarks every day in an automated fashion, and they run on these servers that are dedicated to us for inference benchmarking, and we sweep across so many different configurations and optimization types. Then what it creates is, and all the results are public and all the configurations are public, so now we have the Pareto optimal curve because a lot of times when people are comparing inference performance, they're taking a suboptimal curve or point for someone else and comparing it to their optimal one, it's like, well, yeah, I can stick, if I drove a Porsche versus some race car driver, obviously I'd drive it slower.

17:45Dylan Patel:It's the same thing with inference benchmarking. What we did is we created open source containers for the optimal points across every point on the interactivity, i.e. how fast is it responding to me versus batch size, i.e. how many users am I simultaneously serving curve. Now, anyone who wants the optimal point can just go to inferencex, download it, and run that as the optimal point, and they can check every day if they want, or they can even to auto download the most optimal point for that model, and their inference performance will be near peak. Is that curve like the most important curve in your opinion?

18:15The throughput interactivity curve is the most important one? Yeah, I think most things in hardware,

18:22Dylan Patel:infrastructure, model, application layer, everything is downstream of that curve, right? Is it something that needs to be super, super fast, super low latency, and I don't really care about the cost, so I make batch size very low, and I use techniques like speculative decoding or multi-token prediction heavily and there's so many possible techniques there? Or is it something where actually I'm batch processing a ton of documents and I don't really care about all these things. I don't use these techniques that actually are worse on cost efficiency to help you with speed for an individual user because I just want to pack a bunch of users.

18:52Dylan Patel:I don't care if the document takes all night to process. And right now, the way we treat AI infrastructure is it's like one size fits all. But over time, we're going to get to the point where there's stuff where you have batch workloads or you need instant response. and there's the whole curve that's going to matter for users. And so we see this with Anthropic, right? Cloud code fast mode costs way more than regular mode. And same with OpenAI's priority queue thing. Sorry, dumb question. How does cost factor into the start? So if I, let's say, imaginary example, I have a batch size of 100. Okay.

19:25Dylan Patel:And I can do 10 tokens per second per user. So in total, I'm doing 1 ,000 tokens per second off of that one piece of compute. That's one side of the curve, super slow, 10 tokens per second. Other side is I have 500 tokens per second, but I can only have one user. And so maybe 250 tokens per second, one user. And then there's points on the middle that are more freudal optimal. The average person actually wants like 50 or 100 tokens a second, and maybe the number of users I can batch together. So the curve is, okay, 1 ,000 tokens total per second, or 250 tokens total per second, depending on how many users I batch.

20:02Dylan Patel:and there's a curve in the middle. And so ultimately some workloads will actually want the 4X cost decrease because the same unit of hard work can do 1 ,000 versus 250. And some users, I'll pay 4X more because I don't care about the price. I care about time because the person using the tokens is expensive or the feedback loop that I have here is expensive. If you had to guess, you choose the timeframe, 10 years or 15 years, what percent of inference compute do you think will happen in space? It can be 0%, 50%, 99%. This is a tough one. You choose the time frame, whatever time frame. So I think the non-consensus, or at least against SpaceX thing, I love SpaceX by the way, and I totally would buy the IPO if I could buy stocks.

20:47Dylan Patel:Non-investment advice. Non-investment advice. Thank you. Thank you. Not investment advice. From Sequoia there. I don't think that space data centers will really matter in the next three to five years. With that said, I think in 20 years, I think the vast majority of compute will be going in space. And so the real factor there is sort of, what's the cost? It's the time frame. It's the time frame. It's the cost of building power on terrestrial land and how much power are you going to be able to do on terrestrial land? And I think obviously my views of where inference, how many gigawatts or terawatts are devoted to inference is, it's a crazy curve for me personally.

Read the full transcript

21:25Dylan Patel:What's your forecast? How many gigawatts are? Yeah, I think by 2030, just open-air and Anthropic will have over 100 gigawatts combined. And then you'll add Meta and Google and so on and so on and so forth. It's a humongous amount of compute that will be dedicated to inference. And by 2040, it'll be terawatts, right? The curve of productivity that we're going to get and so inference deployments is going to be huge. And so if you look at 2040, I think probably more than half of the incremental compute will be going in space. But if you look at 2030, I think it's sub-1%. Do you think intelligence per watt has been increasing?

21:58And then it seems like there's still a giant gap between where we are intelligence per watt versus human biology. And so do you think we are to close that gap? And if so, where is that game going to come from?

22:09Dylan Patel:Yeah, I think it often depends on what you're doing too. Like a TI-84 is way more intelligence per watt in terms of doing math than us. That's like 30 years old, right? So obviously, this is like a dumb sort of General intelligence. Yeah, but general intelligence wise, so one of the things inference x does is we also measure the power and cost of all of this hardware. We offer not just throughput versus interactivity, we offer cost versus interactivity. We offer power versus interactivity. So as far as has intelligence per watt been increasing, I mentioned it's been a 60x cost decrease for same benchmark level.

22:43Dylan Patel:We've also seen the same on intelligence per watt. It's not been exactly 60x, it's been closer to 40x. Some of the efficiencies are non-power ways, but there's been a humongous improvement in intelligence per watt on an annual basis, at least so far this year, last year, year before, year before. And I expect that to continue. As far as where we are from the human brain, we're many orders of magnitude away. Thankfully, it doesn't really matter. We can devote a lot of power to computers, much easier to power computers than human brains. We have sickness, disease, and food preferences. Sleep. Exactly.

23:18Let me just ask one more question on the general theme. In my opinion, in terms of intelligence per watt or intelligence per dollar, any of these metrics, I think there's three levels of input. You can get hardware improvements where the hardware is more efficient. You can get low level systems optimizations, like kernel level improvements, major multiplication libraries, things like that. Or you can get high level, model level algorithmic improvements at the highest level. To me, it seems like in the last three years, most of the gains have come from hardware level and some from the model level.

24:09Do you agree with that? Do you think that's what it'll look like in the future? Do you think there's a bunch of juice to squeeze in the kernel level.

24:17Dylan Patel:Yeah, Sean, I completely disagree with you, by the way. Great, that's why I'm asking the question. Okay, so I think one way is to look at it as these three different layers. And in that sense, like, okay, from Hopper to Blackwell, which is all we've had over the last three years, roughly 30x improvement on DeepSeq, on the most optimized deployment, which is, you can see on inference x, there's about a 30x improvement. But over the last three years, we've had way more improvement intelligence per watt. A lot of that coming from the model layer, right? If you look back three years, it's GPT-4. Now it's like, you know, maybe like Quen, one of the smaller Quen models that's like, you know, 27B parameters total and like 2 billion active is like way better.

24:58Dylan Patel:And so you've got this huge improvement on model layer. You've got this pretty sizable improvement on hardware, but it's that co-design layer. And I think that's what's important, right? If you look at the architecture of, you know, any of these models, but DeepSeq is the most famous one, at least that's public and people have seen it. Yeah, DeepSeek got huge efficiency gains from like co-optimization or kernel level, optimizing memory. Yes, I think it's like kernels, of course, but it's actually you build the hardware architecture for the chip. So if you look at the shapes of all the experts in DeepSeek, V3, they were all optimized for Hopper.

25:31Dylan Patel:And if you look at for V4, they're optimized for Blackwell and Huawei's chip. And what's interesting is despite the fact that TPUs are objectively an amazing chip, and they run all of DeepMind and they do all the training for Anthropic as well, on the pre-training side at least. TPUs suck at running DeepSeek, but they are really, really great at running other kinds of models that don't run well on NVIDIA. There is some level of such deep optimization that has been done, whether it be shapes, network I.O., patterns, how you do the collectives, how you do things around the arithmetic intensity of the attention mechanism.

26:05Dylan Patel:all these different things are co-optimized between the model and the hardware and the infrasoftware in between. And it's hard to say you can disentangle the games there. Do you think that my understanding is that China has done this a lot better than the West the last few years? And the DeepSea was one of the first models to really do this? I don't necessarily think so. I think it's more so that the West doesn't tell people what they do. Right. Like OpenAI didn't tell people that, you know, GP40 was how sparse it was, what the shape size was, all these things. But GP40 is roughly the same size, slightly smaller than DeepSeq V3.

26:44Dylan Patel:And 4.0 came out, you know, a little bit earlier, right, if I recall correctly. So is your view that like all three of these things have been happening simultaneously at like roughly the same rate and the most, the biggest gains are when you just co-optimize? Yeah, I would say there's been more gains on the model layer than on the software infrastructure layer and the hardware layer. But there's been innovations on every layer, and really the biggest gain and the beauty of the best labs is when they co-optimize all three. And that's what, like, you know, when Anthropic is, you know, even though they use many different kinds of hardware, they don't really inference too much on TPUs.

27:20Dylan Patel:They mostly train on TPUs. And they inference a lot on Tranium and GPUs. And GPUs are more the jack-of-all-trades. They've optimized their hardware, they've optimized their model, they've optimized everything so they can do that. Whereas OpenAI, prior models were optimized for Hopper more, now they're more optimized for Blackwell. And you step forward through time, these labs, and the same with Google, right? They've optimized, Gemini 2 was really optimized for the TPU V6E, or Gemini 3 was, and then Gemini, the next Gemini that's coming out is really optimized for TPU V7. and so sort of like a lot of these things are being co-optimized and actually when you pull that model and put it running on the old hardware it's really not that great and so I think a lot of this co-optimization is is the most important thing it's called software hardware co-design and that's what's like really exciting about like you know sort of what what you know I think my day-to-day is like you know great you get to look at one layer there's all these innovations happening here there's all these innovations happening on every layer but the real breakthrough innovation is when you leapfrog a few layers, you co-optimize and co-design them, and now all of a sudden, you've taken what could have been a 2x here, 2x here, 2x here, and instead of being multiplicative to 8x, it's actually 100x because you've optimized it across all three layers.

28:34Dylan Patel:That's what's really exciting about what you see at the labs, what you see at a company like NVIDIA, who's not co-optimizing on the model layer per se, but a little bit from the model layer all the way downstream to silicon, or you look at a company like TSMC, they're co-optimizing not just fabrication, but all the way from the components and the consumables and the tools, all the way upstream to what the designs, their chips, or the customers are telling them, is this co-optimization across many layers of the abstraction stack. There will always be bottlenecks somewhere in that optimization that they're lagging behind and then need to get pulled forward.

29:08Dylan Patel:And band-aids to cover that up. If you had to predict at any level of the stack, it can be literally anywhere, what are some of the bottlenecks you're tracking most acutely the next year? And not necessarily in the supply chain, not in scale, but in terms of the actual... And it can be in the supply chain too, but just is it memory improvements? Is it that just scaling loss? So memory is an easy one that everyone's talked about, but I'm not going to talk about it from a supply chain angle. I'm talking about it from a technology angle, right? Memory capacity and bandwidth have been improving very slowly.

29:48Dylan Patel:The NAND cell was invented like 25 years ago. The DRAM cell was invented like 40 years ago. And there's been no major breakthrough in cell, like what a NAND cell is. Obviously, NAND is like a very simple gate or DRAM cell. There is stuff that could come down the pipeline that could be hugely innovative. But even over the last five years, all we've really done is make the HBM more stacks faster. But actually, there's new innovations coming in the next few years where instead of stacking the HBM separately from the chip, you stack the memory directly on the chip and that makes your bandwidth explode.

30:20Dylan Patel:And so there's interesting companies in that space and interesting POCs that companies are trying to do there. I think memory bandwidth is one of the biggest. Another one is for the history of silicon, basically for the last two decades at least, how many watts a chip is can be easily predicted just by looking at it for for a data center or desktop chip, it peaks up at one watt per millimeter squared. And so if a chip is 100 millimeters squared, generally the power consumption's around 100 or a little bit less. And if you look at the newest NVIDIA silicon, the newest TPU silicon, it's still on that range of one watt per millimeter squared.

30:55Dylan Patel:So chips are now getting to 1400 watts, next generation is 2000 watts for NVIDIA with Rubin and such, and you move forward to Rubin Ultra, it's gonna be like 4000 watts or something like that. But really, there's increasing the amount of silicon. What's exciting is we're now finally doing things, and it's in development right now, where you actually can pump the amount of power into the silicon to be way more, more than one watt per millimeter squared. And now that all of a sudden means you need less silicon. Obviously, it's running at higher power and it's less efficient in some cases. But you reduce the amount of silicon and you're able to overcome certain problems.

31:29Without running at thermal issues?

31:31Dylan Patel:Thermal issues. issues, there's electrical interference issues, there's all sorts of different issues that crop up and that's why it's a hard engineering problem, that's why we've stuck at about one. But what's exciting is the world is trying to change these things. I think interesting in a different part of the supply chain is people talk about energy is hard and we have energy bottlenecks and it's like, yeah, but there's actually very simple solutions one could think of. take the millions of diesel engines for trucks that the US has the capacity to make. You can very trivially convert them to be using for gas in the assembly line, and then stick them up to an electrical motor, like back driving it so the electrical motor generates electricity, rather than the electrical motor causing the rotation of the wheel, for example, but doing it the opposite direction.

32:20Dylan Patel:Now you've generated electricity by pumping gas into something that US can make millions of. And then, okay, well, that sounds like a pain in the ass to service, right? Because now you have to have hundreds of these on a data center site. Well, actually, you can just pull people out of car mechanic shops and have them run around and repair truck engines. Actually, it's actually pretty trivial. I don't want to say it's trivial. I couldn't do it. I think you're making a really good point, which is that because the West wasn't really thinking about semi-controversy even hardware more broadly the last 20-30 years we didn't have like much innovation we'd have the best minds like thinking about how do you improve these why would you why would you want to go work in hardware when you can uh you can make ads still ads yeah exactly um okay i'm dying to ask nvidia versus uh tpu what are your thoughts um i think i think like everyone wants to pick one or the other for this but it's really like a function of like look you know you look two years from now google's gonna make 10 plus million tpus and through their supply chain and NVIDIA is going to make, you know, many more, a million, tens of millions of GPUs and both are going to be a hundred plus billion dollar, you know, well, Google's going to be a hundred plus billion dollars, you know, of TPU created a year and NVIDIA will be, you know, 500 plus or, you know, whatever it is.

33:32Dylan Patel:I'm not making a specific estimate. It's not a revenue forecast. This is just a thought experiment. Our research has that. Sean, you've been media trained so well. Absolutely, you know, getting ready for the SpaceX IPO. Are you guys big in SpaceX? Yes. Okay, so that makes sense. We're very lucky to be very large investors. Awesome, awesome. So I would say the case of Google TPUs versus NVIDIA GPUs, they both have points that are really in their favor. NVIDIA will be like, oh, we have switches, and we're general purpose. And TPUs will be like, well, we're more optimized. We're actually more energy efficient.

34:06Dylan Patel:And our network is actually more optimized for certain types of network architectures. And so you have these counterpoints that both would really get into. And I could, with a straight face, argue with you that GPUs are way better than TPUs, or TPUs are way better than GPUs. But it comes down to hardware software co-design. So actually, the way OpenAI's models are headed, it would be a terrible decision for them to use TPUs, potentially. And the way that Anthropic and Google's models are headed, it's actually a terrible decision, potentially, for them to train with GPUs. I mean, it'd be fun for them to train.

34:36What's the fundamental difference there?

34:37Dylan Patel:There's various things, right? Like the size of the matrix multiply unit is different, as a very simple thing. And therefore, the shape of the matrix multiply you do, the attention mechanism you use, the way that attention mechanism is structured, the way the experts are structured. So OpenAI and Anthropic are converging to very different model architectures? I think they have quite different model architectures, in fact. Interesting. OpenAI's are much more sparse, and that has benefits. And then Anthropics are still sparse, but more dense in general, and that has different benefits. And there's many other things, right?

35:08Dylan Patel:The network topology, right? Nvidia, all of their chips are connected to switches, NVLink switches. For Google, they have no switch. But what they've done is they've been able to do, you know, Nvidia, the NVLink can only connect 72 GPUs. For Google, their ICI can connect 8 ,000 chips at super high bandwidth, but you have to pass through other chips to get there because there's no switch. And so there's like, there's trade-offs there, there's positives and negatives, and that influences the model architecture. It's not necessarily that you should claim one is better than the other because at the end of the day, how do you say that this is better than that when you can't measure them in isolation because it also extends up to the model layer?

35:49But I remember for a long time thinking, one, the programmability of NVIDIA and then just CUDA as such a big moat. It seems to me that narrative has kind of changed, at least in my mind for the last three or six months. model companies no longer care about if we have to write custom kernels for this other chip, so be it. We'll work with four or five chips if we have to. Claude and Codex are actually quite good at doing a lot of that optimization work. And so it seems like some of the, and then it's not like there's 10 ,000 model companies that each need programmability. There's on the order of tens, maybe, model companies.

36:23And so it seems to me that the fundamental premise of tens of thousands of big customers that need CUDA compatibility, it seems that kind of thesis is changing.

36:34Dylan Patel:Yeah, I mean, certainly the CUDA mode and software mode is at least partially disentangled because models are just great at coding and all software gets commoditized in that case. I do think there is some level of open source and what people call the CUDA mode is not actually anything to do with CUDA, but it's like the fact that DeepSeek, Kimi, and Zipu.ai, and Alibaba, and Tencent, all these companies, Xiaomi had an awesome model recently, their models are a co-design for GPUs. And therefore, if I want to run them on TPUs, actually, in some cases, they don't run really well on TPUs. Now, Google just has to create their own open source model ecosystem or open source models themselves.

37:15Dylan Patel:So they have the Gemma models. And so you end up with like, well, that's not really CUDA as a moat. It's that the downstream product is more optimized for NVIDIA. And in these cases, these companies are open sourcing them, or like Nemo is just open sourcing it. And then the users of it, for example, the open, the inference API providers, the RL companies that are trying to take open models and customize them for companies' business use cases. All these different companies are downstream of the fact that like, okay, well, I guess I need to use NVIDIA because the ecosystem uses NVIDIA, even though I don't particularly care about writing CUDA kernels, because the models are great at that, but it's like the shape of like, well, this expert, the DMOD is this, and the hidden dimension, blah, blah, blah, is this, right?

37:55Dylan Patel:And so therefore, it's better to run on NVIDIA GPUs than it is on TPUs and vice versa, right? If Google were to actually open source really good models, you know, this would be the same thing, right? People would take their models and they'd be like, oh, wow, these don't run that well on NVIDIA GPUs. I should actually just rent TPUs or buy TPUs and do it on there. For small teams, you're going to want to use all the open source software like VLMSG, LANG, PyTorch, all that stuff. But the big labs, they don't necessarily need to use all that, right? OpenAI forked PyTorch long ago and Anthropic and all these other people don't necessarily rely heavily on the open source implementation of these things.

38:29Dylan Patel:They've forked things or built it on their own already. And so they don't need to rely on the open source. And therefore now it's more like, I'll choose the best hardware and I'll co-design my model and infrastructure software through and through for that hardware that is the best and most cost efficient. And I'll have AI help me write all that software. What do you think of Cerebrus? I think Cerberus is a really innovative company. I think in some spots of the market, they're really, really good. Very fast inference. I think that's a big market. We use fast mode almost exclusively at Semi Analysis.

39:02I love how disciplined you've been about accounting for, I don't know if that was one exhibit you did or if you do it consistently, but accounting for the dollar spent in the ROI on each task. It's an awesome analysis.

39:13Dylan Patel:Yeah, we do it pretty diligently. So thank you. That was the dark GDP article that we wrote. And also track everyone's token spend by day. And if someone's spiked up, I'm like, what did you do? It's like, okay, thank you for telling me that that seems worth it. Cool, on with my day. I think fast mode is obviously worth a lot for high-end tasks, right? I can just see so many different use cases where super fast tokens are worth it. I can also see the flip side where there's a lot of use cases where super fast tokens aren't needed and therefore the market won't pay for them and they'll use GPUs and TPUs instead.

39:45Dylan Patel:I think the big risk for Cerebris is I mostly think the best models are the ones that you want to use fast mode on and small models you necessarily might not use fast mode on. I could see that being wrong with financial markets maybe or something like that, like a Jane Street high frequency trading or something like that, or medium frequency trading. But ultimately, running really large models at really long context is very difficult on SRM-based chips like Cerebris, like Grok. And so now it all of a sudden is like, you know, what happens then if like the models get too big, right? If OpenAI's model is not, you know, on the order of, you know, hundreds of billions parameters or, you know, low trillion parameters, but it's actually 10 plus trillion parameters.

40:23Dylan Patel:Now, all of a sudden, I don't think that that will fit on Cerebris, right? And then if that doesn't with a long contacts length, right? If you have a million contacts length, now that makes it really difficult to justify, you know, and so far we've seen the bulk of revenue and usage at the labs be on their best model. Even when the model price has gone up, we've seen that. There's some data that shows that even though Fable just released today, they've had incredible amounts of people switch to Fable and Mythos, sort of that next tier model, even though it's way more expensive. And so... And that's volume by dollars totally, but without volume by tokens?

40:57Dylan Patel:Well, I guess who cares about volume by tokens? It's about the dollars. Fair enough. Right? If I don't care that there's, you know, I don't know, 200 ,000 Mini Coopers or Toyota Camry sold. If, you know, I don't know, 4.150s are 5X ASP and they sell only half as much. Okay, right enough. And therefore, the most lucrative market is pickup trust in America. Mostly being facetious, but like... I do think this is one of the things that you've done so well and differentiates you from almost everyone else is that you care so much about the economics in addition to the technology. And I think very few people bridged those two things well.

41:32Yeah, thank you.

41:33Dylan Patel:I think it's really fun inside of semi-analysis because we have 90 people and a big chunk of them are technologists, engineers across the whole supply chain. And then a big chunk is people who are formerly at hedge funds. And you see these arguments, people are like, oh, well, that doesn't matter. And then someone's like, well, but cost. And then the engineer's like, no, no, no, but this technology is the closest. You see this organically fight it out. And we're pretty informal. And given the fact that I was a forum moderator, you can imagine what this internal slack looks like. enjoying it. You don't wrestle with a pig because a pig enjoys it.

42:06Exactly. Just on this topic, before I go into the next question, are there like trigger topics in semis for you? You know, like if someone's like, which is like such a meme, you think this person must be a moron. Like if, you know, if it's like, oh, you like memory is the bottleneck.

42:26Dylan Patel:I mean, it's true, but like, I think moreover, the one that really gets me is people are like, AI has no ROI. It infuriates me, right? Like there's like, what's the ROI? Or like denying model progress, right? There's these people that are like, models aren't getting better. They're not reasoning. They can't think. They're going to dead end and plateau. And it's like, bro, the line has been up and to the right in terms of capabilities this entire time. And they're like, look, this benchmark didn't improve. That's because it said 90%. Look at the new benchmarks. Yeah, you saturated. Now they're skyrocketing, right?

42:58Dylan Patel:I think that's more so the issue and challenge. I think semis are really complex and I don't fault people for lacking understanding of it. I learn stuff every day about the semiconductor supply chain from people and I've been studying it for arguably 18 years since I started moderating the forms when I was 12. right? Like, you know, arguably been studying it for that long, but even then, like, and it's like live, breathe, and that's all I care about. But there's so many layers of the abstraction stack. It's like, like I learned about a new chemical that does like a hundred billion dollars of sales, like yesterday.

43:33Dylan Patel:And I'm like, whoa, didn't know this one existed and what process it did. And it's like, but it's like, you know, you learn about things all the time. It's like, okay, a hundred billion dollar sales in a, you know, a couple hundred billion dollar industries, whatever. But like, you know, it's like, but it's essential. It's essential. And it's like, actually every chip requires it. It's like, wow, I guess there are a thousand process steps. And it's like, oh yeah, you like semiconductors name every process step. It's like, no, come on. What I think is the most funny is when people have all the facts in front of them, and then they get the conclusion completely wrong.

44:01Dylan Patel:And that's - That happens in our job all the time too. Yeah. Yeah. I mean, I think my attitude is not to be mad that you do that. It's to do it as fast as possible. I think the industry, because it's so, it's just like AI is the most important thing in the world right now, and there's so many near-term bottlenecks, we talk a lot about the near-term. Are there longer-term things that you're really excited about? Like, say, on a 10-year time frame. We talked about orbital data centers, but like Silicon Tonics, you think they're underrated or overrated on a 10-year time frame? Are there other things that on a 10-year time frame - Yeah, I mean, I think on a space, I think space is like super crazy awesome in the 10-year time frame that I'm, you know, for space data centers and all these sort of mining asteroids and all these things, which is, you know, super excited about the vision of SpaceX, right?

44:44Dylan Patel:Again, not investment in advice, before you hop in. I think on the semiconductor side, tremendous market movements and tremendous things can happen just when things happen one year later or sooner. And so that's all technology that in terms of co-packaged optics, everyone knows it's going to happen by the end of the decade. The debate is 27, 28, 29, 2030, but some point along there, it's going to happen. I think the more interesting thing is there's companies like... Did you guys invest in Naveen Rouse company? We did. Okay. Yeah. So I think like he's trying to innovate on like the silicon layer on the software abstraction layer and the model layer simultaneously.

45:19Dylan Patel:And he fully understands that it's not a, like a, you know, we're going to do this in a few years. It's not a two-year timeframe. Yeah. It's not a few-year timeframe. It's a long-term bet. Um, and like stuff like that is like, okay, we're going to bring like potentially like analog compute with energy-based models and like all this crazy shit all at once. It's like, that's exciting. Probably won't work, but you know that's exciting and i i like really look forward to definitely work quickly yeah definitely will work quickly is what i should say i believe in david and like you know i i met him very you know i think he's one of the first people i met in the industry um funnily enough like in 2020 or 2021 um actually 2020 it says something about him i think he's someone in my experience he's always trying to i baited him on the internet i baited him on the internet he's always trying to help the younger generation.

46:06I think he's trying to identify talent. And he's so ahead of his time with Mosaic. I remember getting pictures. No, he was 2019.

46:12Dylan Patel:I was still anonymous then actually. I baited him on the internet and he started replying. And then I just took it to DMs and then took it to a call. And like, that was the first person who was like really important that I talked to in the entire semiconductor industry. That's funny. But yeah, sorry to interrupt. That's funny. What do you think is the end state of the ecosystem? Do you think every lab, every hyperscaler just has its own chips? Like, Trainium seems like it's now working, right? So do you think we end up with every lab, every hyperscale has own chips, at least for inference, and then maybe for training you go to NVIDIA or whoever?

46:42What do you think is the end state?

46:44Dylan Patel:I think everyone will try and they won't stop trying. I think ultimately, you know, supply chains matter. What technology you can bring in matters. And more and more as the industry gets bigger, supply chain diversification happens. You know, right now everyone's chip more or less looks the same. It's a big logic compute die in the center, and there's some HBM on the right and left and on the top and bottom. Top side is networking, and then the bottom side is PCIe and other I.O. And that is the exact same structure for Tranium, TPU, NVIDIA chips, and most of the startups. Not Grok and 3Risk, they're doing weird shit, but that's cool.

47:23Dylan Patel:I think as you step forward, we're going to get more bifurcation of hardware architecture and model architecture, and therefore people are going to co-optimize them. and some of them will end up in local minimas. If it's gradient descent, people are trying to go to the most optimized solution. Some people will race to a local minima, and then the question is, how do you scoot back over to the absolute minima? And to some extent, NVIDIA will always be more general purpose than anyone else's chip in general, at least on a parallel AI compute basis because they have so many customers who care about different things and who will always give them feedback in the design.

48:01Dylan Patel:The minima will always be better than them, but is that minima a local minima? Like is the TPU or Tranium or GroK or Cerebris or whoever's design optimized awesomely for here, but in the end state actually you got to go over here and so they're wrong. And maybe they make a great time, they're great for a little bit of time, but then they end up being wrong. It's like that's the real question. And so I think there will be a big market for general purpose AI compute because you talk to people at labs, they don't even know what architecture they're going to be doing in a year. like right like they literally don't know what architecture they're going to be doing in a year they have bets they have many research bets and and that's this exciting thing but they don't know where it's going generally they like know what hardware they have they're trying to co-optimize but ultimately like if a new breakthrough happens on model architecture it's like just replace the tension mechanism with something else right who knows or you know all of a sudden you know something happens the best hardware will change and therefore like are people going to make five-year investments on hardware solely on an ASIC that is more specialized, or are they going to have some bucket of more general-purpose compute?

49:05Dylan Patel:And so you see this with like, Google's paying$11 an hour per GPU to XAI for GPUs, right? Like, that's insane, right? That's a very high amount of, obviously, compute is limited and so on and so forth, but it's like very like insane, but at the same, you know, despite the fact that they have TPUs. And so there's like some questions there, why do they do that? Google actually has three different design programs for TPUs. They're making a TPU with Broadcom. That's a different architecture than a TPU with MediaTek. That's a different TPU than the architecture that I won't disclose by research. But they're making different architectures.

49:38Dylan Patel:It's not just like, oh, they're making TPUs with a couple vendors and it's the same architecture. It's different architectures. And the third one is a very different architecture from the first two. And so I think people recognize that the local minima can happen, and therefore, I think everyone will have their own ASIC program. I think everyone will deploy billions of dollars of their own ASICs, tens of billions of dollars. In the case of Google, hundreds of billions of dollars a year of their own ASICs. But ultimately, they're also going to have workloads that don't use TPUs. Some of the Google bets that are not Gemini DeepMind actually primarily use GPUs.

50:10Dylan Patel:They don't use TPUs. Some of them also primarily use TPUs. It's a bit of a broad thing, but maybe for drug discovery or for Waymo, you might not want to use TPUs. I don't say which one it is, but there's different architecture bets and different paths for AI. AI for science may have different algorithmic patterns than general intelligence AGI models. And so I think we'll see diversity continue to proliferate. And because the market has gotten so big, niches will be carved out. And so that makes it possible for companies to have their niche and actually make money, even if the majority of the pie goes to NVIDIA and TPU and training them.

50:46Okay, love that. Can we talk about the data center build out? Like one, it seems like, I mean, by all accounts, if you look at the charts, like dollars per computer, we are in the middle of a crazy compute crunch. And it seems like it's both a demand and supply side crunch, right? Like demand for long-rise agents skyrocketing, supply, all these data center build outs are delayed. Do you think this is we're in a compute crunch for the foreseeable future? Or do you think it alleviates at some point?

51:10Dylan Patel:Yeah, so every quarter, we're deploying vastly more compute than the prior quarter and there's more data centers built than the prior quarter. This year, there's going to be 20 gigawatts, even accounting for the delays. And next year, there's going to be more than 30 gigawatts accounting for the delays. Of course, delays happen on everything, right? Anything hardware can have a delay. That's just the reality of life. Are we going to have a compute crunch for the rest of our lives? It depends on what happens with models. But the TAM for Mythos, Mythos 5, Fable 5 is not just like 2x that of Opus.

51:43Dylan Patel:The model is so much better and it can do so many more tasks that the Tamford is way larger than that. And yet compute in the world did not double in the last six months. From Opus or maybe like seven or eight months since Opus 4.5 would launch to now, huge 4.6, 4.7, 4.8 were improvements, but Fable and Mythos were like a huge step function improvement. The world's compute did not double in that or quadruple or whatever in that same time frame. But the demand for useful tasks that can be done by AI, the number of useful tasks and the value of them that can be done by AI has. And so now the question is, what happens?

52:18Dylan Patel:Well, obviously Anthropic in Q2 is profitable. They're net income profitable, excluding stock-based compensation. And I think by Q3, they may even be profitable, including stock-based compensation. That's how profitable they're getting. And their margins on an Opus token, at least Opus 4.8, token is like north of 80 % for the API price. They've got a lot of deals where their total corporate gross margins gets clawed down a little bit because of like how they do bedrock deals and vertex deals and things like that. But ultimately, their per token margin is so high. Well, then if you don't have the capability, they have the capability to pay.

52:57Dylan Patel:Ultimately, every GPU they buy at above market rate, you know, they also bought GPUs at above market rate from SpaceX, which is below the rate of Google, but that's because they signed earlier. It's something that other companies, maybe a venture-backed company or a company that's not really got positive margins, can't necessarily do. What is the cost-benefit ratio? It's like every GPU I rent, because I'm out of compute capacity, I can immediately turn around and sell tokens on it, or every TP or every trinium. I can immediately sell tokens on it at a positive margin. If I'm running 75 % gross margin and I double the cost of the compute, it's fine.

53:30Dylan Patel:I'm still running 50 % gross margin. Spinning up more compute nodes is not really necessarily a human requiring task for them if they're renting them. And so ultimately it's like, well, my NOI still goes up. And so I'm going to rent GPUs at whatever price, at some level, whatever price I can pay. I have almost a reverse question of like, at some point, does this compute build out go bump at night? Earlier today, I think there was a tweet like Crusoe publicly said one of their customers had asked to halt construction on one of the data center build outs. It seems like everybody in the ecosystem is so levered right now to like, we got to build, we got to build, we got to build.

54:03High leverage, high growth to me is like, makes me very, very nervous as an investor.

54:08Dylan Patel:Like, hold on. High leverage, high growth means small amount of equity has huge upside. You're not a debt investor. You're an equity investor. You got to go to the school of private equity. Levered buyouts only. I actually come from a school of private equity. Oh, awesome. She forgot the school. She's been a VC for too long. No, I just do revenue multiples. Do you see any signs of that? Are you worried about that? I see what you mean, right? And that sort of goes back to the model point, right? Obviously, if the model's expanding the total economic valuable work, that's sort of the dark GDP report that we did and you mentioned earlier.

54:46Dylan Patel:If the work that these models can do does not expand faster than the compute capacity, then that tide turns, right? Over the last six months, that tide has been very much levered in this direction of the models can do more work or are expanding their TAM of work they can do faster than the compute is increasing, and so prices go up. It's very possible that all of a sudden, model progress stops. You talk to anyone at Anthropic or OpenAI, maybe they're drinking the Kool-Aid, but you talk to basically all of them, they're like, no, no, no, model progress still go up. So ultimately, current methods could stall somewhere.

55:23Dylan Patel:I'm not sure where that would be. It seems like we have line of sight to model improvement, rapid model improvement. And in fact, models are improving faster than they were six months ago or a year ago because there's, I would call it recursive self-improvement, but basically the models are helping write all the info and launch the next model sooner and sooner and sooner. So you've got this pseudo recursive self-improvement loop going. And so the models are getting better and better and better faster. But ultimately, capital is a big problem, which is why Google raised capital, you know, they've got a godly amount of SpaceX, right?

55:56Dylan Patel:They own like 5 % of the company. I think a little more, but yeah. I think at one point they had like 10%. Larry Page invested a billion dollars at a$10 billion valuation, got 10 % of the company. It got diluted, like all this, but that was one of the greatest investments of all time. Good job, Larry. Yeah. So they know they have like a hundred billion dollars in the bank that they can sell in nine months or whatever from the lockup. And they have all the gross profit they do. And yet They still modeled that and they're like, we need to raise capital. And so they did an offering and it's like, that's insane.

56:25Dylan Patel:So that tells you how much they think they need to spend. But capital is like really, you know, you know, Meta did announce that they're going to do a raise, stock tanked, people don't like it. But, you know, that's all these companies are going to raise capital, whether it be debt or equity. At some point, money spigots will have to, you know, slow down. But right now, every GPU that Amazon adds, they're making higher revenue or every TP or Tranium, whoever anyone adds, is making gross profit. I'm going to do a little bit of a tee up on this to turn it into a question for you. But as we talk about this, for me, the thing that's going through my head is almost an alternative hypothesis for the Crusoe example.

57:06I'm using an analogy in oil. In oil, Saudi Arabia has way lower costs per barrel to produce oil than a lot other countries. There's also the purity of the oil. A lot of Saudi has generally very low contaminants in their oil, which makes the refining easier, all of this. The question for me is when you look at for every gigawatt that's being put in the ground, I've called the 20 gigawatts coming online today. How much homogeneity do you see in those gigawatts? Is it something like, and you can tell me whatever metric you think is right, but are Are Google's gigawatts two times more valuable than, say, most NeoClouds because they have optical switches and they have like they've been doing it for a long time and like they know how to do power smoothing?

57:55Because I think this could be the alternative hypothesis that some of the people that are it's like the people that are good at building data centers, they should they should just do it to the max because there's so much demand and there's so much better than it. but then maybe we're starting to see the early signs of the people that are not as good at it kind of getting hit a little. So I don't know the reality. I'm just curious how you think about that.

58:18Dylan Patel:So far, there are metrics for this, right? So Tranium sells at sub$10 billion per gigawatt rental rate to Anthropic and to OpenAI. GPUs, at least before the craziness of the last six months, usually went around$12 to$13 billion per gigawatt. So the rental rate, and this is from a neocloud versus Amazon even. And now when Amazon sells GPUs, they'd also be 13 or so. And my understanding of that also is that those numbers, Amazon subsidized that a little bit. So that it's like, I actually think the numbers were even, I think the disparity is even more. It's less than 10, it's less than 10, but there's like some weird, basically how much.

58:56I think there's, yeah. And like, look, my understanding, obviously, Anthropic played a big role in making training useful in terms of writing all the libraries, et cetera. And so everything I hear is that Trainium's really freaking good hardware, and it's getting way better. And obviously Anthropik's now using it a lot. So hopefully we would see that price go up per year long.

59:20Dylan Patel:The deal they did was actually there was a floor mechanism, and if it didn't do well, it would be cheaper and then to the point where it's cancelable. And if it did really well, the price is kind of higher. but effectively less than 10, right, is where Tranium shakes out at. Whereas GPUs, I mean, the SpaceX deal, again, was like 25 or something crazy, billion dollars per gigawatt or$25 million per megawatt, right, a year rental rate with Google. I was like, that's a crazy divergence. Now, obviously, if Amazon was selling Tranium today, it'd probably be more expensive than 10 because of the compute shortages.

59:53Dylan Patel:But you do see this already in the sense of with data centers, oftentimes the rental price of a data center if you're doing co-location right not compute in there but just power here's the data center you you price it generally on a dollars per kilowatt per month and so they used to be 60 per kilowatt hour per month and now you see things transacting at anywhere from like 120 to 160 but different quality data centers this actually you've i've seen data centers go as high as 200 when the customer is not such a great credit rating and And then the data center is a pretty good one. And I've seen stuff go as low as 100 still, or in India go as low as 80 because the grid's not reliable, the internet connection's not great, and it's a pretty mid data center, but at least it's a data center.

1:00:37Dylan Patel:And so you see this huge discrepancy there already. In the case of data center construction, usually the pitfalls, they just fail. There's a lot of people who claim they're like four guys. They're like, yeah, I bought some turbines. I put the money down for them. I'm going to build a data center. and then they get delayed, delayed, delayed and fail. So you have to like probability weight, time weight, time lag, the teams that suck versus don't. And sort of, you know, our data center model does that. We kind of track every data center and try and do this for every single one based on, you know, equipment that they're using and all these things.

1:01:10Dylan Patel:One of the things you mentioned about Google is, you know, in a gigawatt data center, they actually put like 1.5 gigawatts of hardware. And because they have such understanding all the way from workload to, you know, they're able to slosh the power around. And so instead of, you know, constantly a gigawatt of compute, which typically runs at 60 % or 70 % utilization in terms of power consumption, not utilization of the hardware. Someone's always renting it. They're now running it at that 60 % to 70 % means it's at a gigawatt, and you're using the full gigawatt. You see people doing deals with, including Google, with utilities where they're like, oh, well, I know this grid can sustainably take a gigawatt, but except for three days of the year, you can actually do two gigawatts.

1:01:50Dylan Patel:So give me two gigawatts, and then just tell me to turn off. And so they'll do that. And so these sorts of tricks, and then you need to have supreme management of workload, backup power, all these things, generators on site to figure out how to actually keep it two gigawatts sustainably. When people do this, they're able to charge more. Whether it be I'm actually selling two gigawatts despite only having one gigawatt because those three days I'm able to deal with via battery, gas, et cetera. Or I figured out how to build power on site. Now I have a gigawatt where no one else does. And so I'm able to do it quickly.

1:02:20Dylan Patel:It's not necessarily transacting for a higher price. it's that I'm selling more gigawatts. And sometimes there are levers where you're selling more gigawatts, where each gigawatt is selling at a different price. I think it's more on the data center and energy layer. It's more about just having it versus not, and then that being delayed or not. It's more binary. But on the compute side, I do think there's a lot more interesting work there. A gigawatt given to Anthropic is objectively worth more revenue than a gigawatt given to OpenAI. And it seems that both of them could sell every gigawatt that that they have right now, given rate limit problems and token max limit and all these sorts of things that open the Ananthropic, especially since Codex 5.5 came out, it's much better.

1:02:59Dylan Patel:And then likewise, if you gave a gigawatt to SpaceX, they'd turn it on. My guess, my suspicion is that they probably make better use of the hardware than most people. Just like, I think people underestimate how much networking experience they have from Starlink in particular and also how much there's like power management experience they have via from tesla yeah people like brett mayo are like incredible like they're good yeah and so i think for me that's actually i think probably the thing that might i don't actually know the answer but i think that might be missing from the analysis a lot of people i think i think it's also the fact that when core weave builds a gigawatt even though their gpu compute is objectively better than amazon or google or microsoft's in terms of performance we've tested the performance and then reliability.

1:03:50Dylan Patel:The problem is Google sells it six months before they have it up and they need to turn around and take that paper that they signed to get debt with that credit backing and then turn around so they can actually pay for the PO that they've already issued, you know, for the order that they've already issued. Whereas SpaceX was like, no, no, no, this is running now. Buy it, right? And it's a big discrepancy when you have a balance sheet to do that versus not. And that also helps your revenue per megawatt like be much higher. Why does the NeoCloud opportunity even exist? Because if you had asked me five years ago, I would have said the hyperscalers are going to own this.

1:04:21And you mentioned just now core weight has better performance than the hyperscalers. Why does this opportunity exist maybe at the macro level and then in the execution level?

1:04:29Dylan Patel:Yeah. So in 2023, I wrote a report that had Amazon really hate me. It was called Amazon Cloud Crisis. So I talked about how Amazon was the best cloud because they had their Nitro NICs, which offered like tenant isolation. All the hypervisor ran on the NIC and then you could sell all the cores and they had custom SSDs that they made and they'd buy the raw NAND and they'd have lower cost because they'd buy the raw NAND and build their own SSDs. And they had their custom Graviton CPUs and that drove down cost per core. And so they had all these things that enabled them to sell more cores, have better security, good networking.

1:05:03Dylan Patel:But this was all for the traditional CPU, better storage, for the traditional cloud world. But in the AI cloud, a lot of this stuff hurt performance. These Nitro Nix were bad for performance, still are worse performance, although they've caught up a lot because they've had a couple iterations to improve them, but they're still worse for performance. A lot of the security stuff doesn't matter because it's not like I'm time-splicing users or splicing a socket into many users. It's like no one rents a single GPU in an 8GPU server. No one rents a single GPU in a 72GPU rack. They rent the whole rack, and in fact, they rent many of the racks.

1:05:36Dylan Patel:And then there's no like, oh, I rent for six hours and I give it back. It's everyone has these long-term contracts. So the mechanics of the GPU rental market meant that a lot of the expertise of the hyperscalers fell away, and a lot of the expertise that they did have were actually, some of them were detrimental, right? Network performance. For Google and Amazon, they had custom networks that were better for traditional CPU and for the stuff that they were doing, but actually worked for AI. And then in other cases, it's like, well, you know, Microsoft would save money by building their own data centers, but their data center teams were not actually that great, and so when it came time to run, you know, When it was predictable building, it was fine.

1:06:14Dylan Patel:When it came time to actually double your forecast for the year, it's like they fell on their face, and they had to go get a bunch of neocloud capacity. I think so performance, I think time to market's another one. These massive organizations, no one's getting rich from building this data center faster. But you look at Crusoe, for example, Chase, and all the other people at the team. I was going to name some people at the team, but I'd rather not. All these people are getting rich if they fucking deliver this compute faster. They're hyper-levered equity owners. Hey, look, they're also all coming from Bitcoin.

1:06:50And you're not supposed to say that.

1:06:52Dylan Patel:I mean, a lot of the data center, like their main data center guy came from Microsoft. I don't know. I'm just teasing. But it's like you learn a lot when you're in a very high fluctuation market. How much do you think was Jensen playing 4D chess? Jensen absolutely hates a world where all the hyperscalers have all the power. There's a reason he's blowing money on random AI labs that I don't even know if it makes sense to, but he's blowing money and pumping them up and going to everyone around the world and saying, you should invest in this company because he wants to create a multipolar world. That's why he loves Chinese labs because he wants to create a multipolar world.

1:07:28Dylan Patel:A world where open-anthropic and Google models are the only models is one in which he's screwed. A world in which the hyperscalers are the only ones building compute is one he's screwed in. And so, of course, he needs to point the allocation gun at NeoClouds, help backstop their clusters, do anything and everything. Because while today a GPU sold to Crusoe and a GPU sold to CoreWeave and a GPU sold to Google and Amazon are all the same price for him, five years from now, Crusoe and CoreWeave existing means Google TPU will be weaker and means Amazon Tranium will be weaker and more inference being done with non-closed source model labs is better for him.

1:08:09Dylan Patel:So I think the NeoCloud ecosystem is, these people that are Wild West, these Neo Labs as well, a lot of them have investments from NVIDIA. It's the Wild West. Some will fail, many will fail, but some will emerge as really great teams. Whether it be, oddly, Crusoe, who's a bunch of crypto guys who then started building data centers and doing flare gas stuff, or CoreWeave, who initially was a bunch of New York hedge fund guys but then they they like built you know there were a lot of people who didn't bubble up like them started around the same time just failed right and so i think you know i gotta say both those teams are phenomenal yeah there's a lot of credit and it's like that's your point but yeah i mean my point is like he he you know you throw it's like you throw a bunch of like bait into the water and the best fish will figure out and survive right um and sort of the same way with the neoclouds and and and he hopes the neolabs as well we'll see if any of the neolabs really bubble up, but like, you know, Thinking Machines has a few hundred million dollars of ARR, right?

1:09:02Dylan Patel:That's pretty impressive, even though they've had, you know, in the media, it's like, oh, they've lost all this talent. It's like, well, but Tinker is doing a few hundred million dollars of ARR. Like, that's pretty impressive for out of the gate, a product that's less than six months old or whatever. And, you know, we hope the same happens to other Neo Labs. And so, you know, he wants a multipolar world. Truly, congratulations on the success. Thank you. Thank you. Just the last thing I'll say is, I've seen a little bit of this. I think the public, they can probably tell from listening to you how hard you work, but it's clear you've just been working your ass off for more than a decade, and it led to the last few years of being in the right place, right time.

1:09:38But it's unbelievable what you've accomplished, and I know it's just the beginning. Thank you so much. Thank you for doing this. Awesome.

1:09:57Thank you.

From the publisher

Dylan Patel, founder of SemiAnalysis, argues the biggest gains in AI don't come from faster chips, they come from software-hardware co-design. Optimizing the model, the kernels, and the silicon together turns a 2x here and a 2x there into 100x. He explains why DeepSeek's experts were shaped for Nvidia's Hopper (and why TPUs struggle to run it), why OpenAI's sparser models and Anthropic's denser ones pull them toward different hardware, and why the so-called CUDA moat was never really about CUDA. Dylan breaks down InferenceX, his living benchmark that runs the latest models on over $50M of donated hardware daily, tracking a roughly 60x annual drop in cost per unit of quality. He makes the case that inference will be a bigger market than oil, that the compute crunch persists because models expand the value of useful work faster than compute grows, and why Jensen Huang is bankrolling neoclouds to engineer a multipolar world.

Hosted by Shaun Maguire and Sonya Huang, Sequoia Capital

More from Training Data

All 110 episodes
Why Hardware-Software Co-Design Is AI's Real 100x: Dylan Patel of SemiAnalysisTraining Data · 1 h 10 min
Listen in VO