In short
Dwarkesh Podcast - Episode Notes
Episode Overview Title: Dylan Patel & Jon (Asianometry) – How the Semiconductor Industry Actually Works Description: A deep dive into the semiconductor industry and the trajectory towards hardware scaling for Artificial General Intelligence (AGI) by the end of the decade. Guests:
- Dylan Patel - Founder of [Semianalysis](https://www.semianalysis.com/), a leading publication on AI hardware.
- Jon Y - Creator of [Asianometry](https://www.youtube.com/@Asianometry), a popular YouTube channel focused on semiconductors and business history.
Key Themes and Discussions
- Semiconductor Industry Insights
- Espionage in Semiconductor Development
- Discussion on the potential espionage activities by nations like China to catch up in semiconductor technology.
- Insights into what information or blueprints from leading companies (like TSMC, Nvidia) would be most valuable.
- China's Semiconductor Strategy
- The role of centralization in computing resources in China as a strategy to accelerate advancements in AI.
- The impact of sanctions on the semiconductor industry and how China could potentially circumvent these limitations.
- Economic Dynamics
- Market Predictions
- Discussion on the anticipated $1 trillion+ investment in data centers by the end of the decade.
- The focus on the economic viability of semiconductor production and the implications of Moore's Law.
- The stratification of the semiconductor industry and why it remains so.
- Technological Advancements
- Challenges in Semiconductor Manufacturing
- Exploration of the complexity of semiconductor manufacturing and the challenges in maintaining efficiency and innovation.
- The need for advanced research in memory technologies and how it affects overall performance in the chip industry.
- AGI and Hardware Scaling
- The potential shift in AI training paradigms and how this could impact the demand for different types of chips.
- Discussion on the evolving architectures of AI models and the implications of different hardware designs.
- Future Predictions
- Scaling to AGI
- Predictions for the future state of AI models and their training requirements by the year 2028.
- The possibility of achieving one exaflop (1e30) performance in AI training, and the challenges that accompany this goal.
- Personal Stories and Insights
- Journey of the Guests
- Dylan's Background: Transition from a hobbyist and Reddit poster to running a consulting firm focused on semiconductor research.
- Jon's Journey: From a YouTube creator with a focus on semiconductors to engaging with a broader audience on the historical context of technology.
Timestamps of Key Discussions
- 00:00:00 - Xi's Path to AGI
- 00:04:20 - Liang Mong Song
- 00:08:25 - How Semiconductors Get Better
- 00:11:16 - China Can Centralize Compute
- 00:18:50 - Export Controls & Sanctions
- 00:32:51 - Huawei's Intense Culture
- 00:49:21 - Mind-Boggling Complexity of Semiconductors
- 01:04:36 - Architectures Lead to Different AI Models
- 01:16:24 - Scaling Costs and Power Demand
- 01:37:05 - Are We Financing an AI Bubble?
- 01:50:20 - Starting Asianometry and Semianalysis
- 02:06:10 - Opportunities in the Semiconductor Stack
Conclusion This episode provides in-depth insights into the semiconductor industry, discussing both current dynamics and future possibilities. It highlights the complexities of manufacturing, the role of geopolitical factors, and the journey of two individuals who have made significant contributions to understanding and communicating the intricacies of hardware and AI. The conversations underscore the importance of innovation, investment, and strategic direction in the rapidly evolving tech landscape.
For further details and to listen to the full episode, visit [Dwarkesh Podcast](https://www.dwarkesh.com).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Today I'm chatting with Dylan Patel who runs Semianalisis and John who runs the Asian amateur YouTube channel. Does he have a last name? No I do not. No I'm just kidding. John why? What is it? I'm John why? Wait why is it only one letter? Because why is the best letter? Why is your face covered? Why not? Because seriously why is it covered? Because I'm afraid of looking myself get older and fatter over the years. Oh my god. But Mrs. Here say it's like anonymity, right? Anonymity. Okay. By the way, so you know what Dylan's middle name is? Actually, no. I don't know what he told me, but what's my father's name?
0:40I'm not gonna say it, but I remember. You could say it, it's fine. Sanjay? Yes. What's his middle name? Sanjay? That's right. Wow. So I'm the War Cache Sanjay Patel. He's Dylan Sanjay Patel. It's like literally my white name. Yeah. It's unfortunate my parents decided between my older brother and me to give me a white name and I could have been dark -ish Like you know amazing it would have been if we had the same day Like butterfly effect at all that we probably would have all would have turned out the same way But like maybe it would have been even closer. We would have met each other sooner, you know Warcastle and different tell in the world.
1:13Yeah All right first question if you're a jeezing ping and you're scaling pilled what is it that you do? Don't answer that question John That's bad for AI safety. I would basically be contacting every foreigner. I would be contacting every Chinese national with family back home and saying, I want information. I wanna know your recipes. I wanna know, I wanna know. I wanna kind of like a AI lab foreigners or hardware foreigners. Honey potting open AI. I would basically, like, this is totally off cycle, but like, this is off the, off the reservation. But like, I was doing a video about Yugoslavian nuclear program.
1:45What? Nuclear weapons program. Start it. absolutely nothing. One guy from Paris. And then one guy in Paris, he showed up and he was like, and then he had, who knows what he did. He knows a little bit about making atomic nuclear weapons, but like he was like, okay, well, do I need help? And then the state's secret police is like, I would get to everything. And then like, I shouldn't do that. I was getting you everything. And for like a span of four years, they basically, they drew up a list. What do you need? What do you But what are you gonna do? What is it gonna be for? And they just state police just got everything.
2:21If I was running a country and I needed catch up on that, that's the sort of thing that I would be doing. So, okay, let's talk about the SP and Arch. So, what is the most valuable piece of, if you could have this blueprint, like this one megabyte of information, do you want it from TSMC, do you want it from Nvidia, do you want it from OpenAI, what is the first thing you would try to steal? I mean, I guess you have to stack every layer, right? And I think the beautiful thing about AI is because it's growing so freaking fast, every layer is being stressed to some incredible degree. Of course, China has been hacking ASML for over five years, and ASML is kind of like, oh, it's fine, the Dutch government's really pissed off, but it's fine.
3:01I think they already have those files in my view. It's just a very difficult thing to build. I think the same applies for like fab recipes, right? They can poach Taiwanese nationals very, not that difficult, right? Because TSMC employees do not make absurd amounts of money. You can just poach them and give them a much better life. And they have, right? A lot of smics employees are TSMC, you know, Taiwanese nationals, right? A lot of the really good ones, high up ones especially, right? And then you go up like the next layers of the stack and it's like, I think, I think, yeah, of course there's tons of model secrets.
3:37But then like, you know, how many of those model secrets do you not already have and you just haven't deployed or implemented organized. That's the one thing I would say is China just hasn't, they clearly are still not scale -pilled in my view. So these people are, I don't know, if you could hire them, it would probably worth a lot to you. Because you're building a fab that's worth tens of billions of dollars. And this talent is like, they know a lot of shit. How often do they get poached? Do they get poached by like foreign adversaries or do they just get poached by other companies within the same industry, but in the same country.
4:13And then, yeah, why doesn't that sort of drive up the wages? I think it's because it's very compartmentalized. And I think back in the 2000s, prior to TSB4, Smith got a big. It was actually much more kind of open, more flat. I think after that, there was like after the Among song and after all the Samsung issues and after all the, the Smiths rise, when they're literally saw. I think you should tell that story, actually, the TSMC guy that went to Samsung and SMAC and all that. I think you should tell that story out. There are two stories. There's a guy if he ran a semiconductor company in Taiwan called Worldwide Semi Conductor, and this guy, Richard Chang, was very religious.
4:50I mean, all the TSMC people are pretty religious. But like he particularly was very fervent and he wanted to bring religion to China. So after he sold his company to TSMC, huge cooper TSMC, he worked there for about eight or nine months, and he was like, all right, I'll go to China. Because back then there was the relations between China and Taiwan were much more different. And so he goes over there, at Shanghai says, we'll give you a bunch of money, and then Richard Chang basically recruits half of a whole bunch, it's like a conga line of Taiwanese with good lines, just like they get on the plane and they fly on over.
5:21And generally, that's actually a lot of acceleration points within China semiconductor industry. It's from talent flowing from Taiwan. And then the second thing was the Among Song. The Among Song was a nut. And I've met him, I've not met him. I've been people who work with him. And they say he is a nut. He is probably on the spectrum. And he does not care about people. He does not care about business. He does not care about anything. He wants to take it to the limit. The only thing, that's the only thing he cares about. He worked from TSMC, literal genius, 300 patents or whatever, 285, goes, works all the way to like the top, top tier.
5:57And then one day, he decides, he loses out on some sort of power game within TSMC and gets demoted. And he was like head of R &D, right? Or something? He was like one of the top R &Ds, he was like second or third place. And it was for the head of R &D position basically. More of the head of R &D position, he's like, I can't deal with this. And he goes to Samsung and he steals a whole bunch of talent from TSMC. Literally, again, Konga line goes to just emails, people say, we will pay. At some point, some of these people were getting paid more than the Samsung Chairman, which, and not really comparable.
6:28But like, you know what I mean. So there goes the Samsung Chairman usually, like, like part of the family that owns Samsung. Correct them, okay. So it's like kind of relevant. So it's a bet that he goes over there and he's like, well, I'm like, we will make Samsung into this monster. We forget everything, forget all of the stuff you've been trying to do it, like incremental, toss that out. We are going to the leading edge and that is it. They go to the leading edge, the guys like... They win Apple's business. They win Apple's business. They win it back from TSMC. Or did they win it back from TSMC?
6:59They win it back from TSMC. They win it back from TSMC. They had a big portion of it. And then TSMC, Morastang is like, at this time, was running the company and he's like, I'm not letting this happen, because that guy talks to work for as well, but also Goddamn brilliant. And also, very good at motivating people. He's like, we will work literally day or night, sets up what is called the Nightingale Army, where they split a bunch of people and they say, you are working R &D night shift. There is no rest at the TSMC Fab. You will go in, there is, how does you go in? There will be a day shift going out.
7:35They call it the, it's like, you're burning your liver. Cause in Taiwan, they say like, if you get old, like as you work, you're sacrificing your liver. They call it the liver buster. So they basically did this night and gale armory for like a year or two years. They finished Fin Fett. They basically just blow away Samsung. And at the same time, they sue the Almond Song directly for stealing trade secrets, Samsung basically separates from Nell Monksong and Nell Monksong goes to Smith. And so Samsung, like at one point, was better than TSMC. And then yeah, he goes to Smith and Smith is now better than, well, are not better, but they caught up rapidly as well after.
8:14Very rapid. That guy's a genius. That's the guy's a genius. I mean, I don't even know what to say about him. He's like 78 and he's like, beyond brilliant, does not care about people. Like, what is research to make the next process us know to look like, is it just a matter of like 100 researchers go in, they do like the next N plus one, then the next morning, the next 100 researchers go in. It's experiments. They have a recipe and what they do, every recipe, a TSMC recipe is the culmination of a long, long years of like research, right? It's highly secret and the idea is that what you're going to do is that you go, you look at one particular part of it and you say experiment, run experiment, is it better?
8:56is it not? Is it better or not? Kind of a thing like that. You're basically, it's multi variable problem that each every single tool, sequentially you're processing the whole thing. You turn up knobs up and down on every single tool. You can increase the pressure on this one specific deposition tool. And what are you trying to measure? Is it like does it increase the yield? Or like what is it that? It's not, it's yield, it's performance, it's power. It's not just a one, it's not just better or worse, right? It's a multi variable search space. And what are these people in those such that they can do this?
9:24Is it they'd understand the chemistry and physics? So it's a lot of intuition, but yeah, it's PhDs in chemistry, PhDs in physics, PhDs in EE, brilliant geniuses people, and they don't even know about the end chip a lot of times. It's like, oh, I am an etch engineer, and all I focus on is how hydrogen fluoride etches this. And that's all I know. And if I do it at different pressures, if I do it at different temperatures, if I do it with a slightly different recipe of chemicals, it changes everything. I remember someone told me this when I was speaking. Like, how did America lose the ability to do this sort of thing?
9:58Like, etch and hydrofluoric and acid all of that. I told them, like, he told me basically it was like, it's very apprentice, master apprentice. Like, you know, in Star Wars, SIF, there's only one, right? Master apprentice, master apprentice. It used to be that there is a master, there's apprentice, and they pass on their secret knowledge. This guy knows nothing but etch, nothing but etch. Over time, the apprentice has stopped coming. And then in the end, the apprentice has moved to Taiwan. And that's the same way it's still run. Like you have the NTU and the NTHU, Ching Hoi University, National Ching Hoi University.
10:30There's a bunch of masters, they teach apprentices, and they just pass this secret knowledge down. Who are the most AGI -filled people in the supply chain? Is there anybody like the podcast that I can have my phone call with Collette right now? Okay, go for it. Sorry, sorry. Could you mention to the podcast that Nvidia has got a guy who's calling Dylan for the to update him on the earnings call? Well, it's not that it's not exactly that. but go for it, go for it. So Dylan is back from his call with Jensen Huang. Just not with Jensen, Jesus. What did they tell you, huh? What did they tell you about next year's earnings?
11:02No, let's just color around like a hopper black well like margins. It's like quite boring stuff. I'm sure. For most people, I think it's interesting though. I guess we could start talking about it in video. You know what, before we do, we'll go back to China. There's like a lot of points there. All right, we covered the chips themselves. How do they get like the 10 gigawatts data center up? What else do they need? So I think there is a true question of how decentralized do you go versus centralized, right? And if you look in the US, right, as far as labs and such, the OpenAI, XAI, you know, Anthropic, and then Microsoft having their own effort, Anthropic having their own efforts despite having their partner and then Meta.
11:43And you know, you go down the list, it's like there is quite a decentralization. And then all the startups, like interesting startups that are out there doing stuff. There's quite a decentralization of efforts. Today in China, it is still quite decentralized, right? It's not like, Alibaba, Baidu, you are the champions, right? You have like DeepSeek, like, who the hell are you? Does government even support you? Like doing amazing stuff, right? If you are Xi Jinping and scale -pilled, Interesting. You must now centralize the compute resources, right? Because you have sanctions on how many Nvidia GPUs you can get in.
12:15Now, there's still north of a million a year, right? Even post October last year sanctions. There's still more than a million H20s and other hopper GPUs getting in through, you know, other means, but legally like the H20s. And then on top of that, you have every domestic chips, right? But that's less than a million chips. So then when you look at it, it's like, oh, well, we're still talking about a million chips. The scale of data centers people are training on today slash over the next six months is 100 ,000 GPUs, right? Open AI, XAI, right? These are like quite well documented. and others, but in China, they have no individual system of that scale yet, right?
12:54So then the question is like, how do we get there? You know, no company has had the centralization push to have it cluster that large and train on it yet at least publicly like well known, and the best models seem to be from a company that has got like 10 ,000 GPUs, right? Or 16 ,000 GPUs, right? So it's not quite as centralized as the US companies are and the US companies are quite decentralized. If you're zezing, being in your scale -pilled, do you just say XYZ company is now in charge and every GPU goes to one place. And then you don't have the same issues with the US, right? In the US, we have a big problem with like being able to build big enough data centers, being able to build substations and transformers and all this that are large enough in a dense area.
13:37China has no issue with that at all because their supply chain adds like as much power as like half of Europe every year, right? like, or some absurd statistics, right? So they're building transformers, substations, they're building new power plants constantly. So they have no problem with getting power density and you go look at Bitcoin mining, right? Around the three gorgeous dam, at one point at least, there was like 10 gigawatts of Bitcoin mining estimated, right? Which, we're talking about, you know, gigawatt data centers are coming over, you know, 2627 in the, or 26th year in the US, or 27, right?
14:13Yeah. You know, sort of, this is an absurd scale relatively, right? We don't have gigawatt data centers, you know, ready, but like China could just build it in six months, I think, around the three gorgeous dam or many other places, right? Because they have, they have the ability to do the substations, they have the power generation capabilities, everything can be like done like a flip of a switch, but they haven't done it yet. And then they can centralize the chips like crazy, right? Now, oh, all million chips that Nvidia shipping in Q3 and Q4, the H20, let's just put them all in this one data center.
14:42They just haven't had that centralization effort. Right. Well, you can argue that like the more you centralize it, the more you start building this monstrous thing within the industry, you start getting attention to it. And then suddenly, you know, low and behold, you have a little bit of a little worm in there. Suddenly, where you're doing your big training run, oh, this GPU off, oh, this GPU, oh, no, oh, no, no. I don't know if it's like that. Well, is that a Chinese accent, by the way? Okay, just to be clear, John is East Asian. East Asian descent. Half Taiwanese, half Chinese. Right, that is right.
15:17But I think, I don't know if that's as simple as that to like, cause training systems are like fire, like they're water, is it water gated, firewalled, what is it called? Not firewalled. I don't know, there's a word for that. Where they're not like, they're, airgapped, airgapped. I think they're going through like, all the like four elements. They're like, they're earth protected water. If you're using big, your scale -pilled. You kind of like, you snuck the air vendors, fuck the air vendors, you know. We got the avatar, right? Like, you have to build the avatar, okay. I think that's possible.
15:54The question is like, does that slow down your research? Do you like crush like cracked people like DeepSeek who are like clearly like not being, you know, influenced by the government and put some like idiot like, you know, idiot bureaucrat at the top. Suddenly, he's all thinking about like, you know, all these politics and he's trying to deal with all these different things. Suddenly, you have a single point of failure. And that's bad. But I mean, on the flip side, right? Like, there is like, obviously immense gains from being centralized because of the scaling loss, right? And then the flip side is, compute efficiency is obviously going to be hurt because you can't do, you can't experiment and like have different people lead and try their efforts as much if you're less centralized.
16:36A more centralized. So it's like, there is a balancing act there. The fact that they can centralize, I didn't think about this, but that is actually like, cause even if America as a whole is getting millions of GPUs a year, the fact that any one company is only getting hundreds of thousands or less means that there's no one person who can do a training run as big in America as if like China as a whole decides to do one together. The 10 gigawatts you mentioned near the three -word just down, is it like literally like, how widespread is it? Like a state? Is it like one wire? Like how I think like between not just the damn itself, but like also all of the coal.
17:14There's some nuclear reactors there, I believe as well. Between all of NN like renewables like solar and wind, between all of that in that region, there is an absurd amount of concentrated power that could be built. I don't think it's like, I'm not saying it's like one button, but it's like, hey, within X -mile radius, right? Is more of the correct way to frame it. And that's how the labs are also framing it, right? And then they started. And in the US, if they started right now, like how long does it take to build the biggest AI data center in the world? You know, actually, I think the other thing is, could we notice it?
17:51I don't think so, because the amount of factories that are being spun up, the amount of other construction, manufacturing, et cetera, that's being built A gigawatt is actually like a drop in the bucket, right? Like a gigawatt is not a lot of power. 10 gigawatts is not an absurd amount of power, right? It's okay, yes, it's like hundreds of thousands of homes, right? What, yeah, millions of people, but it's like, you got 1 .4 billion people, you got like most of the world's like extremely energy intensive, like refining and like, you know, rare earth refining and all these manufacturing industries are here.
18:23It would be very easy to hide it. It would be very easy to just like shut down like, I think the largest aluminum mill in the world is there, and it's like, it's like North to five gigawatts alone. It's like, oh, what could we tell if they stopped making aluminum there, and instead started like making, you know, AIs there, or is making AIs there? Like, I don't know if we could tell, right? Because they could also just easily spawn like 10 other aluminum mills, make up for the production and be fine, right? So like, there's many ways for them to hide compute as well. To the extent that you could just take out a five gigawatt aluminum refining center and like build a giant data center there, then I guess the way to control Chinese AI has to be the chips because like everything else, so like, how do you, like, just like walk me through how many chips would they have now, how many will they have in the future?
19:07Well, well, like, how many is that in comparison to US and the rest of the world? Yeah, so in the world, I mean, the world we live in is they are not restricted at all in like the physical infrastructure side of things in terms of power, data centers, et cetera, because their supply chain is built for that, right? And it's pretty easy to pivot that. whereas the US adds so little power each year, and Europe loses power every year, the Western sort of industry for power is non -existent in comparison, right? But on the flip side is, quote unquote, Western including Taiwan, manufactured chip manufacturing is way, way, way, way larger than China's, especially on leading edge where China theoretically has, depending on the way you look at it, either zero or a very small percentage share, right?
19:49And so there you have, you have, You have equipment, way for manufacturing, and then you have advanced packaging capacity. And where the US can control China? So advanced packaging capacity is kind of a shot because the largest advanced packaging company in the world was Hong Kong headquarters. They just moved to Singapore, but that's effectively in a realm where the US can't sanction it. A majority of these other companies are in similar places. So advanced packaging capacity is very hard. If it had packaging as useful for stacking memory, stacking chips on co -os, things like that. Then the step down is wafer fabrication.
20:28There is immense capability to restrict China there, and despite the US making some sanctions, China, in the most recent quarters, was like 48 % of ASMR's revenue. And like 45 % of applied materials, and you just go down the list. So it's like, obviously it's not being controlled that effectively, but it could be on the equipment side of things. the chip side of things is actually being controlled quite effectively, I think, right? Like yes, there is shipping GPUs through Singapore and Malaysia and other countries in Asia to China. But the amount you can smuggle is quite small. And then the sanctions have limited the chip performance to a point where it's like, this is actually kind of fair.
21:09But there is a problem with how everything is restricted, right? Because you want to be able to restrict China from building their own domestic chip manufacturing industry that is better than what we ship them. You want to prevent them from having chips that are better than what we have. And then you want to prevent them from having AI's better, the ultimate goal being, and if you read the restrictions, very clear, it's about AI. Even in 2022, which is amazing, at least the Commerce Department was kind of AI -pilled, was like, is you want to restrict them from having AI's worse than us, right?
21:40So starting on the right end, it's like, okay, well, if you want to restrict them from having better than us, you have to restrict chips. Okay, if you want to restrict them from having chips, you have to let them have at least some level of chip that the West also, that is good better than what they can build internally. But currently, the restrictions are flipped the other way, they can build better chips in China than we restrict them in terms of chips that Nvidia or AMD or Intel can sell to China. And so there's sort of a problem there in terms of the equipment that a chip can be used to build chips that are better than and what the Western companies can actually ship them.
22:15John, don't seem to just think the expert controls are kind of a failure. Do you agree with them or? That is a very interesting question because I think it's like, well, thank you. What do you... Sharkfish, you're so good. Sharkfish, you're the best. I think failure is a tough word to say because I think it's like, what are we trying to achieve, right? Like, instead, they're talking about AI, right? When you do sanctions like that, you need such a deep knowledge of the technologies. Just taking lithography, right? If your goal is to restrict China from building chips and you just like boil it down to like, hey, lithography is 30 % of making a chip, so are 25%.
22:55Cool, let's sanction lithography. Okay, where do we draw the line? Okay, let me ask. Let me ask. Let me figure out where the line is. And if I'm a bureaucrat from the lawyer at the Commerce Department or what have you, Well, obviously I'm going to go talk to ASMR. And ASMR is going to tell me, this is the line, because they know, hey, well, this, this, this is, there's like some blending over. There's like, they're like looking at like, what's going to cost us the most money, right? And then they constantly say, like, if you restrict us, then China will have their own industry, right? And the way I like to look at it is, like, chip manufacturing is like, like 3D, Chast, or like, you know, a massive Jigsaw puzzle.
23:31And that if you take away one piece, China can be like, oh, yeah, that's the piece. Let's put it in, right? And currently this export restrictions, year by year by year, they keep updating them. Ever since like 2018 or so, 19, right? When Trump started, and now Biden's accelerated them, they've been like, they haven't just like, take a bat to the table and like, break it, right? Like it's like, let's take one jigsaw puzzle out, walk away, oh shit, let's take two more out. Oh shit, right? Like, you know, it's like, instead if they like, you either have to go kind of like, like full back to the frickin' like table, Sash wall, or chill out, right?
Read the full transcript
24:08Like, and like, you know, let them do whatever they want. Cause the alternative is everything is focused on this thing and they make that. And then now when you take out another two pieces, it's like, well, I have my domestic industry for this. I can also now make a domestic industry for these. Like, you go deeper into the tech tree or what have you? It's a very, it's art, right? In a sense that there are technologies out there that can compensate. Like if you believe the belief that lithography is a linchpin within the system is it's not exactly true, right? At some point if you keep pulling, keep pulling a thread, other things will start developing to kind of close that loop.
24:44And like I think it's a, it is, that's why I say it's an art, right? I don't think you can stop Chinese semiconductor industry, for the semiconductor industry from progressing. I think that's basically impossible. So the question is the Chinese government believes in the primacy of semiconductor manufacturing They used they've believed it for a long time, but now they really believe it, right? To some extent the sanctions have made China believe in the importance of the semiconductor industry more than anything else So from an AI perspective What's the point of export control then because even if like if they're gonna be able to get these?
25:18Like if you were like concerned about AI and they're gonna be able to build centralized though, right? So that's the big question is are they centralized? And then also, there's the belief, I'm not sure if I really believe it, but like prior podcasts, there have been people who talked about nationalization, right? In which case, okay, now you're talking about your favorite exam videoously. Well, I think there's a couple. And I love a little poll. I'm like, you know, no, but I think there have been a couple where people have talked about nationalization, right? But like if you have, you know, nationalization, then all of a sudden you aggregate all the flops and it's like, no, there's no fucking way, right?
25:51China can be centralized enough to compete with each individual US lab. They could have just as many flops in 25 and 26 if they decided they were scale -pilled. Just from foreign ships, for individual model. In 2026, they can train a 2027, they can release a 2027 model by 2026. Yeah, and then a 2028 model, 2028 model in the works. They totally could just with foreign ships supply. Just a question of centralization. Then the question is, do you have as much innovation and compute efficiency wins or what have you get developed when you centralize, or does like, anthropic and open AI and XAI and Google like all develop things and then like secrets kind of shift a little bit in between each other and all that, like, you know, you end up with that being a better outcome in the long term versus like the nationalization of the US, right?
26:39If that's possible and like, or, you know, and what happens there. But China could absolutely have it in 26, 27 if they just have the desire to. And that's just from foreign chips, right? And then domestic chips are the other question, right? 600 ,000 of the, a send 910b, which is roughly like 400 terra flaps or so. So if they put them all in one cluster, they could have a bigger model than any of the labs next year. Right? I have no clue where all the send 910b's are going, right? But I mean, well, there's rumors about like some, they are being divvied up between the like major Alibaba, bite dance, buy -do, et cetera.
27:18And next year more than a million. And it's possible that they actually do have, you know, 1 E30 before the US because data center is not as big of an issue. 10 gigawatt data center is gonna be I don't Think anyone is even trying to build that today in the US like even out to 2728. Well, they're focusing on like linking many data centers together So there's a possibility that like hey come 2028 2029 Chinat can have more flops delivered to a single model even ignoring sort of even once the centralization question is solved, right? Because that's clearly not happening today for either party. And I would bet if AI is like as important as, you know, you and I believe that they will centralize sooner than the West does.
28:01So there is a possibility, right? Yeah. It seems like a big question then is how much could to smick either increase the product, like increasing amount of waifers, like how many more waifers could they make and how many of those waifers could be dedicated to the night. Because I assume there's other things they wanna do with these seven kind of things. Yeah, so there's like two points parts there too, right? So the way the US's sanctioned smick is really like stupid kind of is that in that they've like sanctioned a specific spot rather than the entire company. And so therefore, right, smick is still buying a ton of tools that can be used for their seven nanometer and they're, you know, call it 5 .5 nanometer process or 6 nanometer process for the 910C, which releases later this year, right?
28:41They can build as much of that as long as it's not in Shanghai, right? And Shanghai has anywhere from 45 to 50 high -end immersion lithography tools is what's like believed by intelligence as well as like many other folks. That roughly gives them as much as 60 ,000 wafers a month of 7 nanometer, but they also make their 14 nanometer in that fab, right? And so the belief is that they actually only have about like 25 to 35 ,000 of 7 nanometer capacity. Wafers a month, right? Doing the math, right? Are the chip die size and all these things? Because probably also use the chiplets and stuff so they can get away with using less leading edge wafers, but then their yields are bad.
29:25You can roughly say, you know, something like 50 to 80 good chips per wafer. with their bad yield, right? With their bad yield. Right, with their bad yield. Because it's hard, right? You know, even if it was like, you know, everyone's knows the number, right? Like a thousand steps, even if you're 99 % for each. Like 98 or 98 % like in the end, you'll still get a 40 % yield. You know, overall interesting. I think it's like, even if it's like 99, if I think it's like, I think it's, if it's six sigma of like, or of like perfection and you have your 10 ,000 plus steps, you end up with like yield is still dockshit by the end, right?
30:00like yeah, that is a scientific measure, dog shit percent. Yeah, yeah, as a multiplicative effect, right? So yields are bad because they have hands tied behind their back, right? Like, they are not getting to use UV, whereas on seven animators, I'll never use UV, but TSMC eventually started using UV, initially they used UV, right? It doesn't mean they actually were control succeeded because they have bad yields because they have have to use like success. Again, they still are determined. Successes mean they stop. They're not stomping. Going back to the yield question, right? Like, oh, theoretically, 60 ,000 wafers a month, times 50 ,200 dies per wafer with yielded dies.
30:45Holy shit, that's millions of GPUs. Right? Now, what are they doing with most of their wafers? They still have not become skill -pilled, so they're still throwing them at, like, let's make 200 million Huawei phones. Right? Like, oh, okay, cool, I don't care. Like, as the West, you don't care as much, even though like Western companies will get screwed like Qualcomm and like, you know, or an immediate tech Taiwanese companies. So obviously there's that. And the same applies to the US, but when you flip to like, sorry, I don't know what I was gonna say. I was gonna say. I was gonna say. I nailed it.
31:18We're keeping this in. That's fine, that's fine. Hey everybody, I am super excited to introduce our new sponsors, Jane Street. They're one of the world's most successful trading firms. I have a bunch of friends who either work there now or have worked there in the past. And I have very good things to say about those friends, and those friends have very good things to say about Jane Street. Jane Street is currently looking to hire its next generation of leaders. As I'm sure you've noticed, recent developments in AI have totally changed what's possible in trading. They've noticed this too. And they've stacked a scrappy chaotic new team with tens of millions of dollars of GPUs to discover signal that nobody else in the world can find.
32:00Most new hires have no background in trading or finance. Instead they come from math, CS, physics, and other technical fields. A particular relevance to this episode, their deep learning team is hiring CUDA programmers, FPGA programmers, and ML researchers. Go to jainstreet .com slash doarkesh to learn more. And now back to Dylan and John. 2026, if they're centralized, they can have as big training runs as anyone, a US company. Oh, the reason why I was bringing up Shanghai, they're building seven nanometer capacity and Beijing. They're building five nanometer capacity and Beijing, but the US government doesn't care.
32:39And they're importing dozens of tools into Beijing. And they're saying to the US government in a smell, this is for 28 nanometer obviously. This is not bad. And then obviously you know, like in the background, we're making five nanometer here. Are they doing it because they believe in AI or because they want to make Huawei phones? You know, Huawei was the largest TSMC customer for like a few quarters actually, before they got sanctioned. Huawei makes most of the telecom equipment in the world, right? You know, phones of course, modems, but of course, accelerators, networking equipment. You know, you go down the whole like video surveillance chips, right?
33:11Like you kind of like go through the whole gambit. A lot of that could use seven and five nanometer. Do you think the dominance of Huawei is actually a bad thing for the rest of the Chinese tech industry? I think Huawei is so fucking cracked that it's hard to say that, right? Like, Huawei out competes Western firms regularly with two hands tied behind their back. Like, what the hell is Nokia and Sony Ericsson, like trash, right? Like compared to Huawei, and Huawei is not allowed to ship sell to European companies or American companies, and they don't have TSMC and yet they still destroy them. And same applies to the new phone, right?
33:51It's like, oh, it's as good as a year old Qualcomm phone on a process node that's equivalent to four years old, right, or three years old. So it's like, wait, so they actually engineered us with a worst process node. So it's like, oh wow, okay. Like, you know, Hawley is like crazy cracked. Why do you think that culture comes from? The military because it's the PLA. It is that we it is generally seen as an arm of the PLA, but like How do you square that with the fact that sometimes a PLA seems to mess stuff up? Oh, like filling water and rockets. I don't know if that was true There is there is like that like like crazy conspiracy Not care conspiracies like you you don't know what the hell to believe in China Especially as a not Chinese person, but like nobody knows even Chinese people don't know what's going on in China There's like, you know, like all sorts of stuff, like, oh, they're filling water in their rockets.
34:40Clearly, they're like in constant. It's like, look, if I'm the Chinese military, I want the Western world to like believe I'm completely incompetent, because one day, I can just like destroy the fuck out of everything, right? With all these hypersonic missiles and all this shit, right? Like drones and like, no, no, no, no, no, no, we're filling water in our missiles. These are all fake. We don't actually have 100 ,000 missiles that we manufacture in a facility that's like superhyper advanced and Rathie on his stupidest shit because they can't make, you know, missiles nearly as fast. right? Like I think like that's also like a flip side is like how much false propaganda is there, right?
35:11Because there's a lot of like no, smick could never, smick could never. They have, they don't have the best tools. But but but then it's like, mother fucker, they just shipped 60 million phones last year with this chip that performs only one year worse than like what Qualcomm has. It's like proof is in the pudding, right? Like, you know, there's a lot of like, cope if you will. I just wonder where it comes from. I do really do just wonder where that culture comes from. Like there's something crazy about them where they're kind of like everything they touch they seem to succeed in. And like I kind of wonder why.
35:40They're making cars. I wonder if it's going on there. I think I like if like supposedly like if we kind of imagine like historically like do you think they're getting something from somewhere? What do you mean? Espionage? Yeah. Like obviously. Like East Germany in the Soviet industry was basically was just a it was like a conveyor belt of like secrets coming in and they're just use that to run everything. But the Soviets were never good at it. They could never mass How would Espionage explain how they can make things with different processes? I don't think it's just Espionage. I think there's literally...
36:10It has to be something else. They have the Espionage without a doubt, right? Like, ASML has been known to be hacked a dozen times, right? Or at least a few times, right? And they've been known to have people sued who made it to China with a bunch of documents, right? Not just ASML, but every fucking company in supply chain. Cisco code was literally in like early Huawei, like routers and stuff, right? Like you go down the list, it's like everything is... but then it's like, no, architecturally, the Ascend 910B looks nothing like a GPU, it looks nothing like a TPU, it is like its own independent thing, sure they probably learned some things from some places, but like, it is just like their good at engineering.
36:43It's 996, like wherever that culture comes from, they do good, they do very good. I know, another thing I'm curious about is like, yeah, we're their culture confirms, but like, how does it stay there? Because with American firms or any other firm, you can have a company that's very good, but over time it gets worse, right? like Intel or many others. I guess Huawei just isn't that old with company, but it's hard to be a big company and stay good. That is true. I think it's like, but I think a lot, a word that I hear a lot with regards to Huawei's struggle, and China has a culture of the communist parties that's really big on struggle.
37:19I think Huawei in the sense they sort of brought that culture into the way they do it, like you said before, right? They go crazy because they think that in five years they're gonna fight the United States. And so literally everything they do, every second is like, their country depends on it. It's like, it's the Andy Grovey and mindset, right? Shout out to the based intel, but only the paranoid survive, right? Paranoid Western companies do well. Why did Google really screw the pooch on a lot of stuff and then why are they like researching kind of now? It's because they got paranoid as hell, right?
37:53But they weren't paranoid for a while. If Huawei is just constantly paranoid about like the external world and like, oh fuck, we're gonna die, oh fuck, like, you know, they're gonna beat us. Our country depends on it. We're gonna get the best people from the entire country that are like, you know, the best at whatever they do. And tell them, you will, if you do not succeed, you will die. Or like, my people die. Your family will die. Your family will be enslaved and everything will be terrible. By the evil western fags, right? Even Western, like, like, like, like, like, like, not capitalists. They don't believe in cut.
38:22They don't say that anymore. But some know like, like, you know, everyone is against China. China is being, it's been defiled, right? And like, they're saying, like, if you, that is all on you, bro, like you can't do that. And then like you, if you can't get that fucking radio to be slightly less noisy and like transmit, like 5 % more data, like the great postfire all over again, the British are coming and they will steal all the, So all the trinkets and everything, like that's on you. Uh -huh. Why isn't there more vertical integration in this interconnected industry? Well, like, why are there like this subcomponent requires a Satherson component from this other company, which we're guys have come on in from then the company like, why is more of it not done in house?
39:01The way to look at it today is it's super, super stratified in every industry has anywhere from one to three competitors. And pretty much the most competitive it gets is like 70 % share, 25 % share, 5 % share, in in any layer of manufacturing chips, anything, anything chemicals, different types of chips, but it used to be vertically integrated. Or the very beginning it was integrated, right? And what happened was, it was the funniest thing it said, you know, yet companies that used to do it all in the one, and then suddenly, sometimes a guy would be like, I hate this. I think I know how to do better.
39:36Spins off, does his own thing, starts his company, goes back to his old company, says, I can sell you a product spinner, right? And that's the beginning of what we called the semiconductor manufacturing and equipment industry. Basically, it was seven years, right? Like everyone made their own equipment. 60s. 60s, like they spin off all these people. And then what happened was that the companies that accepted, you know, these outside products and equipment got better stuff. They did better. Like you can talk about a whole bunch, like there are companies that were totally vertically integrated in semiconductor manufacturing for decades.
40:04And they are still good, but they're nowhere near competitive. One thing I'm confused about is like, the actual foundries themselves, there's like fewer and fewer of them every year, right? So there's maybe more companies overall, but the final people make the way for users less and less. And then it's interesting in a way similar to the AI foundation models where you need to use the revenues from a previous model or your market share to fund the next round of ever more expensive development. When TSMC launched the Foundry industry, right? And when they started, there was a whole wave of Asian companies that funded semiconductor foundries of their own.
40:46You had Malaysia, with Siltera, Singapore, with Charter. You had, there was a white semiconductor where I talked about earlier. There's loans from Hong Kong, bunch in Japan. Bunch in Japan, they all sort of did this thing. And I think the thing was that when you go into leading edge, when the thing is that it got harder and harder, which means that you had to aggregate more demand from all the customers to fund the next node, right? So technically in the sense that what it's kind of is aggregating all this money, all this profit to kind of fund this next node to the point where now, like there's no room in the market for an N3.
41:21Like technically you could argue that economically, you can make an argument that like N2 is a monstrosity that doesn't make sense economically and what should not exist in some ways. without the immense single concentrated spend of like five players in the market. I'm sorry to like completely derail you, but like there's this video where it's like this is an holy concoction of meat slurry. Yes. What? Sorry, there's like a video that's like, ham is disgusting. It's an unholy concoction of like meat with no bones or collagen. And like, I don't know, like to use like, the way he was describing two dead -eaters, kind of like that, right?
42:00It's like the guy who pumps his right arm from so much and he's like super muscular. The human body was not meant to be so muscular. Like, what's the point? Like why is through an intermediate or not justify it? I'm not saying N2 is like, N2 specifically, but say N2 is a concept. The next node should technically, like right now, there will come a point where economically, the next node will not be possible. Like at all, right? Unless more technology spawn, like AI now makes, yeah, yeah, yeah, yeah. One nanometer or whatever, there was a long period of 16A viable. Right? So, like, four days of viable and what's as in like, money worth it.
42:37So every two years, you get a shrink, right? Yeah. Like clockwork, Moore's Law. And then five nanometer happened. It took three years, holy shit. And then three nanometer happened. It took three years, or no, sorry, it's a three nanometer. Five, it took three years. Holy shit. Like, is Moore's Law dead? Right? Like, because TSMC didn't, and then what did Apple do? Even on the third year of three of, of, or sorry, when's three nanometer finally launched, they still only, Apple only move half of the iPhone volume to three nanometer. So this is like, now they did a fourth year of five nanometer for a big chunk of iPhones, right?
43:11And it's like, oh, is the mobile industry petering out? Then you look at two nanometer and it's like gonna be a similar, like, very difficult thing for the industry to pay for this, right? Apple, of course, they have, you know, because they get to make the phone, they have so much profit, they can funnel into like more and more expensive chips, but finally, that was running out, right? It was how economically viable it was to Nanometer, just for one player, TSMC, ignore Intel, ignore Samsung. Just in, because Samsung is paying for it with memory, not with their actual profit, and then Intel is paying it from their former CPU monopoly.
43:45Probably equity money. And now, if I've got equity money in, debt and subsidies. And they need people salaries. But anyways, there's a strong argument that like funding the next node would not be economically viable anymore if it weren't for AI taking off. And then generating all this humongous demand for the most leading edge chip. So how big is the difference between 7 to 5 to 3 nanomere? Like is it a huge deal in terms of who can build the biggest cluster? So there's this simplistic argument that like, oh, moving a process node only saves me x percent and power, right? And that has been pedering out, right?
44:22You know, when you move from like 90 nanometer to 80 something, right? Or 70 something, right? It was like, it was like, it was too extra, right? Danard scaling was still intact, right? But now when you move from 5 nanometer to 3 nanometer, first of all, you don't double density. SVM doesn't scale at all. Logic does scale, but it's like 30%. So all in all, you only saved like 20 % in power per transistor. But because of like data locality and movement of data, you actually get a much larger improvement in power efficiency by moving to the next node, then just the individual transistors power efficiency benefit.
44:54Because, for example, you're multiplying a matrix that's like 8 ,000 by 8 ,000 by 8 ,000, and then you can't fit that all on one chip. But if you could fit more and more, you have to move off chip less, you have to go to memory less, etc. So the data locality helps a lot too. But the AI really, really, really wants new processed nodes because of A, a power used is a lot less now, higher density, higher performance of course, but the big deal is like, well, if I have a gigawatt data center, I can now how much more flops can I get? If I have two gigawatt data center, how much more flops can I get?
45:26If I have a 10 gigawatt data center, how much more flops can I get, right? And like you, you look at the scaling, and it's like, well, no, everyone needs to go to the most recent process node as soon as possible. I want to ask the normy question, for like everybody's, well, I want to phrase it that way. Okay, I want to ask a question that's like, I'm not really a nori. Not for your nerves. I think John and I could communicate to the point where you wouldn't even know the fuck it's on. Okay, suppose Taiwan is invaded or Taiwan has an earthquake, nothing is shipped out of Taiwan from now on. What happens next?
46:02The rest of the world, how would it feel to impact a day in a weekend, a month in a year in? I mean, it's a terrible thing. It's a terrible thing to talk about. I think it's like, can you just say it's all terrible? Everything's terrible. Because it's not just like leading edge, people are worth focused on leading edge. But there's a lot of trailing edge stuff that people depend on every day. I mean, we all worry about AI. The reality is you're not gonna catch your fridge, you're not gonna catch your cars, you're not gonna get everything. It's terrible. And then there's the human part of it. It's all terrible.
46:31It's depressing. And I live there. I think day one market crashes a lot. You're gonna think about, I think the big biggest companies, Magnificent 7, and whatever that gets called, are like 60, 75 % of the S &P 500, and their entire business relies on chips, right? Google, Microsoft, Apple, Nvidia, you know, you go down the list right there, they all met up, right? They all entirely rely on AI. And you would have a tech reset, like extremely insane tech reset, by the way, right? So market would crash, week a day in, a couple of weeks in, right? People are preparing now, people are like, oh shit, let's start building fabs with fuck all the environmental stuff.
47:12Wars probably happening. But the supply chain is trying to figure out what the hell to do to refix it. But six months in, the supply of chips for making new cars, gone or sequestered to make military shit, right? You can no longer make cars. And we don't even know how to make non -Semiconductor -induced cars. This unholy concoction with all these chips. right? You are like 40 % chips now. Like it's just chips on in the tire. There's like there's like 2000 plus chips. Every Tesla door handle has like four chips at it. It's like what the fuck? Like why? Like like but like it's like it's like shitty like microcontrollers and stuff but like there's like 2000 plus chips even in an in an ice vehicle like internal combustion engine vehicle right and every engine has dozens of dozens of chips right.
47:57Anyways just all shuts down because not all of there's some in Europe, there's some in the US, there's some in Japan, there's some in Japan. They're gonna bring in a guy to work on Saturday until four. Yeah, yeah, I mean, yeah. So you have like TSMC always builds new fabs. That old fab, tweak production up a little bit more and more, and new designs move to the next, next, next node, and old stuff fills in the old notes, right? So ever since TSMC's been the most important player, and not just TSMC, there's UMC, there's PSMC, there's a number of other companies there, Taiwan's share of total manufacturing has grown every single process node.
48:33So in like 130 nanometer, there's a lot and including many chips from Texas Instruments or analog devices or NXP, like all these companies, 100 % of his manufactured in Taiwan by either PSMC or UMC or whatever. But then you step forward and forward and forward like 28 nanometer, 80 % of the world's production of 28 nanometers in Taiwan. Oh fuck, right? And everything in 28 nanometers, is like, what's made on 28 nanometer today? Tons of microcontrollers and stuff, but also like every display driver I see, like cool, like even if I can make my Mac chip, I can't make the chip that drives the display.
49:07Like, you know, you just go down the list, like everything, no fridges, no automobiles, no weed whackers because that shit has, my toothbrush has fucking Bluetooth in it, right? Like why? I don't know, but like, you know, there's like so many things that like just like poof, where tech reset. We were supposed to do this interview like many months ago, and then I like have like delaying, because I'm like, ah, I don't understand any of the shit. But it is a very difficult thing to understand. But I feel like with AI, it's like, it's not that, like, you've just spent time. You've spent the time. But I also feel like it's less complicated.
49:37It feels like it's a kind of thing where in an amateur way, you can pick up what's going on in the field. In this field, the thing I'm curious about is how does one learn the layers of the stack? Because the layers of the stack are like, there's not just the papers online. You can't just look up the tutorial on how the transformer works or whatever. It's like, it's like, it's like, I mean like, many layers of really different things. They're like 18 year olds who were just cracked at AI. Right. Right. And like, there's high school dropouts that get like jobs at open AI. This existed in the past, right?
50:07Pat Gelsinger, current CEO of Intel, went straight to work. He was, he like grew up in the Amish area of Pennsylvania and he went straight to work at Intel, right? Because he's just cracked, right? That is not possible in semiconductors today. You can't even get like a job at like a tool company without like a, at least like a freaking master's in chemistry, right? and probably a PhD, right? Like, like, of the, like, 75 ,000 TSMC workers, it's like 50 ,000 have a PhD or something insane, right? It's like, okay, this is like, there's like some, there's like a next level amount of like, how specialized everything's gotten.
50:39Whereas today, like, you can take like, you know, Shulto, you know, he, when did he start working on AI? Not that long ago. Not to say anything about Shulto, he's like, no, he's, but he's cracked. He's like, oh, mega cracked it like what he does. What he does, you could pick him up and drop him into another part of the AI stack. First of all, he understands it already. And then second of all, he could probably become cracked at that too, right? Whereas that is not the case in semiconductors, right? You won, you like specialize like crazy. Two, you can't just pick it up. You know, like, Shultoy, I think what did he say?
51:11He like just started like, He was a consultant in McKinsey, and at like night, he would like greed papers about robotics. And like run experiments and whatever. Yeah, and then like he was like, like people noticed who the hell is this guy and why is he posting this? I thought everyone who knew about this was at Google already, right? It's like come to Google. That can't happen in semiconductors, right? Like it's just not conducively, it's not possible, right? One archive is a free thing. The paper publishing industry is like out of warrant everywhere else and you just cannot download IEEE papers or SPI papers or other organizations.
51:46And then two, at least up until like late 2022, a really early 2023 in the case of Google, right? I think what the Palm inference paper, up until the Palm inference paper, before that, all the good best stuff was just posted on the internet. After that, you know, it's kind of a little bit clamping down by the labs, but there's also still all these other companies making innovations in the public. That, and like, what is state of the art is public? That is not the case in semi -conductive. Semi -conductive has been shut down since the 1960s, 1970s, basically. I mean, like, it's kind of crazy how little information has been formally transmitted from one country to another.
52:21Like, the last time you could really think of this was like 19, maybe the Samsung era, right? So then how do you guys keep up with it? Well, we don't know it. I don't personally. I don't think I know it. I mean, I... You don't know it. It's crazy because like, there was a guy. There's like, I spoke to one guy, he's like a PhD in etch or something. The world, one of the top people in etch and he's like, man, you really know like lithography, right? I'm just like, I don't feel like I know lithography, but then you've talked to people who know lithography. You've done pretty good work in packaging, right?
52:50Nobody knows anything. They all have shelman imnesia. They're all in this single well, right? They're digging deep. They're digging deep for what they're getting at. But they don't know the other stuff well enough. And in some ways, nobody knows the whole stack. Nobody knows the whole stack. The stratification of just manufacturing is absurd. The tool people don't even know exactly what Intel and TSMC do in production and vice versa. They don't know exactly how the tools optimize like this. And it's like how many different types of tools there are? Dozens. And each of those has like an entire tree of like all the things that we've built, all the things we've invented, all the things that we continue to iterate upon.
53:28And then like here's the breakthrough innovation that happens every few years in it too. So if that's the case, if like nobody knows a whole stack, then how does the industry coordinate to be like, you know, in five into years we want them to go to the next process, which has gait all around. And for that, we need X tools and next technology's developed by whatever. That's really fascinating. It's a fascinating social kind of phenomenal, right? You can feel it. I went to Europe earlier this year. Dylan was like, had allergies. But like, I was like, talking to those other people. And you can just, it's like gossip.
54:01It's gossip. You start feeling the, you start feeling people could coalescing around like something, right? Early on, we used to have like, like Semitech, where people, all these American companies came together and talked and they came and they hammered out, right? But Semitech, in reality, was dominated by a single company, right? And then, you know, nowadays is a little more dispersed, right? You feel like it's like, it's like a, it's a blue moon arising kind of thing. Like, they are going towards something, they know it, and then suddenly the, the whole industry is like, this is it. Let's do it.
54:33But I think it's like, God came and proclaimed it. We will shrink density to X every two years. So Gordon Morse, he made an observation, and then it didn't go nowhere as it went way further than he ever expected because it was like, oh, there's a line of sight to get to here and here. And he predicted like seven, eight years out, like multiple orders of magnitude of increases in transistors, and it came true. But then by then the entire industry was like, this is obviously true, this is the word of God. And every engineer in the entire industry, tens of millions of people, like literally, this is what they were driven into.
55:05No, not every single engineer didn't believe it. But like people were like, yes, to hit the next shrink, we must do this, this, this, right? And this is the optimizations we make. And then you have this obstratification, every single layer, and abstraction layers, every single layer through the entire stack to where people, it's an unholy concoction, I mean, you'd say in this word, but like, no one knows what's going on because there's an abstraction layer between every single layer. And on this layer, the people below you and the people above you know what's going on. And then like beyond that, it's like, okay, I can try to understand, but not really.
55:38But I guess I didn't answer the question of, when I already asked her whatever, I don't know, was it 10, 20 years ago, I watched your video about it where they're like, we are UV, it's like, this is, we're gonna do UV instead of the other thing, and this is the path forward. How do they do that if they don't have the whole sort of picture of like different constraints, different trade -offs, different blah, blah, blah. They kind of argue it out. They get together and they talk and they argue. And basically at some point, a guy somewhere says, I think we can move forward with this. Semicaductors are so siloed, and the data and knowledge within each layer is a, not documented online, at all.
56:16Right. Documentation. Because it's all siloed within companies. B, it is, there's a lot of human element to it, because a lot of the knowledge, like as John was saying, is like apprentice master, apprentice master, type of knowledge, or I've been doing this for 30 years, and there's an amazing amount of intuition on what to do, just when you see something, to where AI can't just learn semiconductors like that. But at the same time, there's a massive amount of talent shortage and ability to move forward on things. So the technology used on most of the equipment in semiconductor tools, fabs, runs on Windows XP.
56:56Right? Like that each tool has like a Windows XP server on it or like, you know, like all the chip design tools Like have like sent host sent host like version six, right? And like that's old as hell, right? So like there's like so many like areas where like why is this so far behind at the same time? It's like so like hyper optimized. That's like the the tech stack is so broken in that sense. They're afraid to touch it They're afraid to touch it. Yeah, because it's an unholy amalgamation. It's unholy. It should not be work It should not work. This thing should not work. It's literally a miracle So you have all the abstraction layers, but then it's like, one is there's a lot of breakthrough innovation that can happen now stretching across abstraction layers.
57:33But two is because there's so much inherent knowledge in each individual one, what if I can just experiment and test at a thousand x velocity or a hundred thousand x velocity. And so some examples of where this is already like shown true is some of Nvidia's AI layout tools, right? And Google as well, like laying out the circuits within a small blob of the chip with AI. Some of these like RL design things, some of the, there's a lot of like various like simulation things. What is that design or is that manufacturing? It's all design, right? Most of it's design. Manufacturing has not really seen much of this yet.
58:07Although there is starting to come in. Inverse lithography maybe. Yeah, ILT and say, yeah, maybe I don't know if that's AI, that's not AI. Anyways, there's like tremendous opportunity to bring breakthrough innovation simply because there is so many layers where things are unoptimized, right? So you see all these single digit, low double digit advantages just from RL techniques from Alphago's type stuff, or not RL from Alphago, but like 5, 6, 7, 8 -year -old RL techniques being brought in. But generally, AI being brought in could really revolutionize the industry, although there's a massive data problem.
58:47So can you give the possibilities here in numbers in terms of maybe like a flop per dollar or whatever the relevant thing here is? How much do you expect in the future to come from process and order improvements? How much from just like how the hardware is designed because of AI? If you like how to decide, what time specifically for like GPUs? Yeah, like we have to disaggregate future improvements. I think I think you know it's first it's important to state that semiconductor manufacturing and design is the largest search space of any problem that humans do because it is the most complicated industry that anything that humans do and so you know when you think about it, right there's there's one e 10 one e 11 right 100 billion transistors yeah on on leading edge chips right blackwell has 220 billion transistors or something like that So what is and those are just on off switches and then think about every permutation of putting those together contact ground, et cetera, drain source, blah, blah, blah, with wires, right?
59:49There's 15 metal layers, right? Connecting every single transistor in every possible arrangement. This is a search space that is literally almost infinite, right? You could like, the search space is much larger than any other search space that humans know. And it's like, I'm not sure if the search, like, what are you trying to optimize over? Well, useful compute, right? What is, you know, if the goal is optimize intelligence per pica jule, right? And intelligence is some nebulous nature of like the what the model architecture is. And then Pikachu is like a unit of energy. How do you optimize that?
1:00:21So there's humongous innovations possible in architecture, because vast majority of the power on a H100 does not go to compute. And there are more efficient like compute, you know, ALU's or Ethmic Logic Unit like designs, right? But even then, the vast majority of the power doesn't go there, right? The vast majority of the power goes to moving data around. Right? And then when you look at what is the movement of data? It's either networking or memory. You know, you have a humongous amount of movement relative to compute and a humongous amount of power consumption relative to compute. And so how can you minimize that data movement and then maximize the compute?
1:01:05There are 100x gains from architecture. Even if we like literally stop shrinking, I think we could have 100x gains from architectural advancements. Over what time period? That the question is how much can we advance the architecture? Right, the challenge, the other challenge is like, the number of people designing chips has not necessarily grown in a long time, right? Yeah, like company to company shifts, but within the semiconductor industry in the US, and the US designs the vast majority of leading edge chips, the number of people designing chips has not grown much. What has happened is the output per individual has sort because of EDA, electronic design assistance tooling.
1:01:44Now, this is all still like classical tooling. There's just a little bit of inkling of AI in there yet. Right? What happens when we bring this in is the question and how you can solve this search space somehow with humans and AI working together to optimize this. So it's not most of the power is data movement. And then the compute is actually very small. To flip side, the compute is, first of all, a compute can get like 100x more efficient just with like design changes, and then you can minimize that data movement massively. So you can get a humongous gain in efficiency just from architecture itself.
1:02:18And then process node helps you innovate that there, right? And power delivery helps you innovate that. System design, chip to chip networking helps you innovate that, right? Like memory technologies, there's so much innovation there. And there's so many different vectors of innovation that people are pursuing simultaneously to where like Nvidia, gender, gender, Jen will do more than two X performance per dollar. I think that's very clear. And then like hyper scalers are probably gonna try and shoot above that, but we'll see if they can execute. There's like two narratives you can tell here of how this happens.
1:02:50One is that these AI companies were training the foundation models who understand the trade -offs of like how much is the marginal increase in compute versus memory worth to them and what trade -offs do they want between different kinds of memory. They understand this, and so therefore, the accelerators they build, they can make these sort of trade -offs in a way that's most optimal, and also design the architecture of the model itself in a way that reflects where are the hardware trade -offs. Another is Nvidia, because it has, I don't know how this works, but presumably they have some sort of, they're accumulating all this knowledge about how to better design this architecture, and also better search tools for it so on.
1:03:36Who has basically better motier in terms of, will Nvidia keep getting better at design, getting this 100x improvement, or will it be open AI and Microsoft and Amazon and then Thropic or designing the direct accelerators will keep getting better at designing the accelerator? I think that there's a few vectors to go here. One, as you mentioned, and I think it's important to note, is that hardware has a huge influence on the model architecture that's optimal. And so it's not a one way street that better chip equals, you know, the optimal model for Google to run on TPUs, given a given amount of dollars, a given amount of compute is different architecturally than what it is for OpenAI with envidaged stuff, right?
1:04:17It is like absolutely different. And then like even down to like networking decisions that different companies do, and data center design decisions that people do, the optimal, like if you were to say, you know, X amount of compute of TPU versus GPU, compute optimally, what is the best thing? you'll diverge in what the architecture is. And I think that's important to know, right? We can ask about that real quick. The, so earlier we were talking about how China has the H20s or B20s, and there there's like much less compute per memory bandwidth and like the amount of memory, right? Does that mean that Chinese models will actually have like very different architecture and characteristics than American models in the future?
1:04:55So you can take this to like a very large conclude, like leap and it's like all, you know, neuromorphic computing or whatever is like the optimal path and that looks very different than like what a transformer does, right? Or you could take it to like a simple thing which is like the level of sparsity and like course green script varsity, like experts and all this sort of stuff. The arrangement of what exactly the attention mechanism is because there are a lot of tweaks. It's not just like pure transformer attention, right? Or like, hey, demo, like how wide versus tall the model is, right? That's like very important, like demo versus, is a number of layers, right?
1:05:29These are all things that would be different, and I know they're different between, like, say, a Google and an OpenAI and what is optimal. But what really starts to get, hey, if you were limited on a number of different things, like China Invest humongously in computing memory, which is basically the memory cell is directly coupled or is the compute cell, right? So these are things that China's investing hugely and you go to conferences like, oh, there's 20 papers from Chinese companies slash universities about computer memory. Or like, hey, because the flop limitation is here, maybe in video pumps up the on -chip memory and changes the architecture because they still stand to benefit tens of billions of dollars by selling chips to China.
1:06:14Today, it's just like new to American chips, a new to chips that go to the US, but it'll start to diverge more and more architecturally because they'd be stupid not to make chips for China. right? And while they obviously again like has like their constraints right like where are they limited on memory? Oh, they have a lot of networking capabilities and they could move to like certain optical like networking technologies directly onto the chip much sooner than we could right? Because that is what's optimal for them within their search space of solutions right? Because this whole area is like blocked off.
1:06:44It's really interesting to see to think about like the development of how Chinese AI models will differ from American AI Yeah, models because of the, because of these changes or these are the tries to use cases that applies to data. Right? Like American models are very important about like, let me learn from you, right? Let me be able to use you directly as a random consumer, right? That is not the case for Chinese model I assume, right? Because there's probably very different use cases for them. China's crushes the West at video and image recognition, right? I see them all like Albert Gu at, you know, of Cartesian, like state space models.
1:07:19Like every every single Chinese person was like, can I take a selfie with you? Man was harassed. In the US, you see Albert and he's like, it's awesome, he invented state space models, but it's not like state space models are like, like here, but that's because state space models potentially have like a huge advantage in like video and image and audio, which is like stuff that China does more of and that is further along and has better capabilities in, right? So it's like, there are already like, sorry? Because of all the surveillance cameras there. Yeah, that's the, the quiet part out loud, right?
1:07:47But like there's already divergence in like, capabilities there, right? You look at image recognition, China destroys American companies on that, because the surveillance, you have this divergence in tech tree, and people can start to design different architectures within the constraints you're given. And everyone has constraints, but the constraints different companies have are even different. And so Google's constraints have shown them that they built a genuinely different architecture. But now if you look at Blackwell, and then what's set about TPV6, right? They're, I'm not gonna say they're like converging, but they are getting a little bit closer in terms of like how big is the Matt Mul unit size and like some of the like topology and like world size of like the scale up versus scale out network Like there is some like convergence slightly like not saying they're similar yet But like already they're starting to but then there's different architectures that people could go down and path So you see stuff like from all these startups that are trying to go down different tech trees because maybe that'll work But there's a self -fulfilling prophecy here too, right?
1:08:48All the research is in transformers that are very high -rhythmic, madic intensity, because the hardware we have is very high -rhythmic intensity and transformers run really well on GPUs and TPUs. And like, you sort of have a self -fulfilling prophecy. If all of a sudden you have an architecture, which is theoretically it's way better, but you can get only like half of the like usable fobs out of your chip, it's worthless. Because even if it's 30 % you know, compute efficiency when it took twice, it's half as fast on the chip, right? So there's all sorts of like trade -offs and like self -fulfilling prophecies of what do what path do people go down?
1:09:20John and Dylan have talked a lot in this episode about how stupefyingly complex the global semiconductor supply chain is the only thing in the world that Approaches this level of complexity is the Byzantine web of global payments. You're stitching together Legacy text acts and regulations that differ in every jurisdiction In Japan, for example, a lot of people pay for online purchases by ticking a code to their corner store and punching it into a kiosk. Stripe abstracts all this complexity away from businesses. You can offer customers whatever payment experience they're most likely to use wherever they are in the world.
1:09:59And Stripe is how I invoice advertisers for this very podcast. I doubt that they're punching in codes at a kiosk in Japan, but if they are, Stripe will handle it. Anyways, you can head to Stripe .com to learn more. If you are made head of compute of a new AI lab, if like SSI came to you, the Elias Tesco Verneu lab and they're like, Dylan, we give you $1 billion, you are head of compute. Help us get on the map. We're gonna compete with the frontier labs. What is your first step? Okay, so the constraints are you are a US slash Israeli firm because that's what SSI is, right? And your researchers are on the US and Israel.
1:10:38You probably can't build data centers in Israel because power is expensive as hell. And it's probably like risky maybe, I don't know. So still in the US most likely, most of the researchers are here so a lot of them are in the US, right? Like Paul, I'll tour whatever. So I guess you need a significant chunk of compute. You obviously, the whole pitch is you're gonna make some research breakthrough. That's like compute efficiency, when data efficiency, when whatever it is, here we make some breakthrough, but you need compute to get there, right? Because your GPU's per researcher is your research velocity, right?
1:11:12Obviously like data centers are very tapped out, right? None of them are tapped out, but like every new data center that's coming up, most of them have been sold, which has led people like Elon to go through this like insane thing in Memphis, right? I'm just trying to like, I'm just trying to square the circle, yeah. I'm that question. I kid you not in my group house, like group chat, There have been two separate people who have been like, I have a cluster of 800s and I have a long lease on them, but I'm trying to get, sell them off. Is it like a buyer's market right now? Because it does seem like people are trying to get rid of them.
1:11:44So I think like, for the Ilya question is like a cluster of like 256 GPUs or even 4K GPUs, it's kind of, it's kind of cope, right? It's not enough, right? Yes, you're gonna make compute efficiency wins, but with a billion dollars, you probably just want the biggest cluster in one individual spot. Sure. And so like small amounts of GPUs, probably not like, you know, possible to use, right? Like for them, right? Like, and that's what most of the sales are, right? Like you go and look at like GPU list or like vast or like foundry or a hundred different GPU resellers, the cluster sizes are small. Now, is it a buyer's market?
1:12:21Yeah, last year you would buy H100s for like $4 or $3. Like if you, you know, an hour, an hour, right? at a first shorter term or midterm deals. Now, if you want a six -month deal, you can get $2 .15 or less. And the natural cost, if I have a data center, and I'm paying standard data center pricing to purchase the GPUs and deploy them, is a $1 .40 and then you add on the debt because I probably took debt to buy the GPUs or cost equity -caused capital. It gets up to $1 .70 or something. And so you see deals that are the good deals, like Microsoft renting from CoreWeaver, like $1 .90 to $2, right?
1:12:58So people are getting closer and closer to like, there's still a lot of profit, right? Cause the natural rate even after debt and all this is like $1 .70. So like there's still a lot of profit when people are selling in the low twos, like GPU companies, people are deploying them, but it is a buyer's market in a sense that it's gotten a lot cheaper, but cost of compute is gonna continue to tank, right? Because it's like sort of like, I don't know the exact name of the law, but it's effectively Moore's law, right? Every two years the cost of transistors have and yet the industry grew, right? Every six months or three months, the cost of intelligence, you know, like OpenAI and GBDGBD4, what, February, 2023, right?
1:13:37$120 per million tokens or something like that was a roughly the cost and now it's like 10, right? So it's like the cost of intelligence is tanking partially because of compute, partially because of model's compute efficiency wins, right? I think that's a trend we'll see, and then that's gonna drive adoption as you scale up and get, make it cheaper and scale up and make it cheaper. Right, right, right. Anyways, what you're saying, if you're ahead of a computer of SSI. Okay, ahead of a computer of SSI. That's very interesting. There's obviously no free data center lunch, right? In terms of, you know, and you can just, you know, take that based on like the data we see, we have shows that there's no lunch free lunch per se, like immediately today you need the compute for large cluster size, or even six months out, right?
1:14:16There's some, but like not a huge amount, because of what X did, right? X AI is like, oh shit, where are we gonna go? Like, we're gonna go by a Memphis factory, put a bunch of like generators outside, like mobile generators usually reserved for like natural disasters, a Tesla battery pack, drives much power as we can from the grid, tap the natural gas line that's going to the natural gas plant, like two miles away, they can go out natural gas plant, like just like send it and like get a cluster built as fast as possible. Now you're running 100 KGPs, right? I know. And that cost about $5 billion, right?
1:14:48$4 billion, right? Not $1 billion. So scale that SSI has is much smaller, by the way, right? So their size of cluster will be maybe one third or one fourth of the size, right? So now you're talking about 25 to 32K cluster, right? There, you still don't have that, right? No one is willing to rent you a 32K cluster today no matter how much money you have, right? Even if you had more than a billion dollars. So you know, it makes the most sense to build your own cluster one instead of renting it or get a very close relationship like a Open AI Microsoft with Corvieve or open AI Microsoft with Oracle slash Crucio The next step is Bitcoin right So open AI has a data center in Texas right or there it's going to be their data center.
1:15:35It's like the kind of contract and all that Corvieve there is a 300 megawatt natural gas plant on site powering these crypto mining data centers from the company called Core Scientific. And so they're just converting that. There's a lot of conversion, but like the power's already there, the power infrastructure's already there. So it's really about like converting it, getting it ready to be water cooled, all that sort of stuff, and convert it to 100 ,000 GB 200 cluster. And they have a number of those going up across the country, but that's also like tapped out to some extent, because Nvidia is doing the same thing in Plano, Texas for a 32 ,000 GPU cluster that they're building.
1:16:10Is it invidious doing that? Well, they're going through partners, right? Because there's the other interesting thing is the big tech companies can't do crazy shit like Elon did. ESG. Oh, interesting. They can't just do crazy shit like, because this actually do expect Microsoft Google and whoever to like drop their net zero commitments as the scaling picture intensifies. Yeah, yeah. So, so, so, so like this, this like, or what XAI is doing, right? It's not that polluting on the scheme of things, but it's like you have 14 mobile generators and you're just burning natural gas on site on these mobile generators that sit on trucks.
1:16:47And then you have power directly two miles down the road. There's no unequivocal way to say any of the power is because two miles down the road is a natural gas plant as well. There's no way to say this is green. You go to the Corby thing, it's a natural gas plant is literally on site from core scientific and all that. And then the data centers around it are horrendously inefficient. There's this metric called PUE, which is basically how much power is brought in versus how much gets delivered to the chips. Right? And like the hyperscalers, because they're so efficient or whatever, right? Their PUE is like 1 .1 or lower, right?
1:17:19I .e. if you get a gigawatt in, 900 megawatts or more gets delivered to chips, right? Not wasted on cooling and all these other things. This, this like core scientific one is going to be like 1 .5, 1 .6. even like 300 megawatts of generation on site, I only deliver like 180, 200 megawatts to the chips. Given how fast solar is getting cheaper, and also the fact that like, you know, how the reason solar is difficult, the solar is like, you know, you're like, you got to like power the homes at night. Here, I guess it's like theoretically possible to like figure out, you know, only like run the clusters in the day or something.
1:17:56Absolutely not. That is. That is not possible. Because it's so expensive to have these GPUs. Yes. So like when you look at the power cost of a large cluster, it's trivial in some extent, right? Like, you know, like the meme that like, oh, you know, you can't build a data center in Europe or East Asia because the power is expensive. That's not really relevant. What's the re or power is so cheap in China and the US, that's where the only place is you can build data centers. That's not really the real reason. It's the ability to generate new power for these activities. It's why it's really difficult and the economic regulation around that.
1:18:29But the real thing is like, Like if you look at the cost of ownership of a GP of an H100, let's just say you gave me a billion dollars and I already have a data center, I already have all this stuff. I'm paying regular rates for the data centers when I'm paying through the nose or anything, paying regular rates for power and I'm paying through the nose. Power is sub 15 % of the cost. And it's sub 10 % of the cost actually, right? The biggest like 75 to 80 % of the cost is just the servers, right? And this is on like a multi year including debt financing, including cost of operation, all that, right?
1:18:58Like when you do a TCO, total cost of ownership, it's like 80 % is the GPUs, 10 % is the data center, 10 % of the power, rough numbers, right? So it's like kind of irrelevant, right? Whether or not you like, like how expensive the power is, right? You'd rather do what Taiwan does, right? When like power, like what do they do when there's droughts, right? They like, like force people to not shower. They basically reroute the power from, when there was a power shortage in Taiwan, they basically reroute a power from the residentials. And this will happen in a capitalistic society as well, most likely because like, that's the question.
1:19:31You're not gonna pay X dollars per kilowatt hour because to me, the marginal cost of power is irrelevant. Really, it's all about the GPU cost and the ability to get the power. I don't wanna turn it off eight hours a day. Maybe let's discuss what would happen if the training regime changes and if it doesn't change. So like you could imagine that the training regime becomes much more paralyzable where it's like about like coming up with some sort of like search or something that like most of the compute for training is used to come up with synthetic data or do some kind of search and that can happen across a wide area.
1:20:04In that world, how fast could we skate? Just like, let's go through the numbers on like year after year. And then suppose it actually has to be, you would know more than me, but like, suppose it has to be the current regime and like just explain what that would mean in terms of how distributed that would have to be. And then how plausible it is to get clusters of certain sizes is over the next two years. I think it is not too difficult for Ilya's company to get a cluster of 32K and of Blackwell next year. For now, you're gonna be a let's talk about that. Okay, okay, we're not for now. Like 2025, 2026, 2026.
1:20:382026, before I talk about the US, I think it's important to note that there is a gigawatt plus of data center capacity in Malaysia next year now. That's mostly by dance. But in power wise, is there's like, there's the humongous damning of the Nile in Ethiopia, and the country uses like one third of the power that that dam generates. So there's like a ton of power there to like. How much power does that dam generate? Like it's like over a giga lot. And the country consumes like 400 megawatts or something trivial. And it is like, are people bidding for that power? I think people just don't think they can build a data center in fucking Ethiopia.
1:21:11Why not? I wonder if the dam is filled yet, is it? I mean, they have to like, the dam could generate that power, they just don't. Oh, good. Right, like there's a little bit more equipment required, but that's like not too hard. Why don't they? I think there's like, like true security risks, right? If you're China or if you're the US lab like to build a fucking data center with all your IP and fucking Ethiopia. Like you want AGI to be an Ethiopia, like you want it to be that accessible. Like people you can't even monitor like like being the technicians in the fucking data center or whatever, right?
1:21:43Or like powering the data center, all these things, like there's so many like, you know, things you could do to like, You could just destroy every GPU in a data center if you want if you just like fuck with the grid, right? Like pretty, like easily I think. People talk a lot about it in the Middle East. There was 100 KAGB, 200 cluster going up in the Middle East, right? And the US, there's like clearly stuff the US is doing, right? Like, you know, G42 is the UAE data center company, cloud company. Their CEO is a Chinese national, or not a Chinese. He's Chinese, basically Chinese allegiance, but I think Oben and I wanted to use the data center from them, but instead like the US forced Microsoft to like, I feel like this is what happened is forced Microsoft to like do a deal with them so that G42 has a 100K GPU cluster, but Microsoft is like administering and operating for security reasons, right?
1:22:32And there's like Omnivan in like Kuwait, like the Kuwait like super rich guy spending like five plus billion dollars on data centers, right? Like you just go down the list, like all these countries Malaysia has, you know, So, you know, $10 plus billion of like data center, you know, AI data center buildouts over the next couple of years, right? Like, and you know, go to every country, it's like, this stuff is happening. But on the grand scheme of things, the vast majority of the computers being built in the US, and then China, and then like Malaysia, Middle East, and Mike rest of the world. And if you're in the, you know, going back to your point, right?
1:23:03Like you have synthetic data, you have like this search stuff, you have like, you have all these post training techniques, you have all this, you know, all this ways to soak up flops or you just figure out how to train across multiple data centers which I think they have. At least Microsoft and OpenAI have figured out. What do you think they figured it out? They're actions. So Microsoft has signed deals north of $10 billion with fiber companies to connect their data centers together. There are some permits already filed to show people are digging between certain data centers. So we think with fairly high accuracy, we can say we think that there's five data centers, massive, not just five data centers, five like regions that they're connecting together, which comprises of many data centers, right?
1:23:47What will be the total power usage of the? Depends on the time, but easily north of a giga lot, right? Which is like close to a million GPUs. Well, the each GPU is getting more power higher power consumption too, right? Like it's like, you know, the rule of thumb is like GPU, each one hundred is like seven hundred watts, but then total power per GPU all in is like 121300 watts, 1400 watts, but next generation Nvidia GPUs are, it's 1200 watts for the GPU, but then it actions of being like 2000 watts all in, right? So there's a little bit of scaling of power per GPU, but you already have 100k cluster, right?
1:24:22Open AI and Arizona, XAI and MEMFES, and many others already building 100k clusters of H100s. You have multiple at least five, I believe, GB200, 100K clusters being built by Microsoft, slash OpenAI, slash other partners for them. Then potentially even more, 500K GB200s is a giga lot. That's like online next year, right? And like the year after that, if you aggregate all the data center sites, and like how much power, and you only look at NetAd since 2022, instead of like the total capacity at each data center, then you're still like north of multi -gig -a -lot. So they're spending 10 plus billion dollars on at least fiber deals with a few fiber companies, Lumens, AEO, like, you know, a couple other companies.
1:25:07And then they've got all these data centers that they're clearly building 100 K clusters, right? Like, old crypto mining site with Corvieve in Texas or like this Oracle, Crusoe in Texas, and then like in Wisconsin and Arizona and you know, a couple other places. There's a lot of data centers being built up, and providers, right? QTS and Cooper and like, you know, you go down the list, there's like so many different provide and self build, right? Data centers, I'm building myself. So, so, uh, uh, uh, Giga wants, yes. Let's just like, you give the number on like, okay, 2025, Elon's cluster is gonna be the big, like, oh, it doesn't matter who it is.
1:25:42So, so then there's the definition game, right? Like, Elon claims he has the largest cluster at 100 KGPs because they're all fully connected. I don't know who it is, like, I just want to know like, how many, like, I don't know if it's better to to denominate and. 100 ,000 GPUs this year. Okay. For the biggest cluster. For the biggest cluster. Next year. Next year, 300 to 500 ,000, depending on whether it's one side or many, right? 300 to like 700 ,000, I think is a upper bound of that. But anyways, like, you know, it's about like when they tiered on, when they can connect them, when the fibers connect it together.
1:26:13Anyways, 300 to like 500 ,000, let's say, but those GPUs are two to three X faster, right? versus the 100K cluster. So on an H100 equivalent basis, you're at a million chips next year. But then cluster. By the end of the year, yes. No, no, no, well, so one cluster is like, but you know what I mean? The wishy washy definition, right? Multi -site, right? Can you do a multi -site? What's the efficiency loss when you go multi -site? Is it possible at all? I truly believe so. What is it, whether it's, well, what's the efficiency loss as a question, right? Okay, it would be like 20 % loss, 50 % loss.
1:26:45the great questions. This is where you need the secrets, right? And Anthropics got similar plans of the Amazon and you go down the list, right? So then, and then the year after that. The year after that is where, this is 2026. 2026, there is a single gigawatt site, and that's just part of the like multiple sites, right? For Microsoft, the Microsoft FireGear Goat thing happens in 2026. One gigawatt one site in 2026, but then you have a number of others. You have five different locations, each with multiple, some with multiple sites, some with single site. You're easily north of two, three gigawatts.
1:27:22And then the question is, can you start using the old chips with the new chips? And like the scaling, I think, is like, you're gonna continue to see flop scaling like much faster than people expect. I think as long as the money pours in, right? Like that's the other thing is like, there's no fucking way you can pay for the scale of clusters that are being planned to be built next year for OpenAI until unless they raise like $5 ,200 billion. which I think they will raise that end of this year, early next year. A total of 100 billion. Yes. Are you kidding me? No. Oh my God. This is like, Sam has a superpower, no?
1:27:53It's like recruiting and raising money. That's like what he's like a god at. Will ships themselves be a bottleneck to the scaling? Not in the near term. It's more like going back to the concentration versus decentralization point. Yeah. Because the largest cluster is 100 ,000 GPUs. Nvidia's manufactured close to 6 million hoppers, right? across last year and this year. Right, so like what that's fucking tiny, right? So then why is Sam talking about the 7 trillion to build foundries and whatever, like? Well, this is this, you know, like draw the line, right? Like log, log, log lines. Let's fuck, number goes up, right?
1:28:26You know, if you do, if you do that, right? Like you're going from 100K to 300 to 500K, where the equivalent is a million, you just 10x year on year. Do that again, do that again, or more, right? If you increase the pacing, what is do that again? So like 2026, like the number of H -206. If you increase the globally produced flops by 30X, you're on year, or 10X, you're on year, and the cluster size grows by 38 to 5 to 7X. And then you get multi -site going better and better and better. You can get to the point where multi -million chip clusters, IE, they're even if they're like regionally not connected right next to each other, are right there.
1:29:07And in terms of flops, it would be 1E what? 2038, 2039. I think 2038 is like very possible, like 2829. Wow. Okay. And 2030 you said by 2829. Yeah. And so that is literally six orders of magnitude. That's like 100 ,000 times more compute than GPD4. The other thing to say is like the way you count flops on a training run is really stupid. Like you can't just do like parameter, active parameters times tokens times six, right? Like that's really dumb because like the paradigm as you mentioned, right? is like, and you've had many great podcasts on this synthetic data and like RL stuff, post -training, like verifying data and like all these things generating and throwing it away, like all sorts of stuff, search, like inference time compute, all these things like aren't counted in the training flops.
1:29:54So you can't like say 1 ,30 is a really stupid number to say because by then the, you know, the actual flops of the pre -training may be X, but the data to generate the, for the pre -training maybe way bigger or the search inference time, maybe way, way bigger, right? But also because you're doing this adversarial synthetic data where the thing you're weakest at, you can make synthetic data for that. It might be way more sample efficient. So even though the future and the flops will be irrelevant, I actually don't think pre -training flops will be 1 -E30. I think more reasonably, it'll be the total summation of the flops that you deliver through the model across pre -training, post -training, synthetic data for that pre -training data, and post -training data, as well as some of the inference time compute efficiencies could be like, it's more like 1 E 30.
1:30:40Right. So suppose you really do get to the world where it's worth investing. Okay, actually, if you're doing 1 E 30, is that a trillion dollar cluster, a hundred billion dollar cluster? I think it will be multi -hundred billion dollars. And then... But then it'll be... I truly believe people are going to be able to use their of prior generation clusters and alongside their new generation clusters. And obviously like smaller batch sizes or whatever, right? Like, or use that to generate and verify data, all these sorts of things. And then for 2030, right now, I think 5 % of TSMCs and 5 is in video or like whatever percent it is, by 2028, what percentage will it be?
1:31:24Oh, again, this is like a question of like, how scale -pilled you are and how much money will flow into this and how you think progress works. like will models continue to get better? Or does the line slow -over? I believe it'll continue to like skyrocket in terms of keeping that world. In that world, I wouldn't like, oh, not a five nanometer, but like of two nanometer A16, A14, these are the nodes that'll be in that timeframe of 2028, used for AI, I could see like 60, 70, 80 % of it. Like yeah, no problem. Given the fabs that are like currently planned and are currently being built, that is, is that enough for the 130 or will we be able to?
1:31:57I think so, yeah. So then like the chip code is making sense. Because like the chip goes off about like we don't have enough compute. So no, I think like the plans of TSMC on 2 nanometer and such are like quite aggressive for a reason, right? Like to be clear, Apple, which has been TSMC's largest customer, does not need how much 2 nanometer capacity they're building. They will not need A16. They will not need A14, right? Like you go down the list. It's like Apple doesn't need this shit, right? Although they did just hire one of Google's head of system design for TPU, but they are going to make it accelerator.
1:32:34But that's besides the point, and it accelerates. But that's besides the point, Apple doesn't need this for their business, which they have been 25 % or so of TSMC's business for a long time. And when you just zone in on just the leading edge, they've been more than half of the newest node, or 100 % of the newest node almost constantly. That paradigm goes away, right? If you believe in scaling and you believe in like, the models get better, the new models will generate, you know, infinite, not infinite, but like amazing productivity gains for the world and so on and so forth. And if you believe in that world, then like TSMC needs to act accordingly, and the amount of silicon that gets delivered needs to be there.
1:33:10So 25, 26, TSMC is like definitely there. And then on a longer time scale, the industry can be ready for it. But it's gonna be a constant game of like, you must convince them. constantly that they must do this. It's not like a simple game of like, oh, you know, if people work silently, it's not gonna happen, right? Like they have to see the demonstrated growth over and over and over and over again on across the industry. And then you'll notice investors or companies or more so like TSMC needs to see Nvidia volumes continue to grow straight up, right? And oh, and Google's volumes continue to grow straight up and you know, go on down the list.
1:33:49Chips in the near term, right? Next year, for example, are less of a constraint than data centers, right? And likewise for 2026, the question for 2728 is like, you know, always when you grow super rapidly, like people wanna say, that's the one bottleneck because that's the convenient thing to say. And in 2023, there was a convenient bottleneck, COOS, right? The picture's got much, much like cloudier, not cloudier, but we can see that like, No, HBM is a limiter too. Co -Oscis is well -co -oscel, especially, right? Data centers, transformer substations, like power generation, batteries, like UPSs, like CRHs, like water cooling stuff, like all of this stuff is now limitations next year and the year after.
1:34:32Fabs are in 26, 27, right? Like, you know, things will get like cloudy because like the moment you unlock one, oh, like only 10 % higher, the next one is the thing. And only 20 % higher, the next one is the thing. So today, like data centers are like four to 5%. of total US. Of total US. When you think about like as a percentage of US power, that's not that much, but when you think US power has been like this and now you're like this. But then you also flip side, you're like, oh, all this coal's been curtailed, all these like, oh, there's so many like different things. So like power is not that crazy on a like, on a national basis.
1:35:02On a localized basis it is because it's about the delivery of it. Same with the substation transformer supply chains, right? It's like these companies have operated in an environment where the US power is like this or even slightly down, right? And it's like kind of been like that because of efficiency gains because of you know, so you know So anyways like there have been humongous like Weakening of the industry But now all of a sudden if you tell that industry your business will triple next year if you can produce more Oh, but I can only produce 50 % more. Okay fine. You're after that now we can produce three acts as much right you do that to the industry the US Industrial base as well as the Japanese as well as like you know all across the world can get revitalized much faster then people realize, right?
1:35:45Like, I truly believe that people can innovate when given the need to. It's one thing if it's like this is a shitty industry where my margins are low and we're not growing really and like, you know, blah, blah, blah, blah, blah, to all of a sudden, oh, this is the sex, I'm in power and I'm like, this is the sexiest time to be alive and like we're gonna do all these different plans and projects and people have all this demand and they're like begging me for another percent of efficiency advantage because that gives them another percent to deliver to the chips. Like all these things, or 10 % or whatever it is, like you see all these things happen and innovation is unlocked.
1:36:18And you also bring in like AI tools, you bring in like all these things, innovation will be unlocked. Production capacity can grow not overnight, but it will on six months, 18 months, three year time scales. It will grow rapidly. And you see the revitalization of these industries. So, but I think like getting people to understand that, getting people to believe because, you know, We pivot to like, I'm telling you that Sam's gonna raise $5 ,200 billion because he's telling people he's gonna raise this much, right? Like literally having discussions with sovereigns and like Saudi Arabia and like the Canadian pension fund and like not these specific people, but like the biggest investors in the world and of course Microsoft as well, but like he's literally having these discussions because they're gonna drop their next model or they're gonna show it off to people and raise that money.
1:37:03But because this is their plan. If it was like - These sites are already planned and like, the money's not there, right? So how do you like bot plan? It's like without, today Microsoft is taking on immense credit risk. Like they've signed these deals with all these companies to do this stuff, but Microsoft doesn't have, I mean they could pay for it, right? Microsoft could pay for it on the current time scale. Oh, what's, what's, you know, their catbacks going from $50 billion to $80 billion direct catbacks and then another 20 billion across like Oracle Core Web, you know? And then like another like 10 billion across their data center partners, they can afford that, right?
1:37:38to next year, right? But then that doesn't, you know, like this is because Microsoft truly believes in OpenAI. They may have doubts like Holy shit, we're taking on a credit risk. You know, obviously they have to message Wall Street, all these things, but they are not like, that's like affordable for them, because they believe they're a great partner to OpenAI that they'll take on all this credit risk. Now, obviously OpenAI has to deliver, they have to make the next model, right? That's way better. And they also have to raise the money. And I think they will, right? I truly believe from like how amazing 4 .0 is, how small it is relative to four.
1:38:10The cost of it is so insanely cheap. It's much cheaper than the API prices. Lead you to believe and you're like, oh, what if you just make a big one? It's like very clear what's gonna happen to me. On the next jump, they can then raise this money and they can raise this capital from the world. This is intense. That's very intense. John, if he's right, or I don't know, if not him, but like in general, if the capabilities are there, the revenue is there. The revenue doesn't matter. or revenue matters. Is there any part of that picture that still seems wrong to you in terms of like displacing so much of TSMC production, wafers and like power and so forth, does that any part of that seem wrong to you?
1:38:48I can only speak to the semiconductor part, even though I'm not an expert, but I think the thing is like TSMC can do it. Like they'll do it. I just wonder, though he's right in that in a sense of 24 or 25, that's covered. But 26, 27, that's that secret point where you have to say, can the semiconductor industry, and the rest of the industry be convinced that this is where the money is? Like, where's money is? Like, and that means, is there money? Is there money by 24 or 25? How much haven't you done? Do you think the AI industry is a whole needs by 25 in order to keep scaling? Doesn't matter. Compared to smartphones.
1:39:20Compared to smartphone. But I know he says it doesn't matter. I'll get to the line. You keep, I know. What is smartphone? Like Apple's revenue is like $200 something billion. So like, yeah, it needs to be another smartphone -sized opportunity, right? Like, even the smartphone industry He didn't drive this sort of growth. Like it's crazy, don't you think? So today is so far right? The only thing I can really perceive, yeah, go for it. But like that's a lot of it. But you know what I mean? It's not that I want to reel in, David. It's not that I want to reel in. So like a few things, right? The return on invested capital for all of the big tech firms is up since 2022.
1:39:56And therefore, it's clear as day that them investing in AI has been fruitful so far, right? for the big tech firms. Return on invested capital. Like financially you look at the, you look at Meta's, you look at Microsoft's, you look at Amazon's, you look at Google's. The return on invested capital is up since 2022. So it's on AI in particular? No, just generally as a company. Now obviously there's other factors here. Like what is Meta's ad efficiency? How much of that is AI, right? Super messy, that's a super messy thing. But here's the other thing. This is Pascal's wager, right? This is a matrix of like, do you believe in God?
1:40:29Yes or no? If you believe in God, yes or no, like hell or heaven, right? So if you believe in God and God's real and you go to heaven, that's great. That's fine. Whatever. If you don't believe in God and God is real, then you're going to hell. This is the deep technical analysis. You'll subscribe to semi -analysis. This is just, this is just the real thing. Can you imagine what happens to the stock if Sart yes, start talking about Pascal's wager? Don't know, but this is psychologically what's happening. Right? This is a, if I don't, and Sart yes, said it on his earnings call, The risk of under investing is worse than the risk of over investing.
1:41:01He said this word for word, this is Pascal's wager. I must believe I am AGI -pilled because if I'm not in my competitor, does it? I'm absolutely fucked. Okay, other than Zach, who seems... No, not pretty convinced. Zach, Sundar said this on the, on the, on the, on the, on the, on the, Ernst Hall. So, Zach said it, Sundar said it, such as actions on credit risk for Microsoft, do it, he's very good at PR and like messaging, so he hasn't like said it so openly, right? Sam believes it, Dario believes it. You look across these tech titans, they believe it. And then you look at the capital holders, the UAE believes it, Saudi believes it.
1:41:35How do you know the UAE is, I don't believe it. Blackstone believes it. All these major companies and capital holders also believe it, because they're putting their money here. But that's like, how can, like, it won't lie. It can't last unless there's money coming in somewhere. Correct, correct. But then the question is, the simple truth is like, GPT -4 costs like $500 million to train. I agree. And it is generated billions in reoccurring revenue. But in that meantime, opening I raised $10 billion or $13 billion and is building a model that costs that much. Effectively, right? Right. And so then obviously they're not making money.
1:42:09So what happens when they do it again? They release and show GPT -5 with whatever capabilities that make everyone in the world like holy fuck. Obviously the revenue takes time after you release the model to show up. You still have only a few billion dollars or you know, five billion dollars of revenue run rate. You just raised $5 ,200 billion because everyone sees this like holy fuck. This is gonna generate tens of billions of revenue. But that tens of billions takes time to flow in, right? It's not an immediate click, but the time where Sam can convince, and not to Sam, but like people's decisions to spend the money are being made are then, right?
1:42:41Like so therefore, like you look at the data centers, people are building, you don't have to spend most of the money to build the data center, most of the money's the chips, but you're already committed to like, like, oh, I'm just gonna have so much data center capacity by 2027 or 2026 that it's, I'm never gonna need to build a data center again for like, three, four, five years if AI's not real, right? That's like basically what they're all their actions are or I can spend over $100 billion on chips in 26 and I can spend over $100 billion on chips in 27, right? So these are the actions people are doing and the lag on revenue versus when you spend the money or raise the money, spend the money built, there's like a lag on this.
1:43:16So this is like, you don't necessarily need the revenue in 2025 to support this. You don't need the revenue in 2026 to support this. You need the revenue in 2526 to support the $10 billion that OpenAI spent in 23 or Microsoft spent in 23 slash early 24 to build the cluster, which then they trained the model in mid 24, you know, for early 24, mid 24, which they then released at the end of 24, which then started generating revenue in 2526. I mean, like, not only I can say is that you look at a chart with three points on a graph. GPT -1, 2, 3, and then you're like, even that graph is like the investment you have to make in GPT -4 or GPT -3 is 100x.
1:43:53The investment you had to make in GPT -5 or GPT -4 is 100x. So revenue, currently the ROI could be positive, and this very well could be true. I think it will be true. But the revenue has to increase exponentially, not just like, you know, 10th or so. Of course. I agree with you, but I also agree in dealing with this, that it can be a cheap ROI, like, like, send me TSMC does this. Invest $16 million in expects a ROI does that, right? That's, I understand that. That's fine. Lag all that. The thing that I don't expect is that GPT -5 is not here. It's all dependent on GPT -5 being good. If GPT -5 sucks, if GPT -5 looks like, it doesn't blow people's socks off, this is all void.
1:44:38What kind of socks you're wearing, bro? Show them AWS. AWS. AWS. TP -5 is not here. It's late. We don't know. I don't think it's late. I think it's late. I want to zoom out and like go back to the end of the decade picture again. So if you're, if this picture you've made it up, we've already lost John. We've already accepted TP -5 would be good. But yeah, you got it. Yeah. Bro, like life is so much more fun when you just like are delusionally like, you know. You were just ripping balls, bitch. How are we? When you feel the AGI, you feel your soul. This is why I don't live in San Francisco. I have tremendous belief in like, GPD five area because like what we've seen already.
1:45:25I think the public signs all show that this is like very much the case, right? What we see with beyond that is more questionable and I'm not sure because I don't know what, I don't know, right? Like I don't know. We'll see how much they progress. But if things continue to improve, life continues to radically get reshaped for many people. It's also like every time you increment up the intelligence, the amount of usage of it grows hugely. Every time you increment the cost down of that amount of intelligence, the amount of usage increases massively. As you continue to push that curve out, that's what really matters, right?
1:46:04And it doesn't need to be today, it doesn't need to be a revenue versus how much cat -backs. In any time in the next few years, it just needs to be, did that last humongous chunk of CapEx make sense for OpenAI or whoever the leader was, and then how does that flow through? Or were they able to convince enough people that they can raise this much money? You think Elon's tapped out of his network with raising $6 billion? No. XAI is going to be able to raise 30 plus, easily. I think so. You think Sam's tapped out? You think Anthropics tapped out? Anthropics barely even diluted the company relatively.
1:46:36right? Like, you know, there's a lot of capital to be raised in just from like, call it FOMO if you want, but like during the dot com bubble people were spending the private industry flew through like $150 billion a year. We're nowhere close to that yet. Right? We're not even close to the dot com bubble, right? Why would this bubble not be bigger? Right? And if you go back to the prior bubbles, PC bubble, semiconductor bubble, mechatronics bubble, throughout the US, each bubble was smaller. You know, you know, a bubble or not, why wouldn't this one be bigger? How many billions of dollars a year is this bubble right now?
1:47:09For private capital? Yeah. It's like $55, $60 billion so far. For this year, it can go much higher, right? And I think it will next year. Okay, so let me think about this. Do you know the Bungram? You know, at least like finishing up and looping into the next question was like, you know, prior bubbles also didn't have the most profitable companies that humanities ever created investing, and they were debt finance. This is not debt finance yet, right? So that's the last little point on that one. Whereas the 90s bubble was very debt financed. This is like, I was exactly for those companies. Yeah, sure, but it was.
1:47:44Zasters. So much was built, right? You got to blow a bubble to get real stuff to be built. I think it is an interesting analogy where like, with even though the Doc Cumbubble, I've always been bursting like a lot of companies when bankrupt, they in fact did lay out the infrastructure that enabled the web and everything. So you're gonna imagine in an AI, it's like some of a lot of the foundation or whatever, like a bunch of companies will like go bankrupt, but like they will... You could unable the... During the 1990s, that the turn of 1990s, it was a immense amount of money invested in like, memes and like, optical technologies, because everyone expected the fiber bubble to continue, right?
1:48:18That all ended at 2003, to the two of us, you know what I'm saying? Right? And that started in 1994. It hasn't been revitalization since, right? Like that's, you could risk the possibility of... For women, one of the companies that's doing the fiber build -out for Microsoft, the stock like fucking Forex last month or this month. And then how's it done from 2002 to 2000? Oh no, horrible, horrible. But like, we're gonna rip, babe. You could rip that out, maybe? You could freeze AI for another two decades. You sure, sure, possible. Or people can see a badass demo from GPD5, slight release, raise a fuckload of money.
1:48:52It could even be like a devin like demo, right? Where it's like complete bullshit, but like it's fine, right? Like, shit, I should. I didn't know. I didn't know. No, it's fine, it's fine. I did, I don't really care. You know, it's, the capital is gonna flow in, right? Now, whether the, whether the flights are not as like an irrelevant concern on the near term because you operate in a world where it is happening. And being, you know, being, you know, what is the Warren Buffett quote, which is like, you can be, and I don't even know it's Warren Buffett. You don't know who's, you don't know who's just gonna be naked until the tide goes out.
1:49:24No, no, no, no. The one about like, the market is delusional far longer than you can remain solvent. or something like that. That's not Buffett. That's not Buffett. Yeah, yeah, yeah. That's John Maynard Keynes. Oh, shit, that's that old. Yeah. Okay. Okay, so Keynes said it, right? It's like, you can be, yeah, so this is the world you're operating in. Like it doesn't matter, right? Like what exactly happens? Or we have some flows, but like that's the world you're operating in. I reckon that if an AI bubble pops, each one of these CEOs lose their jobs. Sure. Or if you don't invest and you lose, It's Pascalian Wager and you're, that's much worse.
1:50:01Across decades, the largest company at the end of each decade, like the largest companies, that list changes a lot. And these companies are the most profitable companies ever. Are they going to let that list, are they going to let themselves like lose it, or are they gonna go for it? They have one shot, one opportunity. You know, to make themselves into, you know, the whole M &N song, right? I wanna hear like the story of how both of you started your businesses or you're like, Like the thing you're doing now, John, like how did it begin? What were you doing when you started the podcast? You just had a textile company?
1:50:34Oh my god, no way. Please, please. Were you ever, are you joking? I guess if he doesn't want to, we'll talk about it later. Okay, sure. I think like I used to, I mean, the story's famous. I've told it a million times. It's like Asian Optory Startoff is a tourist channel. Yeah. So I would go around kind of like, I moved to Taiwan for work and then doing what? I was working in cameras. And then like I told you the other company you started? It tells too much about me. Oh, come on. I worked in cameras and then basically I went to Japan with my mom and mom was like, hey, you know, what are you doing in Taiwan?
1:51:14I don't know what you're doing. I was like, all right, mom, I will go back to Taiwan and I'll make stuff for you. And I made videos. I would like go to the Chiang Kai Shaq Park. I'm gonna be like, hi mom, this park was this, this. And actually at some point you run out of stuff, but then it's like a pretty smooth transition from that into like history of Chinese history, Taiwanese history, and then people started calling me China nometry. I didn't like that, so I moved to other parts of Asia and now like, and then. So what year did you like start? Like what year was like people started watching your videos?
1:51:44Let's say like a thousand views per video or something. Oh my gosh, that was not, I started the channel in 2017 and it wasn't until like 2018 that 2019 that actually, I labored on for like three years, first three years with no one watching. Like I got like 200 views and I'd be like, oh this is great. And then were you, were the videos basically like the ones you have, but I'm so sorry, backing up for the audience who might not, I imagine basically everybody knows Asian Amitri, but if you don't, like the most popular channel about semiconductors, Asian business history, business history in general, We even like, Jew politics history and so forth.
1:52:20And yeah, I mean, it's like, honestly, I've done research for like different AI guests and different, like whatever thing I'm trying to, be, I'm trying to understand like, how does harder work, how does AI work? It's like, this is like my bed. How does a zipper work? Did you watch that video? No, I watched that one. I think it was a span of three videos. It was like Russian oil industry in the 1980s and how it like funded everything. And then when it collapsed, they were absolutely fucked. Yeah. And then it was like the next video was like, the zipper monopoly in Japan. The next video was about it.
1:52:47It's an monopoly. Strong holding in a mid -tier size. There's like the luxury zipper makers. Asian armature is always just kind of like stuff I'm interested in. And I'm like interested in a whole bunch of different stuff. And I like, like in the channel, for some reasons, people started watching the stuff I do. And I still have no idea why. To be honest, I still feel like it's, I still feel like a fraud. I sit in front of Dylan and he's, I feel like a fraud, legit fraud, especially when he starts talking about 60 ,000 wafers and all that, I'm just like, I feel like I should be known, I should know this, but like, you know, in the end it's, yeah.
1:53:21But then, you know, I just try my best to kind of bring interesting stories out. How do you make a video every single week? Because these are like two a week. You know how long he had a full -time job? Five years, six years. Or sorry, eight textile business. And a full -time job, wait, no. Full -time job textile business and a geometry until like for a long, long time. I literally just gave up the textile business this year. And like, how are you doing research and doing like making a video and like twice a week? I don't know. I like to do these fucking, I'm like fucking talking. This is all I do.
1:53:51And I like to do these like once every two weeks. Sorry. See the difference is Dwarkesh. You go to SF Bay Area parties constantly. Dwarkesh is the, I mean, John is like locked in. Yeah, yeah. He's like locked in 24 seconds. I believe that the agency work ethic. and I've got like the Intel work ethic. I don't, I got the Huawei ethic. If I do not finish this video, my family will be pillaged. He actually gets really stressed about it, I think, like not doing something like on a schedule, yeah. It's very much like, I do two videos a per week. I write them both simultaneously. And how are you scouting out future topics you wanna do, or these are just like, you know, you just like pick up random articles, books, whatever, and then you just, if you find it interesting, you make a video about it.
1:54:34Sometimes what I'll do is I'll Google a country, I'll Google an industry, and I'll Google what the country is exporting now and what it used to export. And I compare that and I say, that's my video. Or I'll be like, or but then sometimes also just a simple is like, I should do a video about YKK. And then it's also just a simple, simple, simple, simple, nice video about it. I do it. I literally just like, do you like keep a list of like, here's the next one, here's the one after that. I have a long list of like ideas. Sometimes it says vague as like Japanese whiskey. No idea what Japanese whiskey is about.
1:55:10I heard about it before, I watched that movie and then so I was just like, okay, I should do a video about that. And then eventually, you get to, you get, you move that. How many research topics do you have in the back burner? Basically, you're like, I'm kind of reading about it constantly and then like in a month or so, I'll make a video about it. I just finished a video about how IBM lost the PC. So right now I'm de, I'm unstressing about that But then I'll kind of move right on to like the videos do kind of lead into others. Right. Like right now this one is about IBM PC. IBM lost the PC. Now it's next is how compact collapse, how the wave destroyed compact.
1:55:43So technically that I'll do that at the same time. I'm dual lining a video about two bits. I'm dual lining a video about. So directed self assembly for semiconductor manufacturing, which I'll read a lot of Dylan's work for. But then like, like a lot of that is kind of like, it's just, it's in the back of my head. And I'm like, producing it as I, as I go. Um, Dylan, how do you work? How does one go from Reddit shit poster to like, running a, like a semiconductor research and consulting firm? Yes. I know. Let's start with the shit posting. It's a long line, right? Like, so immigrant parents grew up in rural Georgia.
1:56:20So when I was eight, I begged for a seven, I begged for an Xbox and when I was eight, I got it, 360, right? They had a manufacturing defect called the Red Ring of Death. There were a variety of fixes that tried them, like putting a wet towel around the Xbox, something called the Penny Trick. Those old didn't work, my Xbox still didn't work. My cousin was coming the next weekend, and he's two years older than me. I look up to him, he's in between my brother and I, but I'm like, oh no, no, we're friends. You don't like my brother as much as you like me. My brother's more like a jockey type, so I didn't matter.
1:56:50He didn't really care that the Xbox is broken, He's like, you better fix it though, right? Otherwise, parents will be pissed. So I figure out how to fix it online. It ends up, you know, I tried a variety of fixes, ended up shorting the temperature sensor and that worked for a long enough until Microsoft had the recall, right? But in that, you know, I, I learned how to do it out of necessity on the forums. I was an early kid, so I liked games, but whatever. But then like, there was no other outlet. Once I was like, holy shit, this Pandora's box. Like what just got opened up. So then I just shit posted on the forums constantly, right?
1:57:22And for many, many years, and then I ended up moderating all sorts of reddits when I was like a tween teenager. And then, like, you know, as soon as I started making money, you know, a group of family business but didn't get paid for working, right? Of course, like yourself, right? But like, as soon as I started making money, like, I got my internship and I was like, 18, 19, right? I started making money. I started investing in semiconductors, right? Like, of course, this is, should I like, right? everything from like, and by the way, the whole way through, like as technology progressed, especially mobile, right?
1:57:53It goes from like very shitty chips and phones to like very advanced every generation they'd add something. And I'd like read every comment, I'd read every technical post about it. And also all the history around that technology and then like, who's in the supply chain and just kept building and building and building? When a college did data science, he typed stuff, went to work on like Hurricane, earthquake, wildfire simulation and stuff for a financial company, but before that, during college, I was still like, I wasn't posting on the internet as much. I was still posting some, but I was like following the stocks and all these sorts of things, the supply chain, all the way from the tool equipment companies.
1:58:28And the reason I like those is because, oh, this technology, oh, it's made by them, you kind of, do you have friends and person who weren't in this shit? Or was it some like? I made friends on the internet, right? Oh, that's dangerous. No, I've only ever had literally one bad experience, and that was just because he was drugged out. Like a one that executes online or a group? Like meeting someone from the internet in person. Everyone else has been genuine. Like you have enough filtering before that point. You're like, you know, even if they're like hyper mega like autistic, it's cool, right? Like I am too, right?
1:59:00I know, I'm just kidding. But like, you know, you go through like the, you know, the layers and you look at the economic angle, you look at the technical angle, you read a bunch of books just out of like, you know, you can just buy engineering textbooks, right? And read them, right? Like, what's stopping you, right? And if you bang your head against the wall, you learn it, right? And then why were doing this? Was there like, did you expect to work on this at some point? Or was it just like, pure interest? No, it was like, it was like obsessive hobby of many years and it pivoted all around, right?
1:59:28Like, at some point, I really like gaming and then I got moved into like, I really like phones and like rooting them and like, underclocking them and the chips there and like screens and cameras and then back to like gaming and then to like data center stuff. like, because that was like where the most advanced stuff was happening. So it's like, I liked all sorts of like telecom stuff for a little bit. Like it's like it like bounced all around, but generally in like computing hardware, right? And I did data science, you know, you could I Said I did AI when I interviewed, but like, you know, but It was like bullshit mouth multivariable regression, whatever, right?
2:00:02Those simulations of hurricanes are of course a lot of fire for like financial reasons, right? Like anyways, You move I moved up to like, you know, I was still, you know, I worked I had a job for three years after college and it's posting and like whatever, I had a blog, anonymous blog for a long time. I'd even made some YouTube videos and stuff. Most of that stuff is scrubbed off the internet, including internet archive because I asked them to remove it. But like, in 2020, I quite quit my job and started posting more seriously on the internet. I moved out of my apartment and started traveling through the US and I went to all the national parks like in my truck slash, like tent slash, also stayed in hotels and motels like three, four days a week, but I'd like, I started posting more frequently on the internet.
2:00:44I mean, I'd already had like some small consulting arrangements in the past, but it really started to pick up in mid 2020, like consulting arrangements from the internet from my persona. Like what kinds of people, investors, hardware companies? Like, there were like, it was like, it was like people who weren't in hardware that wanted to know about hardware, it would be like some investors, right? Some couple of VCs did it, but some public market folks. You know, there was times where like companies would ask about like three layers up in the stack like me because they saw me write some random posts and I'm like, hey, can we add a bubble?
2:01:15But it's all sorts of like random, it was really small money. And then in 2020, like it really picked up and I just like, I don't know, I just arbitrarily make the priceway higher and it worked. And then I started posting, I made a new, I made a newsletter as well. And I kept posting quality kept getting better, right? Because people read it, they're like, this is fucking retarded, like, you know, there's what's actually right? or over more than a decade, right? And then in 2021, towards the end, I made a paid post, someone paid, and for a report or whatever, right? And it ended up doing, I went to sleep that night.
2:01:52It was about photo -resist and the developments in that industry, which is the stuff you put on top of the way for before you put in the Litho tool. Lithography tool. Did great, right? I woke up the next day and I had 40 paid subscriptions. I was like, what? Okay, let's keep going, right? Let's post more paid sort of, like partially free, partially paid, did all sorts of stuff on advanced packaging and chips and data center stuff and AI chips, all sorts of stuff that I was interested in. That was interesting. I always bridged economically because I read all the companies earnings since I was 18 and 28 now, all the way through to all the technical stuff that I could.
2:02:292022 I also started to just go to every conference I could. So I go to 40 conferences a year, not like trade show type conferences, but like technical conferences. Like an art chip architecture, photo resist. You know, AI NIRIPs, right? Like, you know, I see them all like, how many conferences do you go to a year? Like 40. So you like live at conferences? Yes, yeah. I mean, I've been a digital nomad since 2020, and I've basically stopped and I moved SF now, right? But like kind of, kind of not really. You can't say that, the government, the California government. No, no, no, no, I don't live at SF, come on.
2:03:03But I basically do now, right? I'm not going to have an internal revenue service. Oh, you're not joking about that. Like, you're not seriously joking. They're going to send you a clip of this podcast be like 40 % please. I am in San Francisco, like sub four months a year, continuously, you know, exactly 100 and whatever day. Exactly 179 days. Let's go. Right? You know, over the full course of the year. But no, like, you know, go to every conference, make connections at all these like very technical things, like, international, electron, device manufacturing. Oh, lithography and advanced patterning.
2:03:38Oh, like, very large scale integration. Like, you know, all the circuits, conference. So you just go every single layer of the stack. It's so siloed, there's tens of millions of people that work in this industry. But if you go to every single one, you try and understand the presentations, you do the required reading, you look at the economics of it, you are just curious and want to learn. You can start to build up more and more and the content got better. and what I followed, we had better, and then started hiring people in early 2022 as well. Or might have been, yeah, mid -2022 started hiring, got people in different layers of the stack, but now today, like you fast forward, now today, right?
2:04:19Almost every hyper -scaler is a customer, not for the newsletter, but for data we sell, right? Most many major semiconductor companies, many investors, like all these people are like customers of the data and stuff we sell, and the company has people all the way from like X -Simer XASML all the way to like X -Microsoft and like an AI company, right? Like, you know, like, and then through the stratification, you know, now there's 14 people here in like the company and like all across the US, Japan, Taiwan, Singapore, France, US of course, right? Like, you know, all over the world and across many ranges of like, and hedge funds as well, right?
2:04:54X hedge funds as well, right? So you got kind of have like this It's a malgamation of tech and finance expertise, and we just do the best work there, I think. Are you just talking about a monstrosity? An unholy concoction.
2:05:11So when we saw, we saw, we have data analysis, consulting, et cetera, for anyone who really wants to get deeper into this, right? We can talk about people are building big data centers, but how many chips is being made in every quarter of what kind for each company, what are the sub components of these chips, what are the sub components of the servers, right? We try and track all of that, follow every server manufacturer, every component manufacturer, every cable manufacturer, just like all the way down the stack tool manufacturer, and like know how much is being sold where, and how, and where things are, and project out, all the way out to like, hey, where's every single data center?
2:05:49What is the pace that it's being built out? This is like the sort of data we wanna have and sell, and you know, it's the validation is that hyper -scalers purchase it and they like get a lot, right? And like AI companies do and like semiconductor companies do. So I think that's the sort of like, how it got there to where it is is just like trying to do the best, right? And trying to be the best. If you were an entrepreneur who's like, I want to get involved in the hardware chain somewhere. Like what is, like what is, if you could start a business today, somewhere in the stack, what would you pick?
2:06:21John, tell them about your textile business. I think I'd work in memory, something in memory. Cause I think like if you, if this concept is like there, like you have to hold immense amounts of memory, immense amounts of memory. And I think memory already is tapped like technologically to HBM exist because of limitations in DRAM. I said it correctly. I think like it's fundamentally, we've forgotten it because it is a commodity, but we shouldn't. I think it's breaking memory is going to, could change the world in that scenario. I think the context here is that Moore's Law was predicted in 1965, Intel was founded in 68 and released their first memory chips in 69 and 70.
2:07:08And so Moore's Law was a lot of it was about memory. And the memory industry followed Moore's Law up until 2012, where it stopped, right? And it became very incremental gains since then, whereas logic has continued and people are like, Oh, it's dying, it's slowing down. At least there's still a little bit of like, you know, coming, right? You know, still more than 10%, 15 % a year, cager, right, of growth and density slash cost improvement. Memory is like literally like been like since 2012, like really bad. So, and when you think about the cost of memory, you know, it's been considered a commodity, but memory integration with accelerators, like this is like something that, I don't know if you can be an entrepreneur here though.
2:07:44That's the real challenge is because you have to manufacture at some really absurdly large scale, or design something in an industry that does not allow you to make custom memory devices or use materials that don't work that well. So there's a lot of work there that I don't necessarily agree with you, but I do agree it's one of the most important things for people to invest in. I think it's really about where is your where you go at and where can you vibe and where can you enjoy your work and be productive in society, right? Because there are a thousand different layers of the abstraction stack, where can you make it more more efficient working you use, we utilize AI to build better and make everything more efficient in the world and produce more bounty and iterate feedback loop.
2:08:25And there is more opportunity to today than any other time in human history in my view. And so just go out there and try. What engages you because if you're interested in it, you'll work harder. If you have a passion for copper wires, I promise to God if you make the best copper wires, you'll make a shitload of money. And if you have a passion for like B2B SaaS, I promise to God you'll make fuck loads of money, right? I don't like B2B SaaS, but whatever, right? It's like whatever. Whatever you have a passion for, like just work your ass off, try and innovate, bring AI into it and let it, you try and use AI yourself to like make yourself more efficient and make everything more efficient.
2:09:07And I promise you will like be successful, right? I think that's really the view is not necessarily there's one specific spot because every layer of the supply chain has, you go to the conference there, you go to the talk to the experts there, it's like, dude, this is the stuff that's breaking and we could innovate in this way. Or like these five extraction layers, we could innovate this way. Yeah, do it. There's so many layers where this is, we're not at the cradle optimal, right? Like there's so much more to go in terms of innovation and inefficiency. All right, I think that's a great place to close.
2:09:36Dylan, John, thank you so much for coming on the podcast. I'll just give people the reminder, Dylan Patel, semianalysis .com, that's where you can find the technical breakdowns that we've been discussing today. Asian Nonmetry YouTube channel, everybody who already may or of Asian Nonmetry, but anyways. Thanks so much for doing this. This was a lot of fun. Thank you. Yeah. Thank you.
From the publisher
A bonanza on the semiconductor industry and hardware scaling to AGI by the end of the decade.
Dylan Patel runs Semianalysis, the leading publication and research firm on AI hardware. Jon Y runs Asianometry, the world’s best YouTube channel on semiconductors and business history.
* What Xi would do if he became scaling pilled
* $ 1T+ in datacenter buildout by end of decade
Watch on YouTube. Listen on Apple Podcasts, Spotify, or any other podcast platform. Read the full transcript here. Follow me on Twitter for updates on future episodes.
Sponsors:
* Jane Street is looking to hire their next generation of leaders. Their deep learning team is looking for FPGA programmers, CUDA programmers, and ML researchers. To learn more about their full time roles, internship, tech podcast, and upcoming Kaggle competition, go here.
* This episode is brought to you by Stripe, financial infrastructure for the internet. Millions of companies from Anthropic to Amazon use Stripe to accept payments, automate financial processes and grow their revenue.
If you’re interested in advertising on the podcast, check out this page.
Timestamps
00:08:25 – How semiconductors get better
00:11:16 – China can centralize compute
00:18:50 – Export controls & sanctions
00:32:51 – Huawei's intense culture
00:38:51 – Why the semiconductor industry is so stratified
00:40:58 – N2 should not exist
00:45:53 – Taiwan invasion hypothetical
00:49:21 – Mind-boggling complexity of semiconductors
00:59:13 – Chip architecture design
01:04:36 – Architectures lead to different AI models? China vs. US
01:10:12 – Being head of compute at an AI lab
01:16:24 – Scaling costs and power demand
01:37:05 – Are we financing an AI bubble?
01:50:20 – Starting Asianometry and SemiAnalysis
02:06:10 – Opportunities in the semiconductor stack
Get full access to Dwarkesh Podcast at www.dwarkesh.com/subscribe




