China’s catching up to US AI… Here’s why it won’t matter

14 May 2025 · 49 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Notes: Azeem Azhar's Exponential View

Episode Title

China’s catching up to US AI… Here’s why it won’t matter

Guest

Lennart Heim Date Recorded: October 2023 Podcast Description: Azeem Azhar, founder of Exponential View, engages in discussions to demystify the rapid changes brought by AI and other exponential technologies. This episode features researcher Lennart Heim from the RAND Corporation discussing the implications of China's advancements in AI compared to the United States.

---

Episode Summary Lennart Heim argues that while China is indeed making strides in AI capabilities, it ultimately does not threaten the U.S. due to significant disparities in compute resources. The discussion focuses primarily on the following key areas:

Key Points Discussed

  1. Core Thesis
  2. China may catch up with the U.S. in AI model capabilities but this won't have meaningful consequences due to the U.S.'s substantial lead in compute resources.
  1. Importance of Compute
  2. Compute is likened to currency in AI; more compute enables the development of more advanced AI systems.
  3. The concentration of supercomputers is heavily tilted towards the U.S., which holds 60-70% of them.
  1. Investment in AI
  2. The podcast discusses the investment split between research and development (R&D) of AI models and their execution in real-world applications.
  3. Growing demands for test-time compute, or inference, lead to rising costs.
  1. Geopolitics of Compute
  2. The discussion touches on the geopolitical implications of AI compute resources.
  3. The U.S. has better access to advanced chips due to export controls, giving it an advantage over China.
  1. Economic and National Security Concerns
  2. There's a trade-off between economic needs and national security.
  3. The U.S. has to balance its export controls to prevent sensitive technologies from reaching adversaries while still fostering innovation.
  1. Future Predictions
  2. The conversation speculates on how technological changes might affect battlegrounds in AI and compute.
  3. Emphasis is placed on the need for adequate infrastructure to support AI growth on a national and global scale.

Key Insights

  • Compute as a Resource: Compute capacity is essential for AI progress; countries with greater access will likely progress faster in AI development.
  • AI Diffusion vs. Model Capability: It’s not just about having the best model but also about how widely and effectively AI is deployed across various sectors.
  • Long-term Trends: The next few years will see significant shifts in how AI is utilized, particularly with the rise of inference models that could reshape applications.

---

Conclusion In this episode, Azeem Azhar and Lennart Heim explore the complex landscape of AI advancement, national security, and economic competition. They emphasize that although China is advancing in AI, the U.S.'s substantial compute capacity and the implications of export controls maintain its leading position. The implications of these dynamics will shape the future of AI, making it crucial for policymakers to navigate these complex issues carefully.

---

Links

  • Lennart Heim
  • [Twitter](https://twitter.com/ohlennart)
  • [Personal Blog](https://heim.xyz/)
  • Azeem Azhar
  • [Substack](https://www.exponentialview.co/)
  • [Website](https://www.azeemazhar.com/)
  • [LinkedIn](https://www.linkedin.com/in/azhar)
  • [Twitter](https://x.com/azeem)

---

Timestamps

  • 00:00 – Episode Trailer
  • 01:19 – Lennart’s Core Thesis
  • 03:26 – Importance of Compute
  • 07:31 – Investment Split in AI
  • 11:18 – Cost Implications of Test-Time Compute
  • 16:14 – Geopolitics of Compute
  • 21:32 – U.S. Compute Capacity vs. China
  • 25:01 – Economic vs. National Security Needs
  • 31:54 – Technology Change and Future Battlegrounds
  • 35:33 – Managing Compute and Power Concentration
  • 48:19 – Concluding Quick-Fire Questions

---

This structured markdown file captures the essence of the podcast episode, summarizing the discussions while highlighting key concepts and insights for easy reference.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00We are really lucky to be joined by Lennart Heim. He said China is catching up with the U.S. in AI capabilities and that it didn't matter. What's America's big advantage? Having eight times more of the resource, which is the most important. Let's start on why compute matters and what it drives. I think it's good to say compute is like the currency of AI. Every time it's being produced requires computing power. If you have less compute, you've got less AI workers. And population size is pretty important for economic growth. There is this question about a world where compute resources are highly concentrated.

0:31Typically in political systems, we don't like concentration of power. Let's look at the status quo. I think like 60 to 70 percent of all supercomputers are in the U.S. I'm not an antitrust lawyer, but I think this is at least something one should look into. This also explains in a sense why the EU and why India and also the UAE are pushing for their own compute infrastructure. So this sense of this concentration risk is really being noticed increasingly over the last five years. How do you imagine it plays out? I don't think the plan is like every country is going to like redevelop the whole tech stack and come up with a solution.

1:04Again, maybe I'm a little bit too much kumbaya. It would be nice if we can just get together and just say like, here are the standards. Here's how the democratic world is doing AI.

1:18We are really lucky to be joined by Lennart Haim. He is a researcher and information scientist at RAND, the OG of think tanks. He's been an advisor of EPOC, which is an AI research group, which readers of Exponential View know that we rather like. I have been following his work for a long time, and he recently wrote an essay on the China Talk substack arguing two things. He said China is catching up with the US in AI capabilities and that it didn't matter. So that is something to talk about. Lenard, let's get into that key claim. China may match individual U.S. model capabilities this year, but that won't matter for America.

2:02What's America's big advantage? So the argument which I'm putting forward in this essay is basically inspired what happened with DeepSeek. Everybody freaked out. Some people were just expecting there's like this really big lead and it's like really, really hard to build competitive AI. Don't get me wrong, not everyone can do it. But I think like actually more actors than many people assume can build competitive models. In particular, being like running behind is actually easier than many people expect. What I mostly work on is the importance of compute. And I'm just always asking, well, there's these many data centers, there are millions of GPUs, or how are they being used?

2:34Sure, tens of thousands of them are being used for training systems. And this is what we see more and more is AI being deployed, AI being used. And what I took, what I'm trying to put forward to tease is just like, to some degree, like, guys, actually, it's more complex. We cannot only look at model benchmarks and say, this is the country which is winning, this is the country which is losing, or we only have a three-month skip. There's more to it. in particular if you think about like how do export controls bite so what i'm putting forward there is just like actually you use compute for training they only use part of the company for training they will probably build competitive models in china this year particularly with test time compute being on there but given the compute might be more important in the future or you want to deploy it at scale to get your ai agents or whatever the future holds this is eventually where i think export controls will hold and i think where the us at least has a big advantage how to and leverages is another question.

3:21Wonderful. Okay, so we've got a sketch of the argument. We need to do some level setting. Let's start in the engine room on why compute matters and what it drives. So what exactly is compute in this context? Yeah, I think we use compute loosely here to refer to many different things. But I think more broadly, what we mean with this is like processing units, integrated circuits, which do the computations, right? And when we talk about AI, we talk about GPUs, or AI chips. It started in 2012. LXNet used two GPUs. Now we see Elon Musk building clusters with 200 ,000 of them. I think he's currently right now talking about upgrading to a million GPUs.

4:00And I think it's fair to say computers like the currency of AI. You know, when you're in Silicon Valley, people bolster with how much GPUs they have and how big their model is. And again, it just became like this key ingredient for AI for training the systems, but eventually deploying it. It's like the water as it is for the fish or the air as far as like, This is what these systems need to be run and to be developed. Every time I talk to a researcher at a frontier lab, they're always complaining about the lack of compute, even though these labs have more compute than anyone else, because there is so much work to be done.

4:33And the more compute you have, it's not just that you do things faster. It's actually that you can do new things, right? You can do things at a higher degree of complexity, which may develop new capabilities. One of the things I found staggering is that there is this new exponential in the computer industry that is far faster than the Moore's law exponential, which is the amount of compute that is being used on these frontier models. Can you give us a flavor for what that is like today and how that has evolved over the past few years? Absolutely. Yeah, I loved it to bring it up because that's actually how it got started for me.

5:10I think in 2020, I think OpenAI put out the blog post, AI and compute. And they're like, look, we looked at how much computation we needed to train a system, and it's doubling every 3.4 months. I was like, wow, Moore's law is described as like the fastest exponential ever, doubling every two years. So doubling every 3.4 months is quite, quite staggering. And what I then did, which basically also the original story of Epoch, I found some colleagues and we did like the study, I'm like, well, actually, let's add some more systems because the original AI paper wanted like eight or 10 AI systems there, and we then upgraded to a couple of hundreds.

5:41And what we then found, okay, it's not doubling every 3.4 months anymore. It's doubling every six months. That's still crazy, right? Crossing like multiple orders of magnitude over the last few years. Again, going from two GPUs to 200 ,000 or going from like spending a couple of hundred bucks on computational partners to literally millions, potentially in the future soon, billions. And why are people doing this? Well, because the systems get better. It's not like we want to burn money, right? It's just like we train big AI systems, train in more data. And if you do this, you just require more computing power.

6:10And this has just been what's been called scaling loss. And this has been, I think, the easiest way to describe AI progress over the last decade. It really is a sort of a staggering outcome that we have this quite predictable law that as you double the compute capacity, you get not quite a doubling, right? It's a sublinear relationship to the capabilities of the model, but it's predictable. And the power of a predictable upfront relationship is that you can start in some sense to talk about your ROI. You can kind of go to your CFO and say, well, if you give me this X hundred million dollars, I'm pretty confident that I can push the system this far and give us this result.

6:49And then we can work out whether it makes sense financially or strategically. Although I think at the moment that calculus doesn't exist. It's just get the compute, run bigger models, get higher capabilities. Is that your sense of what the world is like right now? I think to some extent. I mean, like, if you need to justify millions being spent on something, that's probably easier than justifying billions, right? So, and if Ope may I and others announced a 500 billion for Stargate, I hope they did some math, right? Again, but the math here is beyond just training bigger models, even if you just stay at current capabilities, we need to deploy them.

7:21More and more people will be using them. AI agents might be working. They're going to need to be running somewhere. And that's, to some degree, the nice thing about compute. You're not making a specific bet saying, oh, large language models are the thing. You're basically saying like, look, having more of something is useful. One of the things that OpenAI, I think Sam Alton said, was that in a few years, the bulk of the compute would be on inference, right? That is the deployment of these models. And certainly in the modeling that Exponential View has done, we had come to a similar conclusion.

7:49What is your assessment of how the AI compute specifically splits now between the research and development of training new models and actually executing them? when we think about these leading frontier labs? It's hard to say because, of course, they publicly don't tell us. There have actually been some papers coming out of Google where they talked about the environmental footprint and NVIDIA, where I think they asked NVIDIA and AWS. The question is like, what is an AI workload? Are your recommender systems you use for Instagram also AI workload? So if you look at it this way, then AI is like computers predominantly being used for deploying systems, right?

8:24This might look different if you look at a frontier lab, right? But I think like over time, you see like this share of deployment getting higher. And before ChatGPT, there was not a lot of deployment. There were not many users, right? There were like some nerds playing around with the API. But since then, again, I think many people say the fastest growing app ever. You need to serve millions of users, probably up into the hundreds of millions of users now. You need tons of GPUs there. And if everybody tries to Ghibli-ify the image, which we saw, they had a GPU crunch. Open AI practically doubled its users in a few weeks because of the Ghibli image redrawing that took place.

8:55By the way, not something I succeeded in. I somehow made a mistake with my prompt every time. But there's also this new shift in how the systems are designed, right? So back in the old days, so ancient history, 2024, the idea was that you did pre-training compute, which was that you trained this enormous model once and then you would have this model that you could run inference against, which was much less compute intensive. Give listeners a sense of just how much compute was required because these numbers are pretty astronomical to train a GPT-4 or a Claude. I mean, we were talking about history, right?

9:30We started with like two GPUs, which was the breakthrough of LXNet, and now we go up. I think the biggest system we know with numbers of GPUs is around 100 ,000, which is like Grok 3. It goes back to the deep sector. How do we measure this in terms of money? Let's say there are two ways of measuring it. There's like how much you need to pay for getting the CPUs if you actually want to buy them. And NVIDIA stock is going up. They have a nice margin. So again, if you buy 100 ,000 chips, Each chip costs you$40 ,000, right? So we go up into the billions to just get the AI chip, which makes up the large majority of the cost for building such a data center.

10:00This might change in the future if the AI chip landscape becomes more competitive right now, but there is a reason right now when NVIDIA is so high on the stock market. So it costs you billions to just get the GPUs. Of course, it's better a lot of times you don't need to actually build your own cluster. You can just go to a cloud computing company and rent the GPUs, right? So you tell them, hey, could I get 30 ,000 GPUs, say, for six months? If you do it this way, then current systems cost them like the three-digit millions for the computing power alone. The three-digit millions, wait, we talked about 100 ,000 chips at 40 ,000 each, which is$4 billion.

10:30So you're saying three-digit millions, in other words, as an amortized cost of... Indeed. The amortizing the chips over five years, the electricity they need to run them, and the very low salaries, no doubt, that the AI scientists are paid. DeepSeek puts out this number, it costed us$6 million. this does only include the armatized cost for buying the gpus it does not include the salaries it doesn't include all of your failures and experiments we at least know from one paper i think it was the science paper where hugging face played a big role where they said the total amount of compute spend was actually 3x the amount we spend on the final training ground you make mistakes do like de-risking training rounds i'm not going to spend 100 billion on new architecture which i haven't tested so like i start with a small architecture scale it up see if it works that's the whole idea right this trial and error and you throw a computer at it and someone you We say, sure, why not?

11:16You know, let's burn it and hopefully we get a good system out of it. One last question on the compute before we move into the sort of the heart of the discussion. So we talked about pre-training compute, which is these really expensive, risky, frankly, training runs that are building the biggest, most complicated pieces of software we've ever known. They run to seven, eight figures in cost and they can go wrong. But we then started to move with the O1 model that OpenAI released to the idea of inference compute or test time compute, which is that you allow the model to provide an answer and then explore many, many more answers before it gives you a final result, which allows a sort of a reasoning-like capability.

11:59So could you just explain the sort of relative weight of test time compute and what it has meant for the traditional prompt, right? How much more compute expensive is an O1 prompt over a traditional GPT-4 prompt. So everything we just said before is so true. We train these big pre-trained models, but now we do another thing. We make these model reasons. I think the best way to think about it is just like people are not surprised. If I get you an hour to think about something, you come up with a better solution than if I get your off-the-cuff answer, right? Yeah. And that's basically, well, this was always the case, right?

12:29Like a couple of years ago, somebody said, oh, if I prompt the model and tell it to think step-by-step, it produces a better result. All I told it to think step-by-step. Again, not surprising. Every teacher knows this. If you have a good prompt, your students produce better results. What we basically did with O1 and this whole new test time paradigm is we help these models to reason. We train them all to reason. So we have our big pre-trained model again, spending millions on this. Then we do another step on it. We can call it post-training. Some people might even call it mid-training now, where we do some type of reinforcement learning on reasoning.

12:59This particularly works well for coding and math. Why is that? There are little truths. We can verify it. There's a right and wrong. I work on policy. Kind of hard to say what's right and wrong here, actually. Well, I mean, your idea is probably all right. Well, yeah, I don't know. Let's see. But you're right. So maths and coding, we can verify them, and they lend themselves very well to this inference time compute. Let's get some numbers out there. I mean, we're talking about, I mean, roughly speaking, you might have a back and forth with a traditional LLM, and it might be a few thousand tokens where a token is roughly a word.

13:30It's a bit less than a word. In a test time compute model, that might run into tens, if not hundreds of thousands of tokens being computed. And every token is a computing operator. I mean, millions, tens of millions of competing operations. Indeed, it's just the same operation, the multiverse things. But ideally, at the end, you get these high-quality tokens, you know, like those who actually have the answer. I think my favorite example is actually the ARP-HEI benchmark, where like 01, when it came out, had this really impressive record. To achieve this record, I think they spent$20 ,000 on compute time on each task.

14:00They produced a couple of, I think, hundreds of thousands of tokens. They probably produced like five times a Harry Potter book to fill in a pixel. Again, this goes back to how LLMs do it. But I think this gives people an idea here. And they did two things. They did this reasoning. It's not that the model reasoned over five Harry Potter books. It just did many, many reasoning attempts. And you're just like, oh, I tried here, I tried here. And then you just look, which is the better one? Ideally, you have some research over it, right? So all of these things are now combined. But the cost can be kind of staggering, right?

14:24And I think the key implication is here to achieve the best capabilities at the beginning. You spend a ton of time on appearance. But again, AI moves forward. AI exponents move forward. Everything gets cheaper over time. Well, everything gets cheaper over time. And you have hinted at this because DeepSeek, as we know, came out with similar quality performance to many American models, but it was much cheaper to run. So what is the role of algorithmic efficiency, software optimization, and general optimization in all of this? Because I think we all know that getting a GPT-4 quality response is 200 or 300 times cheaper today than it was 18 months ago.

15:00I mean, it's staggering. We're talking about exponentials. That is a tremendous exponential. Why is that happening? History of humanity, you know? We do something and then we later learn how to do it cheaper. Just like economies of scale, we build bigger models. And I think that's the case for computer science since forever, right? And I think for AI, it's just staggering because we have these basically brute force approaches. We've got a big model, we do it, and then later we get smarter over it. We have these compute efficiency improvements. And again, the odds are really fast, right? It's roughly 3x per year.

15:28So like basically the cost to achieve a given capability is like 3x cheaper at the end of the year because we just like, again, this is really hard to measure. Do you do it with loss? Do you do it on a specific benchmark? It gets tricky, but the rough trend line is there. And then DeepSeq is, I think, caught many by surprise because they just didn't think about it. They just felt like, oh, it's always cost us 100 million. DeepSeq is perfectly on the trend line, basically, if you look at it, just like it sits directly where it's supposed to sit. The exception might be test time compute. This is actually a bigger increase in terms of compute efficiency in quantistar by improvement, which we've seen.

16:00Okay, so what we've done now is we've covered off the importance of compute, the difference between training and inference compute. We've also identified that we get much, much better, more efficient at getting the same performance over time. So that gets us to this question of the geopolitics of compute, right? So there is this competition between the US and China. We can discuss a bit later as to why there needs to be such a competition. But your argument was that the total compute capacity is what will ultimately give the U.S. an advantage. So is the core idea that it's not about the best frontier, most performant model that people are pursuing for the headlines, but rather about the rate of diffusion, which is an argument that Jeffrey Ding, who's I think a mutual connection of both of ours, has made, or an argument that Eric Schmidt made with a colleague in the New York Times today, that it's really about diffusion that's going to drive the economic and strategic benefits of AI?

16:59Yeah, I think I don't want to go forward and say it is this, right? I'm just saying it's like, look, there has been this hyper focus on benchmark capabilities on chatbot arena to measure the AI US rates. I was like, man, I think it's more complicated. Eventually, it really depends on what you want to measure and what you're worried about. If I care about national security, I measure different things than if I care about like economic diffusion, right? To some degree for national security, I might not even care about how many users are using a chatbot to get better cooking recipes. But I care a lot about if you have, I don't know, a million drones embedded with AI or like two drones embedded with AI.

17:35So the argument which I'm just putting forward is more just like, we need to understand how export controls, which the US government is doing on these AI chips, are actually working. What to expect and what not to expect. I think there's a word where like pure raw model capabilities matter a ton. It depends how you think the AI future is going to go, right? If there's going to be an Oracle in the future and it just tells me the truth to the universe, hell yeah. And that's like, what, 10 ,000 tokens? Amazing. I think that's quite not plausible. I think the future will just look way more complicated.

18:03It's actually about people using the systems. Now we're saying this, like having a car is something, but you actually need to drive it, right? And then it's not you driving, everybody needs to drive it, right? And that's how I expect it to help. And I think the best example to think about is AI agents and workers. People claim there will be in the future these drop-in remote workers, and AI is literally going to do my job because I'm just sitting in front of a computer all day. I mean, sometimes I do talks at whatever, I guess maybe that's my pleasure. These things, if that's true, it just depends on how much computer you have.

18:29I would say that we've run some experiments here. And a few weeks ago, while I was sitting with one of my colleagues, he accesses his AIs through an API rather than through the apps, which means he can count his token use. And he was going off just doing some little tasks. And in a 15 minute period, he used up 400 ,000 tokens, which is about 320 ,000 words, something like that, which is far beyond what you would do as a human talking to any machine, even with a brain interface. And I think as I see enterprise applications, as I go around the world and talk to companies, you increasingly see that they have lots of autonomous or semi-autonomous workflows running where AI systems are talking to AI systems at quite high rates.

19:11And those are well beyond the kind of consumption capability of a human. Or I generally have an AI listening to all of my meetings. And so you're running at tens of thousands of tokens a day before I do anything. So there is clearly this sort of core foundational demand that it's quite hard to get one's head around. And you're familiar with Open Router. I mean, Open Router is a, it's a sort of aggregator of LLM APIs, and it sees a very, very small portion of the market. But their year-to-year growth in token usage has been 38 to 40x, 40-fold, right, based on their public data. There is this burgeoning explosion that I'm sure is going to happen.

19:53And now that Google Gemini, for sake of argument, summarizes every email in Gmail, well, every one of those is across a billion users. Indeed. I mean, there's nothing new, right? Like, I'm talking to you by an iPhone. There's a lot of thinking which went into this iPhone. Lots of tokens, if you want to put it this way. And at the end of it, I get like one product. This is just how things are being produced and they're being like hidden behind layers. And this is what we're just seeing more and more with models. And again, every token that's being produced requires computing power. And then if you have less compute, you have less of this.

20:21If you have less compute, you've got less AI workers. And population size is pretty important for economic growth, right? Take us through the argument that you actually concluded, because you concluded with some things that are born out of data and the extrapolations of that data, ultimately that there would be more compute in the US and not a sort of sufficient amount in China to enable a 2030 economy. Is that a reasonable assessment? I'm not sure what sufficient means. I'm just saying like, if like one country has more, this definitely helps it accelerate faster, right? So the argument which I'm putting forward there is just like, look, guys, it's more complex.

20:53Like what do export contracts do and what they don't do? We have more compute. How can you leverage more compute? Well, actually, you have more AI workers, you can leverage your eye mirror. Again, if you go to the policy discussion, this does not necessarily mean you're winning. You need to use it the right way. What is the right way? That's kind of the question which I'm putting forward, right? Where do we expect it to help where we don't expect it to help? That's the question which I'm asking there. But if we go back to just like we think AI is going to be like the driver of the economy in the future, then having eight times more of the resource, which is the most important, it seems pretty important to me, right?

21:26This will just help you and you might get even the flywheel effect where you just take off to some degree. Have the Biden era export controls been helpful in this set of scenarios that you have posited? Absolutely. And notably, these are not Biden era export controls. Actually, Trump started it. Trump made sure that ASML, this obscure company in the Netherlands, which building the machines, which built the chips, don't go to China anymore. This was in 2018. And then we had entity listing controls on Huawei, one of the leading AI chip producers in China, right, and the related companies there. And then buying just added on top of this with making sure the AI chips don't go there.

22:02And I think this is a large part why we currently see less compute in China versus in the US. I will posit a different reason on that while I acknowledge all of the points you've made. You know, the compute has to be secured by people who know what they are doing. It's not just a case of having billions of dollars. And the US and China both have a small number of firms that do know that. But ultimately, it's expensive and you have to be able to access the capital markets. And the balance sheets of Google and Meta and Amazon are far, far larger and deeper than those of Baidu and Tencent and so on.

Read the full transcript

22:40And so you already had an ability to do more CapEx, buy more chips. and in the context of the moment, we had that moment just before the Ant Financial IPO, which was pulled, where the Chinese government really sat on its entrepreneurs and a chill came into the Chinese market. Their valuations declined, the access to capital declined, the venture capital dried up, the balance sheets weakened. And in a way, I'm sure someone, a historian will look at this, but my argument would be that more constraints were applied at that point. Because even today, with the U.S. hyperscalers and the big tech firms spending$320 billion on CapEx, a little bit more than half of which will go on GPUs, that's above the capabilities of the Chinese big tech firms anyway.

23:27I'm sure that it's been very useful, particularly the ASML thing, which actually, in a sense, prevents China or slows China's ability to build its own semiconductor industry, which has been a strategic priority for many years. but in terms of procuring chips there was just much more money in the u.s absolutely yeah i think i mean to some degree this also might largely explain where we have like more frontier companies here right so i do i would like i would not argue against it that it's not one of the reasons but even if you have the money and you can just buy it like less good chips that's definitely going to hit you in the long run right and again depending on where this goes i think the point about extra controls this has been that confusing thing with deep seek it takes some time to buy it and we had many failures right the first times they got the numbers wrong Those are the chips they train DeepSeek on, right?

24:12So we basically have only working X controls since, what is it, October 2023? It's a little bit more than a year, right? You don't build a cluster every new year. So I expect we'll see like a major impact going forward if certain enforcement and things are being patched. And again, all of this also depends on will NVIDIA's chips continue to get exponentially better, right? To some degree, the good thing about computing is it gets exponentially better. Who cares about the computer built five years ago? Nobody uses that anymore, right? It's like my chips now are like 10 times better and I produce 10 times more of them.

24:42But if this would level off, then this might look actually different. You do hear, I mean, you hear Jensen say, I'm the chief revenue destroyer in the sense that he keeps bringing out new chips that are better than the old version. But you see the hyperscalers saying, we're sweating chips at six years old and seven years old, and they're still sort of fully used. In a sense, they have to say that to the market. What we have here, though, is we have a challenge. We have a real bifurcation here, which is national security needs, which in a way predicated on how good the frontier models can be. And there are the economic needs, which are predicated on the ability to deploy and diffuse high-quality, reliable AI into your firms.

25:20That feels like it's a really difficult pair of objectives to optimize for. You know, there's a line, you know, he who chases after two rabbits will get neither. How do policymakers make sense of these things? Because they're quite conflicting. They are. I mean, this is the concentrate of, right? Security comes at a cost. But this is also known for every company. And every company has ideally a CISO who takes care that your IP is secure, right? But this isn't so much about security coming at a cost, right? Security can be having the equivalent of biological secure levels for your labs, for AI labs, and that is a cost that comes out.

25:59This is something that's slightly unknowable, which is will scaling over five years create this super intelligence that builds self-reinforcing loops that take you up and away. And if it does, it's an existential national security issue. And therefore, you need to think in those terms. The other part of the argument, which is the economic argument, is that it's broad diffusion and use that matters. And it feels like those are really slightly at odds with each other. Yeah, I mean, look, there is like a view you just laid out, like it's a national security, which is like, well, the build AI, it takes off and then it's going to kill us all.

26:31I I think that's a national security risk. I think there are many other national security risks which are between AI is all good and between AI is going to kill us all. There are many different risks, right? We talk about cyber risk. We talk about biobiological misuse. We talk about loss of control scenarios, right? And they all need to be managed. And I think they're actually not too different from like normal security risk, which we're trying to manage all day long. If the US is trying to decide who to sell fighter jets to, pretty similar on this one, right? And then you're totally right. Eventually, if you decide to not sell chips to China, you might lose revenue right that's just the case i think this is just like again what the government began it has been bipartisan consensus why has decided it was like yep we take this revenue in but eventually in the long run that's the thing we want to do here because ai might go to kill us or china might take over the world whatever your favorite threat scenario is i think there's like a polarity of like national security arguments why people are doing this and again yeah my job as a policymaker is like to calculate the costs of both right all i would say is our industry has been doing pretty fine so far, right?

27:31And again, when we talk about export roads to China, that's one market, there's still the rest of the world. The majority of the market is still available. And if you just compare these semiconductor exports, they're minor compared to what's happening right now with the tariffs, right? This is the way bigger economic deal. And it's still being done for many reasons, right? I want to sort of dig into this a little bit because, you know, in a way, there is a cost to export controls. There is a cost to controlling this. The cost comes as follows. it's about where the government chooses to intervene in a market, which has a doctrinal effect on where a nation operates and thinks about itself.

28:05It has a fundamental impact on the certainty and the risk profile of a particular industry and therefore its ability to secure capital. It has a cost of pallor over any country that becomes dependent on, say, for NVIDIA chips where they see that there is a line that is easily movable at whim in pencil about whether they'll be able to get updates or support and whether that investment in CUDA and that infrastructure goes. So it's more than a blanket, sort of simple, clean clinical cut. It's also one that spills over in many different directions. Then on the other side, we have this issue of to what extent does it stimulate innovation in other places, in China in this case?

28:53We haven't really talked about Huawei much, but Huawei came out with this cloud matrix, server unit. It has a sort of performance that's similar to NVIDIA's top-tier system, except it's four times as energy expensive. There are these other calculuses, right, that say, what are the costs that we have to bear? How do you think about that? Let's go back to the Huawei system, I think, because it's a nice example of exactly the point which I'm trying to make. They will be able to build systems which are, on paper, similar to the American system. The difference, just like, as you're saying, it eats way much power.

29:25It requires three times more chips, right? And, like, you have all of these inefficiencies. Does this mean they cannot not train, let's say, a GPT 4.5 model? Absolutely, they can't probably. This is kind of what I'm expecting, right? The difference is just, like, while TSMC is churning out five million chips per year, they're churning out less than a million chips and they need three times more chips to achieve the same performance, right? This is exactly the argument which I'm making here. And then you're right. There are costs to this, right? These export controls, they used to say small yard, high fence.

29:54I don't think that's the case, right? To some degree, when it's originally started in 2022, this was before Chattivity. And I give the government like a lot of credit for foresight here. Like knowing that computer's cool, if like people would have invested in 2022 into NVIDIA, they were still early. So they had this foresight. And then it was just like, oh, we only hit on these advanced AI chips which are being used for servers. Sure, AI was a niche application. But again, if you look at people like us, we believe AI is the future. AI is going to be everywhere. Then this is not a small yard anymore, right?

30:23If I use AI to literally plan my day, cook my food, and everything else gets organized around me, then you have collateral damage, which hits on the broad economy. And this is just a thing which is simply true, right? And this is generally the challenge of governing dual-use technologies on these kinds of things. What I would say, though, is China still has access to AI services. that can still use Microsoft Azure, that can use ChatGPT, right? So it really depends where on the value chain you want to intervene. What the S-government wants to do is like, hey, you shouldn't be able to build your own competitive big systems and like you make human rights abuse and a bunch of other nasty stuff with it.

30:59You want to use AI, which is like monitored by us. And like, again, you cannot build biological weapons. But sure, go for it. That's fine, right? And I think that's the best way to think about it. And here, again, the key is, chips can't go into China. they can access cloud computing right now without any restrictions. And any engineer, it's not like I need a computer in my basement to train a model, I go to a cloud provider. So right now they can't technically do it. You're totally right. They're the whims of the US government if they decide tomorrow to cut them off the cloud. That's a problem.

31:29But again, that's historically been the case for the monetary system the same. This is just how national security works and how you try to govern, right? And those things are trying to be balanced. But I think eventually we can find a good middle ground here. And ideally, it's not only the US doing all of this. It's the U.S. and its allies and partners with the semiconductor industry, the key allies there, building a broader multilateral governance framework, which is building with democratic values. That's at least what I'm aspiring to. Let's talk about technology change in all of this, right?

31:56We're having to hold a couple of distinct ideas in our head. One is this idea of the very powerful national security model. I'm not thinking about one that kills us, but I'm thinking about one that runs the million drones rather than the 20 drones. The other is the economic side of this, which is the ability to scale this out across an entire economy and all the workers and give them all superpowers. One of the things that strikes me over a five-year period of time is that isn't long enough for certain things to change. I mean, the CTO of AMD, Mark Papermaster, has said that by 2030, he expects most inference loads to be run on device, right?

32:30So on the smartphone, because in reality, to undertake 28, 20, 29, 20, 30 tasks, the quality of a smartphone chip and the RAM that's on that device will be enough to do that, to act as a software coding agent and so on and so forth. So to some extent, at that point, the battle is not about these super sexy$40 ,000 GPUs or their competitors. It's about what can exist on device for which China can build lots of devices with great ships, that's the mobile phone industry. But it might also be about new architectures that emerge predominantly for inference, right? Because sometimes we think the training and inference are the same, but they're not.

33:10They require different things of the chip architectures, the memory bandwidth, the latency, and so on. So that also changes the way in which the leakiness or the effectiveness of export controls, given what's needed for the inference and the deployment. Absolutely. Well, I mean, we just talked about history. We talked about going from two chips to 100 ,000 chips. The two chips was not a transformer architecture. At some point, we developed transformers. At some point, we had a new loss function. Me as a technologist, I love when people develop new papers, but I'm interested in the macro trends.

33:41That's why I just love like line going up, compute, Moore's law, more transistors per area. Generally too. How do we do it? Completely different architectures over time, right? If you find the right friction layer, again, you see these exponentials being the case there. Then some people are arguing like, oh, we have a new architecture that does this. And they're like, true. That's what happens all the time. This is what we call increasing algorithmic efficiency and compute efficiency. I think it's quite unlikely that somebody tomorrow pulls out a new architecture and says, oh, look, AGI on a smartphone.

34:10I don't think that's the case. I think we would first build AGI on a big cluster, and then a couple of decades later, we would build it on a smartphone. This is just what we always see for these kinds of things. So the thing to acknowledge here is diffusion, right? A given capability gets cheaper over time. I used to run it on a big server. No, I can run it on my smartphone. That's definitely the case, right? So yes, we will see more workloads being on edge. For privacy reasons, We see Apple being bad on these kinds of things. But even Apple is kind of acknowledging. It's like, damn, there are certain questions you can't do them locally.

34:39You're going to run out of battery power. We sent them up to the cloud and they came up with a really fancy architecture to preserve privacy. That was super fancy. Right? Yeah. And these are the things I'm expecting, right? So I'll usually say, if I want to set a reminder, sure, my LLM on my smartphone can. If I ask for the meaning of life, well, hopefully it goes up to the server and it's processing and generates five books and tells me what the meaning of life is. Right? So I expect this trend to generally continue, that we just see these kinds of things. While, again, one curve is going up, the best capabilities will always be in the bigger servers, roughly, whereas a given capability goes down over time and will diffuse across.

35:14And that's definitely a national security challenge. You're totally right. If GBP5 poses a national security risk, we can control it for, what, a year or two? And then every guy in a garage with five Mac minis potentially could reproduce it. And that's just the thing we need to deal with eventually. Well, and eventually may, in that case, be very soon if we're talking about GPT-5. Okay, we've got a few minutes left. I want to jump into just some thinking about the future and what the future ought to look like. When DeepSeek released its R1 model in V3, we're not much fanfare from DeepSeek, but lots of people were very excited at the time.

35:47One unique feature was that it was an open source and open weights model. And Marc Andreessen, who is a Silicon Valley venture capitalist, really behind the Build America movement, viewed this as a gift. I think I forget the exact text of the tweet, but it was something like, you know, this is a gift to humanity or something similar. And what that does is that helps to diffuse AI capabilities because anyone can get those weights. They can do things with them, customize them, and build their own applications. And at this point, you know, Google and OpenAI have been getting progressively more closed.

36:19So there is this question about, as we look at the future, between a world where compute resources are highly concentrated, concentration lends, of course, to control, but in many places, typically in political systems, we don't like concentration of power, right? We like decentralized power. We like checks and balances. We like some competition. We also like that in economic systems, right, in markets. There is this simple two by two, right? Maybe it's just two boxes, two parts. a world of compute concentration and a world of compute abundance and a world of closed source models living with that concentrated compute and a world of capable open source models living on more available compute.

37:02What are the considerations for each of those scenarios? You know, how should people making their own mind up on what will ultimately be a political choice for us and hopefully think about those two? Well, that's indeed the question to which it will be a political choice for us, right? Or if one needs to intervene. I think the best way to start with is just let's look at the status quo. Most compute is largely concentrated. Where is it concentrated? The large majority sits in industry. This didn't always used to be the case, right? Academia was really strong on traditional high-performance computing.

37:29This lead was eradicated over the last decade, basically. Then secondly, where is this compute concentrated? Within industry, within our hyperscalers. Amazon Web Service, Microsoft Azure, Google Cloud, Oracle, they own a large majority of it. If you look at NVIDIA's shareholder output, they need to tell if it's mostly for big customers. That's what you see. That's simply the case. These are all American companies. Which countries are the most concentrated? Most in the U.S., bad news. We just published a data set where we see, I think, 60 % to 70 % of all supercomputers are in the U.S., right? Like Europe, 10%.

38:04China's another 10%. Something along these lines. I think roughly this is also what I expect. All of AI computer results being there. So that's the status quo, right? So if people just say, like, power to the people, they're worried about more concentration. I was like, look, the status quo, it's already concentrated across all these companies. And there's a good reason for investigation. I think there is ongoing investigations into the cloud stack, right? All of these cloud companies have just been getting bigger and bigger and bigger. Again, I'm not an antitrust lawyer, but I think this is at least something one should look into what we could do about it.

38:32So those are the computer resources being concentrated. Argument against it is like, this doesn't mean who's using it. We both right now use computing power by Substack. Well, I don't think Substack has a cluster. It's probably Amazon Web Services, right? So it's something that I can just access it. The world is just like, what if they cut us off, right? What in the future, I need my AI lawyer, and they turn my AI lawyer off, and I'm just like on my own. That's a big problem, right? Right. And I think that's where we just need to make sure, like, which role will AI play in the future? And then I think it's even more important.

38:58I don't want compute in the future. I want AI. If I have 10 GPUs in my garage, but I don't have the leading AI model, that's the way bigger problem, right? So I think that's the thing which we need to think about, like what kind of services do people need access to, right? And then is it a compute with the AI model or not? If my mom needs a lawyer, I'm sorry, she's not going to set up a cluster and spin up the new DeepSig R2. Ain't going to happen, right? There, I think it's like way more likely you potentially want access to AI systems. Even here right now, whoever can afford the more tokens already has a lead on this, right?

39:30There are more countries, of course, than the US and China. So this also explains in a sense why the EU and why India and also the UAE are pushing for their own compute infrastructure. And that can either be physically owned by local companies that fall under their jurisdiction, or it can be the right kind of contracts with a hyperscaler. And then you've got this emerging tension that we've seen, and certainly Microsoft has come and spoken about, talked about it in the last couple of weeks, regarding Europe and the sense that their assets in other countries, sovereign nations now want to know under whose legal purview does this reside?

40:09How certain can we be of this? So this sense of this concentration risk is really being noticed increasingly over the last five years. I guess one question is, how do you imagine it plays out? You're totally right. You see more and more countries kind of waking up to this computer, right? Like I just said in the beginning, computers survived in San Francisco. I was just in Brussels. It's all survived there. They just announced an AI factory, right? I was sort of hanging around London around the time they announced a future of compute review in the UK. So like, I think people woke up to this compute idea earlier.

40:40And to some extent, it's easier. I think as a government, it's easier for me to set up a cluster than to actually build a leading AI company. This is my worry. I'm just like, guys, I don't want you to overly focus on compute, where there's a kind of force leading people to compute. It's like, wait, are you actually solving any real problem here? You need a whole tech stack strategy. Do you want to build a different AI company? Sure, go for it. Do you need more energyless compute? There's a problem here that I think is really hard for these governments to wind their way through, which is that the AI layer is going to be the fundamental infrastructure layer of the global economy.

41:12AI systems will be our interfaces as citizens, consumers, whatever role we take, to the bulk of the digital services that we need, and increasingly so. And so if the iPhone ended up being a political issue, which in a sense it was, it's something I wrote about in my first book a few years ago, the AI layer will be even more so. And I think that governments will have to figure out how they can articulate what they need and expect and what kind of guarantees they need from that layer, especially as you say, it's concentrated in American hyperscalers delivered admittedly through their European subsidiaries, but ultimately all routes lead back to Washington in some sense.

41:55And that seems like a really difficult, thorny problem for them to address over the next few years. Absolutely. I mean, you were saying it's actually not a new one, right? When Snowden came out, I was like, oh, god damn, all of our data is going over there. We want more sovereign cloud. So to some degree, I think it's not a new problem. It's just repackaged with a new tech layer on top of it. The thing where I think it just differs, I do think AI is a fairly big deal, probably the biggest deal in our century. So we really need to get it right here. So I'm excited about all of these governments having these plans, right?

42:25And I want more techies in there and doing it. But I'm just saying I have not seen a sovereign AI strategy, if you want to call it this way, which actually makes me feel like you're actually tackling the root cause here right and eventually what i believe here is i don't think the plan is like every country is going to like redevelop the whole tech stack and come up with down solution i think eventually what i want again maybe i'm a little bit too much kumbaya it would be nice if we can just get together and just say like here the standards here's how the democratic world is doing ai and be like have these scales right to work on this and like you're not at the whims of somebody just shutting off your ai tomorrow because they just feel like it i don't think that's great i don't think that's how we should go about it That is super kumbaya, so I'm going to have to pick that up.

43:05Number one, what did you have for lunch? I want some of that to feel that great. But number two, I mean, the issue that you have there is this sort of original sin of the internet, which is that we never really figured out who was going to ultimately make the rules. And it was fine in the late 80s and early 90s when it could be, you know, IETF and IANA and so on. But now it's a much, much bigger issue. And the decision ultimately needs to rest from the view of a national government or the EU, which means that for that to be cast iron, those rights have to be subsidiarized by the U.S. government in a way that is non-repudiable to whoever else is using that.

43:41And I actually don't know what that mechanism is anyway. And even if that mechanism existed, I'm not sure I can see the political circumstances in the next few years where it's politically tractable. So let's start with that first question, which is what is that mechanism? I mean, we made progress on these kinds of things, right? There are certain data localization laws. And this at least stops other governments to literally raid the data center, grab the hard drive, and take your data. There's also encryption. So we've got methods to something to defend. And there's nothing new. We've been doing this with different types of cloud acts, initiatives, and more.

44:16So that's already going on. I think for AI, we want the same. And ideally, you just want to be part of the broad AI supply chain. I think to some degree, we're having these discussions because Europe is just like, where is Europe on the tech stack? I see it all the way below with ASML. right where they're producing the machines which produce the chips but across else it takes it looks like not that well i don't think the u.s government can say tomorrow like oh actually taiwan netherlands south korea meh we don't want anywhere then they got a problem right because these are key allies in the semiconductor supply chain so i'd be more way more keen for countries to pick like their part in the supply chain right this is just how it works be it a local community be it on a global scale to pick it there and then localize certain things where they feel like They've got a say or they've got potentially some bargaining there if they eventually needed to.

45:01And then again, I think just building data centers gives you at least the physical control over its computing power. And then we should just ask the question, which part of the tech stack do you want to localize? What is the most important? I don't have the answers here. I'll make two comments about that, right? So the first is I like what you've said about pick a point where you can be the winner and you can be the best. And therefore, everyone needs you in the team. it's a little bit like how the primes in the US, you know, build even a hammer in 39 different states. So that there's always a congressman fighting for your size.

45:29And, you know, the wood is done here and the plastic is done there and the screws are done there. You could, you know, argue that. But of course, as you say, three countries right now are part of that, which is the Netherlands, South Korea for memory and Taiwan for the chips. So that is quite an interesting approach. And you're right, I haven't seen that in any sovereign AI strategy. And you've read more than me. But on the second point about the physical data centers, I think one of the problems there is Moore's law or Huang's law, whatever you want to call it. Because essentially, given the price performance improvement of chips every year on a static amount of CapEx every year, you're sort of adding the same amount of compute as you already had.

46:08And so even if you have physical data centers and you can't swap out the chips because the chips are sort of physical, they have different cooling requirements and energy requirements, you have to build new data centers for the new families of chips. Even if you do have the data centers, in two years, if you get cut off for whatever reason, in two years, they're going to be anemic and they're not going to be able to run anything and they're going to be irrelevant in that new modern AI economy. So you have to be able to guarantee the consistent supply of uprating of the chips. And it's not just going to be hyperscale data centers, these big ones like Stargate and Colossus.

46:43It will be MetroScale and Edge data centers as well that will need kind of constant replenishment. So it feels like that may be an indicator of your commitment to be a customer and therefore create incentives to not be cut off. But it doesn't feel like it's as strong a defense as the proposal that you have, which is let's become the best photoresist in the world, which I think exists in Japan right now. I mean, two points. Let's flip it around. You just said, well, you build a data center and the next year is outdated. Nice thing is, you can enter the race pretty late. You can get a good share of compute.

47:14If I now build a data center with B200 and Vitea's newest chip over H100, here we go. It's way easier. And if you look at these compute shares, and again, me crunching all these numbers, You just see how, like, to some of the volatile they are, right? If you just build a couple of big data centers, you're suddenly in lead. You're right. It's not guaranteed. Next year, it changes again. But it goes both ways here, right? And then you're right. You need to always upgrade and get the newest chips. This is where the semiconductor supply chain, again, plays a key role, right? And again, the U.S. government just identified AI chips.

47:44Oh, nice. They're like these four countries which play, like, this important role here. We got, like, this nice authority. We can just decide where it goes. But again, why did we have all of this outcry, for example, like Poland and the diffusion framework not being one? same with Portugal, it's exactly for this reason. The vibes are off. The vibes are off if you just feel like you're subject to and they can cut you off at any single point in time. And it's a product you literally need to replace every five years, right? And this is where I wish for more bilateral and multilateral to have these engagements, right?

48:11But I think we need to talk a lot about AI here. I think we found this like across many different industries where we found a way and ideally we can work together. That is a positive note, Lenart, for us to bring this conversation to an end. And I'm so disappointed to do it because I'm really enjoying it. So I hope we will get a chance to do this in person or in the near future. Just have one fun quickfire question for you. Don't overthink it. Who makes a more fun dinner companion? An AI governance person or an AI developer? I mean, speaking for me personally, I love the developers. You know, I'm in DC.

48:43We got lots of high level talk with little substance. And you know what I love about engineering? There are truths. Things either work or they don't. And again, this is my bias as an engineer. So right now, maybe I'm just speaking what I need right now, while Gath is spending a lot of weeks here in DC and not being in the Bay Area for a long time. I would just love to hang with an AI developer and ask him a bunch of questions. Well, you know what? The other thing that has worked has been our conversation this afternoon. Thank you so much for making the time. You have a great weekend and thank you to everybody who tuned in.

From the publisher

Lennart Heim, a researcher and information scientist at RAND Corporation, joins Azeem Azhar to unpack a provocative claim: China is catching up with US AI capabilities, but it doesn't matter. 

Timestamps: 

(00:00) Episode trailer 

(01:19) Lennart’s core thesis 

(03:26)   Why compute matters so much 

(07:31)  The investment split between model R&D and model execution 

(11:18)  How test-time compute impacts costs 

(16:14) The geopolitics of compute 

(21:32) Why does the U.S have more compute capacity than China? 

(25:01)  The trade-off between economic needs and national-security needs 

(31:54)  How technology change might shift the battlegrounds 

(35:33)  Dealing with compute and power concentration 

(48:19)  Concluding quick-fire question 

 

Lennart's links: 

Azeem's links:

This was originally recorded for "Friday with Azeem Azhar", a new show that takes place every Friday at 9am PT and 12pm ET. You can tune in through Exponential View on Substack. 

Produced by supermix.io and EPIIPLUS1 Ltd


Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

More from Azeem Azhar's Exponential View

All 44 episodes
China’s catching up to US AI… Here’s why it won’t matterAzeem Azhar's Exponential View · 49 min
Listen in VO