In short
Podcast Summary: Nvidia CTO Michael Kagan - Scaling Beyond Moore's Law to Million-GPU Clusters
Podcast Details
- Title: Training Data
- Episode Title: Nvidia CTO Michael Kagan: Scaling Beyond Moore's Law to Million-GPU Clusters
- Host: Sonya Huang and Pat Grady
- Description: In this episode, Michael Kagan, co-founder of Mellanox and CTO of Nvidia, discusses the transformative impact of Mellanox's acquisition on Nvidia, technical challenges of scaling GPU capabilities, and the future of AI infrastructure.
Key Themes and Concepts
- Cultural Philosophy of Nvidia
- Nvidia operates on a "Win-win" philosophy, focusing on expanding the market rather than competing against others.
- The success of Nvidia is closely tied to the success of its customers.
- The Mellanox Acquisition
- The $7 billion acquisition of Mellanox was pivotal in transforming Nvidia from a chip maker to a leader in AI infrastructure.
- Mellanox’s technology enables Nvidia to overcome limitations posed by Moore's Law, allowing scaling beyond traditional silicon capabilities.
- Scaling Challenges
- Transitioning from single GPUs to large clusters presents significant technical challenges.
- Network performance is crucial; it must not only provide high throughput but also low latency to maximize efficiency.
- Kagan emphasizes that the architecture must be designed with the understanding that failures will occur in a large system.
- GPU Clusters and Network Performance
- Scaling up (increasing performance within a compute node) and scaling out (connecting multiple nodes) are essential strategies.
- Mellanox technology enables efficient communication between GPUs, which is vital for parallel processing and scaling.
- Training vs. Inference
- Training involves backpropagation and weight adjustment, while inference is about generating predictions.
- Generative AI has shifted the demand, creating more load during inference as models engage in iterative processes.
- Future of AI and Data Processing
- Kagan describes AI as potentially being humanity's "spaceship of the mind," capable of enhancing human cognitive abilities.
- AI could lead to new discoveries in physics and other fields by processing data in ways humans might not conceive.
- Energy and Infrastructure Implications
- Discussion of energy consumption limits for data centers, with implications for scalability as energy becomes a bottleneck.
- Liquid cooling technologies are being adopted to enable denser computing environments.
- Predictions on AI Development
- Kagan believes that AI will drastically change how humans engage with technology, allowing for significant productivity increases.
- The partnership with Intel is seen as a fusion of accelerated computing with general-purpose computing, enhancing overall efficiency.
Key Takeaways
- Mellanox's Role: The acquisition has been crucial for Nvidia's transition and growth in the AI sector.
- Networking is Key: For AI workloads, network performance determines overall efficiency, not just raw computing power.
- Future Workloads: As AI evolves, the distinction between training and inference will blur, with both requiring significant computational resources.
- AI’s Potential: Kagan envisions AI as a transformative tool that could enable humanity to solve complex problems more efficiently.
Conclusion The episode highlights the critical role of networking and infrastructure in advancing AI capabilities and the ambition of Nvidia to lead in this transformative era. Michael Kagan offers insights into the technical and philosophical underpinnings of Nvidia's strategy, emphasizing collaboration and innovation as key drivers of success in the AI landscape.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00One of the interesting things about NVIDIA is the culture of Win-win. We are not after taking a bigger piece of existing pie. We are after baking a bigger pie for everybody. And our success is our customer's success. Our success is not the failure of our competition. And I think fusing together conventional computing when human machines and accelerated computing that provided with NVIDIA, it actually gives NVIDIA and Intel channels to the market, or expanding the market and serving the markets that otherwise was more challenging.
1:04We are delighted to hear today from one of the legends of the semiconductor industry, Michael Coggan, the CTO of NVIDIA. Michael was formerly chief architect at Intel and then co-founder and CTO of Mellanox, which NVIDIA acquired for$7 million in March 2019. In the time since, Michael has been a major driver of NVIDIA's dominance as the AI compute platform, in large part due to the role of Mellanox and Interconnect in driving chips beyond Moore's law. The AI race is ultimately a silicon race to squeeze the most intelligence possible out of each unit of silicon. And Michael takes us down a journey about how the compute frontier has evolved from squeezing more transistors onto a single chip to bringing together thousands and hundreds of thousands of chips into a single fabric connected by networking or an AI data center.
1:52Michael has been driving the compute frontier forward for more than four decades, and we're honored to have him on today's show. Okay, we're here with Michael Kagan, the CTO of NVIDIA, currently the world's most valuable company. Michael, thank you for joining us. Thank you. My pleasure. So I thought where we could start. Our partner, Sean, likes to make the case about every six months that NVIDIA would not be NVIDIA without Mellanox. Mellanox is the company that you co-founded some 25 years ago and have been a part of through this day. So can you kind of paint that picture for us why is it that the Mellanox acquisition was so critical to NVIDIA um you know there is a huge transition in the world in terms of computing and need for computing and it grows grows exponentially and one of the things that we usually estimate linearly but the world was exponential and exponential growth now is actually accelerated it used to be like which is a basic silicon and it was you know twice every other year and regardless you know the discussion that more low in terms of physics is not quite running anymore the once the ai kicked in which was in the 2010 2011 and kicked in when gpu from graphic processing unit became general processing unit actually where it's running workloads were the first time the ai workload was run on the gpu taking advantage of the programmability and parallel nature of this machine the requirements for performance started to grow up at much higher coefficient so the models started to go up in terms of size and capacity 2x every three months, which requires now 10x or 16x a year performance growth versus old school of twice every other year.
4:07And in order to grow this scale, you need to innovate and you need to develop solutions at much higher scale than just basic component. And that's where network kicks in. That's where network is. And there's multiple layers of scaling performance that requires high-speed networks and high-performance networks. The one is what we call scale-up. Basically, if you're going back to the CPU days, scaling up was more low, more transistors, and also some advances in the micro architecture like auto-force execution and at some point multi-core and so on and so forth. But so this is the basic building block of computing.
4:54In the GPU world, the basic building block is the GPU. And in order to scale it up more than you can do on a single piece of silicon with the lot of advances that we are doing with the micro architecture and advanced technologies, you actually need to do something on the scale of, on the sort of multi-core CPU, but a much larger scale. And that's what we are doing with Envilink. This is a scale-up solution. So our GPU, what we call GPU today, is a rack-sized machine. You need a forklift to lift it. So, you know, if you order just GPU on Amazon, just don't wonder that you'll show up this huge rack.
5:39Yeah, people think chip, but it's really a system. Right, and that's just one GPU. Yeah. Okay, so basic building block, very basic computer that application software is running on is this GPU. And it is not just silicon, it's not just hardware, it's not just wires, but there is also a software layer that exposes CUDA as the API. And that's what actually enables to pretty much seamlessly scale. I'm simplifying a little bit the story, but seamlessly scale from one component that used to be single GPU all the way up to 72, maintaining the same software interface. And once you get this building block as big as it conceivably can be built in terms of you know, power, cost, efficiency, then you start scaling out.
6:39And scale out means you take many of these building blocks, connect them together, and now on the algorithm level, on the application level, you actually split your application to multiple pieces running in parallel on these big machines. And that's again where network comes in. and there is a, so if you talk about scale up, we basically made memory-like domain to go beyond single compute node and a single GPU and that's actually the first thing that Mellanox technology comes in because before Mellanox acquisition, the scaling up of NVIDIA with the NBLINK was limited to a single node machine. Going outside of single compute node, if there's 72 GPUs, it's actually 36 computers.
7:34Each one has two GPUs and they're wired together, so present all this as a single GPU. Getting connectivity outside of the single node, it's not just plugging in the wire to the connector. It's a lot of software, it's a lot of technology within the network how to make multiple nodes to work as the single machine. And that's where Mellanox first, just immediate in terms of, you know, the way we go upstream, that's the first one. The second one is how do you split the operation across multiple machines. And the way to do it, if I have a task that it takes one GPU to do one second, if I want to accelerate it, I split it to 10 to, you know, thousand pieces and send each piece to different GPU.
8:27And now in one millisecond, I get done whatever I was doing in the second. But, you know, you need to communicate this partial job split, split the task. Then you need to consolidate the results. And every time you are, and you run this, you know, multiple times, Sometimes you have multiple iterations or multiple applications running to us so that there is a part of doing communication, part of doing computation. Now the thing is that you want to split it to as many pieces as you possibly can because that's your speed up factor. But then if your communication is actually blocking you, you waste time, you waste energy, you waste everything.
9:13So what you need to do, you need to have a very fast communication. so you split it to the many many pieces and so each piece takes very little time but then there is a another piece that is communicated and you need to feed it this time so that's that's just pure bandwidth and another thing is that when you tune your application you tune your application so that communication can be hidden behind computation and it means that if communication for some reason gets longer, then everybody waits. So it means that what you need to do in the network, you need to have not only just raw performance like what's called hero numbers.
9:55You know, I can get to that many gigabits per second. I also need to make sure that no matter who communicates to whom, the latency, the time it takes, distribution is very narrow. So if you look at other network technologies or other network products, you know, You go to the hero numbers, sending bits from one place to another. It's basically physics. It's pretty much close to everyone. We are a little bit better, but that's not the big advantage. But when you do it thousands of times, and it takes the same time to do it versus a very wide distribution of other technologies, then the machine becomes less efficient.
10:38So instead of being able to split your job to 1000 GPUs, you can split it only to 10 GPUs because you need to accommodate for the jitter on the network within the communication, within the computation phase. So inherently, network determines the performance of this cluster. And we look at this data center as basically a single unit of computing. Yeah. Okay, single-limit computing means that you look at this, you start architecting your components, your software and your hardware at the point where, you know, this is a data center, this is 100 ,000 GPUs that we want to make them work together. we need to make multiple chips, compute chips two, network chips five.
11:29Okay so this is a scale just you know in terms of you know what what's the impact just and what's the investment you need to make to create this single unit of computing. So that's that's where Mellanox technology came in and And another aspect of this is there is a, we talked about network that connects the GPUs to run the task. But there is a other side of this machine, which is customer facing. So you need to, this machine needs to save, to serve multiple tenants. And this machine needs to run operating system, you know, every computer runs operating system. Another part of the Mellanox technologies is what we call Bluefield DPU, data processing unit, which is actually the computing platform to run the operating system of the data center.
12:27In conventional computer, we have a CPU that runs operating system and runs application software. And there is many things we can talk about, you know, the advantage versus disadvantage. But there are two key things. One is how much time do you spend on your general purpose computing to run the application? You want to maximize it. Another thing is how do you isolate your infrastructure computing from the application computing? Because viruses and cyber attacks and so on and so forth. And being able to run infrastructure computing on the different computing platform actually reduces significantly the attack front, especially in the side channel attacks, versus what happens if you run it on the same computer.
13:23If you remember, there was five or six, well, actually, almost 10 years ago, there was this meltdown and all these cyber attacks on the side channel on CPUs. And this cannot happen or the attack surface is reduced significantly when you run the different. So on the other side of the network, we have also technology. So that's what makes data center to be more efficient. and I, well, I may be not objective, but I do agree that this merger of Mellanox and NVIDIA, and it actually goes both ways. I don't think that networking business of now it's NVIDIA, previously Mellanox could have been growing that significantly.
14:09Yeah. As it grew now, I think we are the fastest growing internet business, you know, let alone NVLink and InfiniBand. Yeah. But just internet business is the fastest growing business ever. What are the things that break as you get to 100 ,000, maybe eventually a million GPU clusters and how do you use software to help design around that? It's a multi-stage challenge. challenge. Okay, one of the things that you need to keep in mind, it's not very obvious for all the engineers that when you design the machine, you think how to operate it. Well, you know, you have these components and they're working and now just let's figure out.
14:56Okay, so the thing is that the hardware component works at 99.999, whatever percent of the time. And it's usually okay if you are dealing with single box with a couple of them but if you are building 100 ,000 component machine, a 100 ,000 GPU machine which means in terms of components there is a millions of them, the chance that everything works is zero. So something is definitely broken and you need to design it both from hardware and from software perspective to keep going, to keep going as efficient as you can to keep, you know, your performance, you keep your power efficiency, and of course, keep the service running.
15:41So, this is challenge number one, even before you go to millions. This challenge actually starts at, you know, a few tens of thousands. That's number one. Number two is when you are running these workloads, it is really important to, sometimes you run single job on the entire data center, and then you need to write the software and you need to provide all the interfaces to the software to place the different parts of the jobs more efficiently. Building networks at this scale is a very different story than building compute network on this scale. It is a very different story than build just general purpose data center network.
16:29General purpose data center network is Ethernet. It's not a big deal. Well, it is a big deal, but it is a different deal. It's you are serving loosely coupled collaborative microservices that create the service that you see as a customer from outside. Here you are running one single application on 100 ,000 machines. Yeah. Is that specific to training workloads or is that also true with inference workloads? It's true for everything. It depends on what scale. And the inference is yet another topic that we touched on. Till recently, training was the key thing. A lot of GPUs and there is a very specific way of training is being done.
17:16You basically copy this or another model on multiple machines or multiple sets of machines and run them, then consolidate the results and so on and so forth. On the inference, the story is a little bit different. But the thing is that you need to provide the hooks on the hardware and on your low-level software, system software, for application and for scheduler to place the job and place the different parts of the job in the most efficient way. And as long as your machine fits a building, which which is about 100 ,000 GPUs now talking about, you know, gigawatts, it's all power driven, okay? The challenge is that for many reasons, you want to split your workloads across multiple data centers.
18:11And sometimes data centers are at a distance of many kilometers, many miles. It may be across the continent. And this, you know, it comes with yet another challenge, which is speed of light. Yeah. Okay. Now the latency variance between different parts of your machine is dramatically different. And what is even more challenging is that when you talk about networks, the congestion on the network is one of the key problems that deteriorate network performance. Yeah. And managing congestion across such a latency difference is not like in old telco days you put some box at the edge of your data center with huge buffers and it's a shock absorber for congestion.
19:05Huge buffer is not good. Bigger is not better. There is a famous statement from a very famous woman. And so we need to, and these buffers are basically, these devices are basically to isolate the external world from the internals. but when you want to run a single workload across data centers that are distant by kilometers, you need to be every machine on one side to be aware of whom does it communicate to, whether it's short communication, long communication and adjust all the communication patterns accordingly. So you don't need these big buffers because big buffers is a jitter. Yeah. And so we have a technology We actually developed it recently, technology.
20:00You know, all the internet network is Spectrum X. And this is the device that we designed and developed based on the Spectrum switch that we put on the edge of the data center and it provides all the information and telemetry needed for the endpoints to adjust for the congestion. Yeah. Can we talk a little bit more about training versus inference? How does the shape of workload differ when you're doing, I guess, backprop is a lot more computationally intensive, forward pass less so, but how does the workload differ? And then are you seeing customer demand start to shift from pre-training towards inference or do you think it's still very training heavy right now?
20:46And if I could just ask a quick follow-up question with that, will people be running inference workloads on the same data centers that they use for training or will these end up being two separate? Because they're different optimizations, people end up using two different sets of data centers. Okay, yeah, that's a great question. And let me start with the first one. So training has two phases. One is inference, which is just for propagation, and then back propagation to adjust the weights. And for data parallel training, it's yet another phase to consolidate the results of the weights update across multiple model copies.
21:29So till recently it was the main driver of the compute because till not very long ago, it's maybe two years, which is ages in the AI era, the inference or AI was mainly perceptional. So you show the picture, that's a dog. You show the photo of the person and here's that's Michael and that's Sonia. So that's a single path and that's it. Then it became generative AI where actually you get the recursive generation. So when you pause the prompt, then it's not just one inference. It's many inferences. Because for every token, when you generate text or generate picture, for every new token, you need to go through the entire machine all over again.
22:29So instead of one shot interference, there is more. And then now there is reasoning, which means machine starts, you know, sort of thinking. Yeah. If you ask me what time is it now, I can tell you it's easy, right? What time is it now? But if you ask me more complicated question, then I need to think. I probably need to wait or compare multiple solutions or multiple paths. And every such a thing is inference. Yeah. Every such a thing is inference. And inference itself has actually two phases. One is much more computer intensive and the other one is memory intensive. It's what we call pre-fill because when you do the inference you have some sort of background, right, which is prompt, which is some relevant data that you need to process and create the context to generate the answer.
23:28And this is very computer intensive. It's not much memory intensive. And the other part is actually generating the answer, which is the decode part of the inference, where you generate token by token. Well, there are some technologies that you can generate more than one token, but it's still, you know, single path is much less than the final answer. So if you combine all these things together, Inference demand for computing is actually not less than training. It's actually even more. And there are two reasons for this. One is that what I explained that there's much more computing than it used to be for the inference.
24:15The other thing is, you know, you train model once, but you infer many times. You know, chat GPT, you know, billion of people or it's almost billion of people, right? customers, they are pounding them all the time in the same model. They trained it once. Now they're making videos. So that's a lot of people doing. Right, right. Now they're making videos and you can generate and then, you know, everybody is doing the inference. My wife, I think she talks to ChatGPT more than to me. Once she discovered that's her best friend. So in terms of things, now to your question about about machines, you can infer on the phone.
24:56Okay, so there is a definitely going to be much smaller scale installations for inference. Yeah, it's like mobile devices. If you look at the data center scale and its efficiency of the programming, programmability, is much more viable than optimizations for hardware. And every hardware instance has its own cost and its own drawback. So as long as you don't identify, and I don't think besides this, We actually did it's very similar GPU, it's same programming model as GPU for Prefill versus Decode. I think I don't remember when it happened, but actually we announced that we are building the GPU SKU that is optimized for Prefill.
26:08So you will have, it can do decode and decode GPU can do pre-fill, but you can equip your data center with the SKUs or that pre-fill versus SKUs that are for decode to optimize for like typical use. But if your workload shifts for more decode or for more prefill, you can use either one of them to compensate. And this is the importance of programmability. The same interfaces for GPUs, it's based on CUDA and UP, which is, that's what made NVIDIA NVIDIA before Mellanox. Can I ask you a question about data center scaling? So for many decades, we had Moore's Law and chips got more and more dense and produced better and better performance.
27:05And then we ran into the laws of physics and chips just couldn't get more dense because their quantum mechanical properties caused them to break down. And so then we had to scale up to the rack level and now we've got to scale out to the data center level. Is there some analogous law of data center scaling that says when data centers get too big, the communication overhead causes the performance to break down? Or just said differently, or maybe said more simply, is there a natural limit to how big data centers can get? I think there is a practical limit of how much energy you can consume within given size of the data center.
27:44But if you were surrounded by nuclear power plants and the energy was available, would the data center itself perform? I don't know. I'm not an expert in the construction even, but if you surround you, there's energy coming in, now the heat is going out. So there is a whole bunch. We are now basically moved pretty much entirely to the liquid cooling. And one of the reasons we did it is to enable much denser compute power. We couldn't build as dense computing as we're building now with air cooling. Yeah. So there's a whole bunch of technologies coming to help this more and more denser. Now, the last big data center, which is like XAI scale is 100 or 150 megawatt.
28:38Now we're talking about gigawatt data centers, people talking about 10 gigawatt data centers. So, you know, there is a looking forward to build much bigger data centers. Are you sending the data centers to outer space? Pretty cool. I think, well, one of the things that determines the speed of data center deployment is, you know, how fast concrete gets stable. Fair enough. So before starting Mellanox, you were at Intel. That's right. 16 years? 16 years. You became chief architect. NVIDIA and Intel recently announced a partnership. Can you share a little bit about what the vision for that might be?
29:21The starting point is that computing changed in the last decade or a little bit more than a decade. NVIDIA started as the accelerated computing company. video games was the first and then it evolved to to ai which is the new way of data processing so you cannot just general one human machine just is not capable of of being used as a platform to solve the problem like you know programming when human machine is just explaining somebody what to do. I can explain many things and I can explain many people what to do, but I can't explain how to distinguish between cat and dog. So there is new challenges that AI solves and you need acceleration there.
30:17And our partnership with Intel is actually fusing accelerated computing with the general purpose computing because general purpose computing is not going away. Everything will be accelerated, but we accelerate the purpose computing, we accelerate the applications.
30:37X86 is the architecture that is dominant there and it would serve greatly both companies. That's actually one of the interesting things about NVIDIA. It's the culture of Win-Win. okay we are not after taking a bigger piece of existing pie we are after baking our bigger pie for everybody and the success our success is our customer success it's not our success is not the failure of our competition our success is success of our customers and success of our ecosystem and I think fusing together conventional computing when human machines and accelerated computing that provide with NVIDIA it's actually it's probably open yet another dimension that I'm not sure what it is but it basically gives you know on the practical you know short-term view is it's this gives Nvidia and Intel channels to the market or expanding the market and serving the markets that otherwise was more challenging.
31:56You mentioned the culture of Nvidia. So when Mellanox became part of Nvidia in 2019, the market cap of the combined company was about$100 billion, which is no joke. But the market cap today is about$4.5 trillion. And so 45x growth in value in six years is pretty phenomenal. How has that changed the culture of NVIDIA? How is NVIDIA different today now that it's one of the most admired companies in the world, if not the most admired versus six years ago? Yeah. But this, you know, when we just joined, Jensen was in Israel and I presented him, you know, that's I believe that one plus one will be 10.
32:36and I actually was off by a factor of four. But Mellanox and NVIDIA, in a sense, it's sort of similar. The culture is very similar to begin with, but there is nothing absolutely similar. And I was the only founder that left in Mellanox after Real resigned, a few months after the acquisition. And my main focus at the beginning, you know, things about what do you think about in the shower was how to make sure that this acquisition will succeed. Yeah. You know, NVIDIA paid$7 billion for a company that I founded. And, you know, it was all the mixed feelings that were there. But once it's done, it's done.
33:30Now I have to make it successful. Yeah. So eventually it worked. Most of the Israeli employees are state. I think 85 or 90 % of regional employees are state. Actually, NVIDIA grew more than 2x in Israel in terms of manpower. So we're growing and we are announcing that we're actually going to build a campus in Israel, new campus for NVIDIA. And so that's where I think the overall merger was very successful. I did my best to make sure it succeeds. And besides the technology that I was looking at, this part of which is sort of technology, but it's technology and theology and there's many other things to make sure that people are comfortable that you know from being in the center of melanox which is the headquarters of israel don't feel left you know somewhere in the far away and jensen is key emphasizes is the networking is the critical part of NVIDIA success.
34:59And he's right. Yeah. So I think it's considered to be the most successful merger in the history of the technology. You guys probably track the things better than I am, but overall I think it was a great move. What are the science fiction things that you spend your time thinking about? just even wondering, like for example, optical interconnects, do you think that will exist? Do you think AI will ever be better at physics than us and better at data science than us? Well, what I'm thinking, you know, if you look about science fiction is how to make history to be experimental science. You know, again, physics do try something and then, you know, see what works and then try something else in history.
Read the full transcript
35:47time goes one direction but you have a good simulation of the world you can make history experiments and we have an Earth 2 climate simulator and with this type of technology we can actually simulate how what we do today will impact the global warming 50 years from now experimental science you try something, you see what happens 50 years later so So that's the science fiction part. Yeah. And, you know, the physics, now we are moving from reasoning and so on and so forth. Now, once we get AI models to understand physics, we actually can learn physics. Yeah. AI can teach us physics because the way we get to the laws of physics is that we observe theoretical physics, right?
36:42you observe some phenomena and you generalize it and you compose the rule that basically the law, the physics law that stays underneath of this phenomena. And AI is really great of generalizing and data processing and observing. So AI can help us to get to know some laws of physics that we don't even imagine now. Yeah. Moore's law was 2x every two years. Huang plus Kagan's law is, what is the slope and how long do you think you can sustain it? Well, the slope is somewhere in the range of 10x or few orders of magnitude a year. And that's what we are doing, by the way, now we are, since about two or three years ago, we accelerated our product introduction from every other year to every year.
37:36Now we introduce a new wave of products every year and it's an order of magnitude higher performance and it's not on the chip level performance, it's on the machine that you can build with this performance. That's what we are looking at, it's a single unit of computing. And how How long it will stay? I don't know. I don't know. But we'll do our best to maintain it as long as needed and probably even accelerate. It's all about exponent. It's all about exponent. It's hard to imagine. You know, if you look at this Mouleau curves or any neural curves, they usually plot it on the logarithmic scale.
38:22So it looks like linear. But that's the wrong thing to look. You can't predict what's going to happen. Who could predict that when iPhone was first introduced or smartphone was first introduced, you know, that's 15 years ago? 2007 was the iPhone. Yeah, 2007. Oh, 17 years ago. Okay, who could imagine that this smartphone, the least used function at least for me is a phone yeah unless it's e-commerce it's uh it's texting it's news it's mail it's basically running running your life from this machine so your authentication your id is there it's so you know now who can imagine what's going to happen uh you know 10 years from now with all these developments that we are we are doing today but we are building the platform for innovation.
39:14What is the, your commentary on who can imagine notwithstanding, what is the most optimistic view of our future with AI that you like to think about? What could AI do for the world 5, 10, 15 years from now? Steve Jobs called computer to be the bicycle of mind. Yeah. Okay, so AI is, it's maybe, I don't know if it's, it's probably spaceship. Because there is a lot of things that I would like to do, but I just don't have enough time, don't have enough resources to do it. With AI, I will have it. And it doesn't mean that, you know, I will do twice as much. Maybe I will do 10 times as much. But the thing is that I will want to do 100 times as much as I want to do today.
40:07And that's where, you know, you go to any project leader and nobody says, you know, I have enough. I have enough manpower. I have enough resources. I don't need any more. Okay. If you give him a resource which is twice as efficient, he will do four times more. Yeah. And he will want to do 10 times more. So it's like electricity changes the world, right? instead of using, you know, in London, you still see these gas lamps and this infrastructure to use the gas as the source of energy. Who could think that, you know, once this electricity was invented, it will change the world that, you know, we can't live without electricity.
40:51The same with AI. Beautifully said. Thank you so much for joining us today. I love this conversation. Thank you. Thank you. Thank you for having me.
41:28you
From the publisher
Recorded live at Sequoia’s Europe100 event: Michael Kagan, co-founder of Mellanox and CTO of Nvidia, explains how the $7 billion Mellanox acquisition helped transform Nvidia from a chip company into the architect of AI infrastructure. Kagan breaks down the technical challenges of scaling from single GPUs to 100K and eventually million-GPU data centers. He reveals why network performance—not just compute power—determines AI system efficiency. He discusses the shift from training to inference workloads, and his vision for AI as humanity's "spaceship of the mind," and why he thinks AI may help us discover laws of physics we haven’t yet imagined.
Hosted by Sonya Huang and Pat Grady




