In short
BG2Pod Episode 17: Welcome Jensen Huang
Podcast Overview
- Title: BG2Pod with Brad Gerstner and Bill Gurley
- Description: A bi-weekly conversation focusing on technology, markets, investing, and capitalism.
- Episode Guest: Jensen Huang, CEO of NVIDIA
- Episode Focus: The discussion centers around the rapid evolution of Artificial General Intelligence (AGI), machine learning acceleration, NVIDIA’s competitive advantages, and implications of AI across various industries.
Chapter Breakdown
- [00:00] Introduction
- Introduction of Jensen Huang and the context of the discussion.
- [01:50] The Evolution of AGI and Personal Assistants
- Huang discusses the predicted arrival of personal assistants powered by AGI.
- [06:03] NVIDIA's Competitive Moat
- Exploration of NVIDIA's strategic advantages and unique approach in the tech landscape.
- [15:51] The Future of Inference and Training in AI
- Discussion on the changing dynamics of AI training and inference.
- [19:01] Building the AI Infrastructure
- Insights into NVIDIA's infrastructure supporting AI advancements.
- [31:35] Inventing a New Market in an AI Future
- The role NVIDIA plays in shaping new markets in AI.
- [38:40] The Impact of OpenAI
- Examination of OpenAI’s influence on AI technology and adoption.
- [43:25] The Future of AI Models
- Predictions and expectations for future AI models.
- [46:44] X.ai and Memphis Supercluster
- Discussion of the ambitious AI supercluster initiative led by Elon Musk.
- [51:21] Distributed Computing and Inference Scaling
- The importance of distributed computing in AI applications.
- [55:54] Inference Time Reasoning and Its Importance
- Huang addresses the significance of real-time reasoning in AI.
- [01:00:46] AI's Role in Growing Business and Improving Productivity
- How AI can enhance business productivity and efficiency.
- [01:08:00] Ensuring Safe AI Development
- Conversations about responsible AI development practices.
- [01:12:31] The Balance of Open Source and Closed Source AI
- The interplay between open-source and proprietary AI models.
Key Concepts and Insights
AGI and Personal Assistants
- Huang predicts that the development of personal assistants powered by AGI is imminent.
- Enhanced personal assistants will improve over time due to advancements in AI technology.
NVIDIA's Competitive Moat
- Huang emphasizes the importance of NVIDIA's full-stack approach, not just focusing on hardware (GPUs) but also on software, libraries, and the entire data pipeline.
- Innovations in machine learning techniques amplify NVIDIA's competitive edge.
Inference vs. Training
- The discussion highlights a shift in AI complexity where post-training and inference are becoming increasingly challenging, requiring innovative solutions.
AI Infrastructure
- NVIDIA is positioned to be the backbone of AI infrastructure necessary for the future of computing.
- The emphasis on the entire data center as a computing unit reflects Huang's vision for integrated AI systems.
OpenAI's Impact
- OpenAI is recognized as a transformative force in AI, with significant partnerships and advancements that extend its influence across various domains.
Future of AI Models
- Huang discusses the growing importance of both open-source and proprietary AI models.
- The balance between these models is essential for fostering innovation while ensuring safety and accessibility.
Distributed Computing
- The conversation touches on the necessity for distributed computing strategies to support the growth of AI applications.
Productivity and Business Growth
- AI is portrayed as a tool for enhancing productivity within organizations, allowing for expansion and innovation rather than layoffs.
Safe AI Development
- Huang stresses the importance of developing AI technologies responsibly, ensuring that safety measures are built into the systems.
Conclusion
- The episode illustrates the rapid advancements in AI technology and NVIDIA’s pivotal role in shaping the future of computing.
- Huang's perspectives provide a compelling vision of how AI will integrate into everyday life, transforming industries and personal experiences.
Host and Guest Acknowledgments
- Appreciation for Brad Gerstner, Bill Gurley, and Clark Tang for facilitating the insightful discussion alongside Jensen Huang.
Key Takeaways
- The future of AI and its applications are expanding rapidly, with significant potential for productivity and innovation.
- Balancing open-source and closed-source models is crucial for a sustainable and innovative AI ecosystem.
- The focus on ethical AI development and safety is paramount in fostering public trust and advancing technology responsibly.
Hashtags jensenhuang #nvidia #bradgerstner #billgurley #clarktang #xai #memphiscluster #elonmusk #openai
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00what they achieved is singular, never been done before. Just a put in perspective, 100 ,000 GPUs, that's easily the fastest supercomputer on the planet. That's one cluster. A supercomputer that you would build would take normally three years to plan. Right. And then they deliver the equipment and it takes one year to get it all working. We're talking about 19 days. MUSIC Jensen's nice glasses. Hey, yeah. You too. It's great to be with you. Yeah. I got my ugly glasses on just like you. Come on, those aren't ugly. They're pretty good. Do you like the red ones better? There's something all your family could love.
0:54Ha ha ha. Well, it's Friday, October 4th, we're at the NVIDIA headquarters just down the street from Altimeter. Welcome. Thank you. Thank you. And we have our investor meeting, our annual investor meeting on Monday. Where we're going to debate all the consequences of AI, how fast we're scaling intelligence. And I couldn't think of anybody better and really to kick it off with than you, as both a shareholder, as a thought partner, kicking ideas back and forth. You really make us smarter. And we're just grateful for the friendship. So thanks for being here. Happy to be here. You know, this year the theme is scaling intelligence to AGI.
1:30And it's pretty mind boggling that when we did this two years ago, we did it on the age of AI. And that was two months before chat, GBT. And to think about all this change. So I thought we would kick it off with a thought experiment and maybe a prediction. Yeah. If I colloquially think of AGI as that personal assistant in my pocket. Thanks. If I think of AGI as that colloquial assistant. I'll give you my pocket. Use to it. Exactly. Yeah. You know, it knows everything about me. That's perfect memory of me. That can communicate with me. That can book a hotel for me, or maybe book a doctor's appointment for me.
2:06When you look at the rate of change in the world today, when do you think we're going to have that personal assistant in our pocket? Soon in some form. Yeah. Yeah, soon in some form. And that that that assistant will get better over time. That's the beauty of technology as we know it. And so I think in the beginning it'll be quite useful, but not perfect. And then it gets more and more perfect over time, like all technology. When we look at the rate of change, I think Elon has said the only thing that really matters is rate of change. It sure feels to us like the rate of change has accelerated dramatically.
2:46Is the fastest rate of change we've ever seen on these questions? Because we've been around the rim like you on AI for a decade now. You even longer. Is this the fastest rate of change you've seen in your career? It is because we've reinvented computing. You know, a lot of this is happening because we drove the marginal cost of computing down by 100 ,000 X over the course of 10 years. More's law would have been about 100 X. And we did it in several ways. We did it by one introducing accelerated computing, taking what is work that is not very effective on CPUs and putting it on top of GPUs. We did it by inventing new numerical precision.
3:33We did it by new architectures, inventing an tensor core. The way systems are formulated MVLINK added insanely, insanely fast memories, HBM, and scaling things up with MVLINK and Infiniban and working across the entire stack. Right. Basically, everything that I describe about how Nvidia does things led to a super Moore's law rate of innovation. Now, the thing that's really amazing is that as a result of that, we went from human programming to machine learning. And the amazing thing about machine learning is that machine learning can learn pretty fast, as it turns out. And so as we reformulated the way we distribute computing, we did a lot of parallelism of all kinds, tensor parallelism, pipeline parallelism, parallelism of all kinds.
4:30And we became good at inventing neo -ogritism on top of that and new training methods. And all of this invention is compounding on top of each other as a result. And back in the old days, if you look at the way Moore's law was working, the soft wall was static. It was pre -compiled as shrink wrapped, put into a store, it was static. And the hardware underneath was growing at Moore's law rate. Now we've got the whole stack growing, innovating across the whole stack. And so I think that's the, now all of the sun we're seeing scaling. That is extraordinary, of course. But we used to talk about pre -trained models and scaling at that level.
5:17And how we're doubling the model size and doubling therefore appropriately and doubling the data size. And as a result, the computing capacity necessary is increasing by a factor of four of a year. Right. That was a big deal. Right. But now we're seeing scaling with post -training and we're seeing scaling at inference, isn't that right? And so people used to think that pre -training was hard and inference was easy. Now everything is hard, which is kind of sensible. The idea that all of human thinking is one shot is kind of ridiculous. And so there must be a concept of fast thinking and slow thinking and reasoning and reflection and iteration and simulation and all that.
6:01And that now it's coming in. I think to that point, one of the most misunderstood things about NVIDIA is how deep the true NVIDIA mode is. I think there's a notion out there that as soon as someone invents a new chip, a better chip that they've won. But the truth is you've been spending the past decade building the full stack from the GPU to the CPU to the networking and especially the software and libraries that enable applications to run on NVIDIA. So I think you spoke to that. But when you think about NVIDIA's mode today, do you think NVIDIA's mode today is greater or smaller than it was three to four years ago?
6:46Well, I appreciate you recognizing how computing has changed. In fact, the reason why people thought, and many still do, that you designed a better chip, it has more flops, has more flips and flops and bits and you know what I'm saying? And you see their keynotes slides and it's got all these flips and flops and a bar chart and things like that. And that's all good. I mean, look, a horsepower does matter. So these things fundamentally do matter. However, unfortunately, that's old thinking. It is old thinking in the sense that the software was some application running on Windows and the software static, which means that the best way for you to improve the system is just making faster and faster chips.
7:38But we realized that machine learning is not human programming. Machine learning is not about just a software. It's about the entire data pipeline. It's about, in fact, the flywheel of machine learning is the most important thing. So how do you think about enabling this flywheel on the one hand and enabling data scientists and researchers to be productive in this flywheel? And that flywheel starts at the very, very beginning. A lot of people don't even realize that it takes AI to curate data to teach an AI. And that AI alone is pretty complicated. And so... And so... And so... Accelerating, again, when we think about the competitive advantage.
8:27Right? It's combinatorial evolving system. Exactly. Exactly. And I was exactly going to lead to that because of smarter AI's to curate the data. We now even have synthetic data generation and all kinds of different ways of curating data, presenting data to... And so before you even get the training, you've got massive amounts of data processing involved. And so people think about, oh, a pie torch, that's the beginning and the end of the world. And it was very important. But don't forget, before pie torches a month amount of work. After pie torches amount of work. And that... The thing about the flywheel is really the way you have to think.
9:06You know, how do I think about this entire flywheel? And how do I design a computing system, a computing architecture that helps you take this flywheel and be as effective as possible? It's not one -size -size slice of an application. Training, does that make sense? That's just one step, okay? Every step along that flywheel is hard. And so the first thing that you should do, instead of thinking about, how do I make Excel faster? How do I make, you know, doom faster? That was kind of the old days, isn't that right? Now you have to think about, how do I make this flywheel faster? And this flywheel has a whole bunch of different steps.
9:43And there's nothing easy about machine learning as you guys know. There's nothing easy about what OpenAI does or X does or Gemini in the team that DeepMind does. I mean, there's nothing easy about what they do. And so we decided, look, this is really what you'd be thinking about. This is the entire process. You want to accelerate every part of that. You want to respect Amdahl's law. You want to, Amdahl's law would suggest, well, if this is 30 % of the time, and I accelerated that by a factor of three, I didn't really accelerate the entire process by that much. Does that make sense? And you really want to create a system that accelerates every single step of that, because only in doing the whole thing, can you really, materially improve that cycle time.
10:29And that flywheel, that rate of learning is really in the end what causes the exponential rise. And so what I'm trying to say is that our perspective about a company's perspective about what you're really doing manifests itself into the product. And notice, I've been talking about this flywheel, you know, entire cycle. That's right. And we accelerate everything. Right now, right now, the main focus is video, a lot of people are focused on physical AI and video processing. Just imagine that front end, the terabytes per second of data that are coming into the system, give me an example of a pipeline that is going to ingest all of that data, prepare for training in the first place.
11:20So that entire thing is could accelerated. And people are only thinking about text models today. But the future is this video models, as well as using some of these text model, like O1, to really process a lot of that data before we even get there. Yeah, yeah. So language models are going to be involved in everything. Well, it took us, took the industry enormous technology and effort to train a language model to train these large language models. Now we're using a large language model in every single step of the way. It's pretty phenomenal. I don't mean to be overly simplistic about this, but again, you know, we hear it all the time from investors, right?
12:03Yes, but what about custom Asics? Yes, but their competitive mode is going to be pierced by this. What I hear you saying is that in a combinatorial system, the advantage grows over time. So I heard you say that our advantage is greater today than it was three to four years ago, because we're improving every component and that's combinatorial. Is that, you know, when you think about, for example, as a business case study, Intel, right? Who had a dominant mode, a dominant position in the stack, relative to where you are today. Perhaps just, you know, again, boil it down a little bit. You know, compare contrast to your competitive advantage to maybe the competitive advantage they had at the peak of their cycle.
12:49Well, Intel extraordinary, Intel's extraordinary because they were probably the first company that was incredibly good at manufacturing, process engineering, manufacturing. And that one click above manufacturing, which is building the chip. Right. And designing the chip and architecting the chip in the X86 architecture and building faster and faster X86 chips, that was their brilliance. And they fused that with manufacturing. Our company is a little different in the sense that, and we recognize this that in fact, parallel processing doesn't require every transistor to be excellent. Serial processing requires every transistor to be excellent.
13:45Parallel processing requires lots and lots of transistors to be more cost effective. I rather have 10 times more transistors, 20 % slower than 10 times less transistor, 20 % faster. Doesn't make sense? They would like the opposite. And so single threaded performance, single threaded processing and parallel processing was very different. And so we observe that in fact, our world is not about being better going down. We want to be very good as good as we can be, but our world is really about much better going up. Parallel computing, parallel processing is hard because every single algorithm requires a different way of refactoring and re -architecting the algorithm for the architecture.
14:32What people don't realize is that you can have three different ISOs, CPU ISOs. They all have their own C compiler as you could take software and compile down to that ISO. That's not possible in Excel ready computing. That's not possible in parallel computing. The company who comes up with the architecture has to come up with their own open GL. So we revolutionize deep learning because of our domain -specific library called KudianN. Without Kudian, nobody talks about KudianN because it's one layer underneath PyTorge and TensorFlow and back in the old days, Cafe and Fiano and now Triton. There's a whole bunch of different frameworks.
15:10So that domain -specific library, KudianN, a domain -specific library called Optics. We have a domain -specific library called KuKwanTum, Rapids, the list of aerial for industry -specific algorithms that sit below that PyTorge layer that everybody's focused on. Like I've heard oftentimes, well, you know, with LLM subs. If I didn't invent that, no application on top could work. You guys understand what I'm saying? So the mathematics is really, what Envite is really good at is algorithm. That the fusion between the science above the architecture on the bottom, that's what we're really good at. There's all this attention now on inference, finally.
15:56But I remember, you know, two years ago, Brad and I had dinner with you, and we asked you the question, do you think your moat will be as strong in inference as it is in training? And I'm sure I said it would be greater. Yeah. And you touched upon a lot of these elements just now, just the composability between, or we don't know that total mix at one point into a customer, it's very important to be able to be flexible in between. That's right. But can you just touch upon now that we're in this era of inference? It was inference, training is inferencing at scale. I mean, you're right. And so if you train well, it is very likely you'll inference well.
16:44If you built it on this architecture, without any consideration, it will run on this architecture. You could still go and optimize it for other architectures, but at the very minimum, since it's already been architected, built on Nvidia, it will run on Nvidia. Now, the other aspect, of course, is just kind of capital investment aspect, which is when you're training new models, you want your best new gear to be used for training, which leaves behind gear that you used yesterday, while that gear is perfect for inference. And so there's a trail of free gear. There's a trail of free infrastructure behind the new infrastructure that's coulda compatible.
17:30And so we're very disciplined about making sure that we're compatible throughout, so that everything that we leave behind will continue to be excellent. Now, we also put a lot of energy into continuously reinventing new algorithms, so that when the time comes, the hopper architecture is two, three, four times better than when they bought it, so that infrastructure continues to be really effective. And so all of the work that we do, improving new algorithms, new frameworks, notice, it helps every single install base that we have, hopper is better for it, ampere is better for it, even Volta is better for it.
18:11And I think Sam was just telling me that they had just decommissioned the Volta infrastructure that they have at OpenAI recently. And so I think we leave behind this trail of install base, just like all computing, install base matters. And in videos in every single cloud we're on -prem and all the way out to the edge. And so the Vila vision language model that has been created in the cloud works perfectly at the edge on a robots. Without modification, it's all coulda compatible. And so I think this idea of architecture compatibility was important for large, it's no different for iPhones, no different for anything else.
18:52I think the install base is really important for inference. But the thing that I really, really, we really benefit from is because we're working on training these large language models in the new architectures of it, we're able to think about how do we create architectures that's excellent at inference someday when the time comes. And so we've been thinking about iterative models for reasoning models and how do we create very interactive inference experiences for this personal agent of yours. You don't wanna say something that I have to go off and think about for a while. You want it to interact with you quite quickly.
19:34So how do we create such a thing and what came out of it was MV link. You know, MV link so that we could take these systems that are excellent for training, but when you're done with it, the inference performance is exceptional. And so you wanna opt -in for this time to first token. And time to first token is insanely hard to do actually. Because time to first token requires a lot of bandwidth. But if your context is also rich, then you need a lot of flops. And so you need an infinite amount of bandwidth, infinite amount of flops at the same time in order to achieve just a few millisecond response time.
20:15And so that architecture is really hard to do. And we invented a Grace Blackwell MV link for that. Right. In the spirit of time, I have more questions about that. But don't worry about the time. Hey guys, hey, hey, listen, Janine. Yeah. Look. Let's do it until the right. There you go. I love it. So, you know, I was at dinner with Andy Jassy here. We know we don't have to worry about that. Early with Andy Jassy earlier this week. And Andy said, you know, we've got training, you know, coming and infreence coming. And I think most people, again, view these as a problem for Nvidia. But in the very next breath, he said, Nvidia is a huge and important partner to us and will remain a huge and important partner for us as far as I can see into the future.
21:02The world runs on Nvidia. Right. So when you think about the custom A6 that are being built, that are going to go after targeted application, maybe the inference accelerator at Meta, maybe, you know, Trainiam at Amazon, you know, or Google's TPUs. And then you think about the supply shortage that you have today. Do any of those things change that dynamic, right? Or are they compliments to the systems that they're all buying from you? We're just doing different things. Yes. We're trying to accomplish different things. You know, what Nvidia is trying to do is build a computing platform for this new world, this machine learning world, this generative AI world, this agentic AI world.
21:47We're trying to create, you know, as you know, what's just so deeply profound is after 60 years of computing, we reinvented the entire computing stack. The way you write software from programming to machine learning, the way that you process software from CPUs to GPU, the way that the applications from software to artificial intelligence, right? And so software tools to artificial intelligence. So every aspect of the computing stack and the technology stack has been changed, you know, what we would like to do is to create a computing platform that's available everywhere. And this is really the complexity of what we do.
22:31The complexity of what we do is if you think about what we do, we're building an entire AI infrastructure and we think of it as one computer. I've said before, the data center is now the unit of computing. To me, when I think about a computer, I'm not thinking about that chip. I'm thinking about this thing. That's my mental model and all the software and all the orchestration, all the machinery that's inside. That's my computer. And we're trying to build a new one every year. Yes. That's insane. Nobody has ever done that before. We're trying to build a brand new one every single year and every single year we deliver two or three times more performance.
23:08As a result, every single year we reduce the cost by two or three times. Every single year we improve the energy efficiency by two or three times, right? And so we ask our customers, don't buy everything at one time, buy a little every year. And the reason for that, we want them cost average into the future. All of it's architecturally compatible. Now, so that building that alone at the pace that we're doing is incredibly hard. Now, the double part, the double hard part, is then we take that all of that and instead of selling it as a infrastructure or selling it as a service, we disaggregate all of it and we integrate it into GCP.
23:49We integrate it into AWS. We integrate it into Azure. We integrate it into X, we integrate, doesn't make sense? And so everybody's integration is different. We have to get all of our architectural libraries and all of our algorithms and all of our frameworks and integrate it into theirs. We get our security system integrated into theirs. We get our networking integrated into theirs, isn't it right? Then we do basically 10 integrations and we do this every single year. Now, that is the miracle. That is the miracle. Why were you, I mean, it's madness. It's madness that you're trying to do this every year.
24:27So what drove you to do it every year and then related to that, you know, Clark's just back from Taipei and Korea and Japan when meeting with all your supply partners who you have decade -long relationships with. How important are those relationships? To again, the commonatorial math that builds that competitive mode. Yeah, that's when you break it down systematically, the more you guys break it down, the more everybody breaks it down, the more amazed that they are. Yes. And how is it possible that the entire ecosystem of electronics today is dedicated in working with us to build ultimately this cube of a computer integrated into all of these different ecosystems and the coordination is so seamless.
25:20So there's obviously APIs and methodologies and business processes and design rules that we've propagated backwards and methodologies and architectures and APIs that we've propagated forward. That have been hardened for decades. Hardened for decades, yeah, and also evolving as we go. But these APIs have to come together. When the time comes, all these things in Taiwan, all over the world being manufactured, they've got land somewhere in Azure's data center, they're gonna come together, they're click, click, click, click. Someone just calls an opening API and it just works. That's right. Yeah, exactly.
26:00It's kind of crazy, right? It's a whole chain. So that's what we invented. That's what we invented. This mass of infrastructure of computing, the whole planet is working with us on it. It's integrated into everywhere. You could sell it through Dell, you could sell it through HPE. It's hosted in the cloud. It's all the way out at the edge. People use it in robotic systems now, human and robots. They're in self -driving cars. They're all architecturally compatible. Pretty kind of craziness. It's craziness. Clark, I don't want to, don't want you to leave the impression. I didn't answer the question.
26:37In fact, I did. What I meant by that when, when I was in T. Racic, is the way to think about, we're just doing something different. Yes. As a company, as a company, we want to be situationally aware, and I'm very situationally aware of everything around our company and our ecosystem. Right. I'm aware of all the people doing alternative things and what they're doing. And sometimes it's adversarial to us. Sometimes it's not. I'm super aware of it. But that doesn't change what the purpose of the company is. The singular purpose of the company is to build an architecture, that a platform that could be everywhere.
27:24That is our goal. We're not trying to take any share from anybody and Vity is a market maker, not share taker. If you look at our company slides, we don't show, not one day does this company talk about market share, not inside. All we're talking about is how do we create the next thing, what's the next problem we can solve in that flywheel? How can we do a better job for people? How do we take that flywheel? They used to take about a year. How do we crank it down to about a month? Yes. What's the speed of light of that? Isn't that right? And so we're thinking about all these different things, but the one thing we're not to, situationally aware of everything, but we're certain that what our mission is, is very singular.
28:09The only question is whether that mission is necessary. Does that make sense? And all companies, all great companies, ought to have that at its core. It's about what are you doing? The only question is a necessary, it's a valuable, is it impactful? Does it help people? And I am certain that your developer, your generative AI start up, and you're about to decide how to become a company, the one choice that you don't have to make is which one of the A6 do I support? If you just supported Kuda, you know you could go everywhere. You could always change your mind later. But we're the on ramp to the world of the AI, isn't that right?
28:51Once you decide to come onto our platform, the other decisions you could defer. You could always build your own A6 later. You know, we're not against that, we're not offended by any event. When I work with, when we work with all the GCPs, the GCPs, we present our roadmap to them. Years in advance, they don't present their A6 roadmap to us. And it doesn't ever offend us. Does that make sense? We create, if you have a sole purpose, and your purpose is meaningful, and your mission is dear to you and is dear to everybody else, then you could be transparent. Notice my roadmap is transparent at GTC.
Read the full transcript
29:30My roadmap goes way deeper to our friends at Azure and AWS and others. We have no trouble doing any of that, even as they're building their own A6. I think, you know, when people observe the business, you said recently that the demand for Blackwell is insane. You said one of the hardest parts of your job is the emotional toll of saying no to people in a world that has a shortage of the compute that you can produce and have on offer. But critics say this is just a moment time, right? They say this is just like Cisco in 2000, we're over building fiber, it's gonna be boom and bust. You know, I think about the start of 23 when we were having dinner.
30:17The forecast for Nvidia at that dinner in January of 23 was that you would do 26 billion of revenue for the year 2023. You did 60 billion, right? The 25 people. It was just, let's let the truth be known. That is the single greatest failure of forecasting the world has ever seen. Right, right, right. Can we all at least admit that? What, what, what, to me? That was my takeaway. I just go. And that was, and that was, we got so excited in November 22 because we had folks like Mustafa from inflection and no one from character coming in our office talking about investing in their companies and they said, well, if you can't, pencil out investing in our companies, then buy Nvidia because everybody in the world is trying to get Nvidia chips to build these applications that are gonna change the world.
31:07And of course, the Cambrian moment occurred with chat GPT and notwithstanding that fact, these 25 analysts were so focused on the crypto winner that they couldn't get their head around an imagination of what was happening in the world. Okay, so it ended up being way bigger. You say in very plain English, the demand is insane for black well that it's going to be that way for as far as you, you know, for as far as you can see. Of course, the future is unknown and unknowable, but why are the critics so wrong that this isn't going to be the Cisco -like situation of overbuilding in 2000? Yeah. Yeah.
31:48The best way to think about the future is reason about it from first principles. Correct. Okay, so the question is, what are the first principles of what we're doing? Number one, what are we doing? What are we doing? The first thing that we are doing is we are reinventing computing. Do we not? We just said that. The way that computing will be done the future will be highly machine learned. Yes. Highly machine learned. Okay, almost everything that we do, almost every single application, word, Excel, PowerPoint, Photoshop, Premiere, you know, AutoCAD, you give me your favorite application that was all hand engineered.
32:29I promise you it will be highly machine learned in the future, isn't it, right? And so all these tools will be, and on top of that, you're going to have machines, agents that you help you use them. Right. Okay, and so we know this for a fact at this point, right? Isn't that right? We've reinvented computing. We're not going back. The entire computing technology stack has been reinvented. Okay, so now that we've done that, we said that software is going to be different. What software can write is going to be different. How we use software will be different. So let's now acknowledge that. Though those are my ground truth now.
33:01Yes. Now the question therefore is what happens? And so let's go back and let's just take a look at how is computing done in the past. So we have a trillion dollars with the computers in the past. We look at, just open the door, look at the data center and you look at it, are those the computers you want doing that? Doing that in the future? And the answer's no. Right, you got all these CPUs back there. We know that what it can do and what it can't do. And we just know that we have a trillion dollars with the data centers that we have to modernize. And so right now as we speak, if we were to have a trajectory over the next four, five years to modernize that old stuff, that's not unreasonable, sensible.
33:37So we have a trillion. And you're having those conversations with the people who have to modernize it. Yeah. And they're modernizing it on GPUs. That's right. I mean, well, let's make another test. You have 50 billion dollars of CapEx, you like to spend, option A, option B, build CapEx for the future, or build CapEx like the past. You already have the CapEx of the past. It's sitting right there. It's not getting much better anyways, and more so largely ended. And so why rebuild that? Let's just take 50 billion dollars, put it into generative AI, isn't that right? And so now your company just got better.
34:13Now how much of that 50 billion would you put in? Well, I would put in 100 % of the 50 billion, because I already got four years of infrastructure behind me, that's the of the past. And so now you just, I just reasoned about it from the perspective of somebody thinking about it from first principles. And that's what they're doing. Smart people are doing smart things. Now the second part is that, so now we have a trillion dollars with a capacity to go build, right? Trillion dollars with a infrastructure. We're about, you know, call it 150 billion dollars into it. Right. Okay? So we have a trillion dollars in infrastructure, infrastructure to go build over the next four or five years.
34:46Well, the second thing that we observe is that the way that software is written is different, but how software is gonna be used is different. In the future, we're gonna have agents, isn't that right? We're gonna have digital employees in our company. In your inbox, you have all these little dots and these little faces. It's in the future, it's in the little icons of AIs, isn't that right? I'm gonna be sending them, I'm gonna be no longer gonna program computers with C++. I'm gonna program AIs with prompting, isn't that right? Now there's no different than me talking to my, you know, this morning I wrote a bunch of emails before I came here.
35:24I was prompting my team, right? And I would describe the context. I would describe the fundamental constraints that I know of. And I would describe the mission for them. I would leave it sufficiently, I would be sufficiently directional so that they understand what I need. I wanna be clear about what the outcome should be as clear as I can be, but I leave enough ambiguous space on a creativity space so they can surprise me, isn't that right? There's no different than how I prompt an AI today. It's exactly how I prompt an AI. And so what's gonna happen is, is on top of this infrastructure of IT that we're gonna modernize, there's gonna be a new infrastructure.
36:03This new infrastructure are going to be AI factories that operate these digital humans. And they're gonna be running all the time, 24 -7. We're gonna have them for all of our companies all over the world. We're gonna have them in factories. We're gonna have them in autonomous systems, isn't that right? So there's a whole layer of computing fabric, a whole layer of what I call AI factories that the world has to make that doesn't exist today at all. So the question is how big is that? Unknowable at the moment, probably a few trillion dollars. Unknowable at the moment, but as we're sitting here building into the beautiful thing is, the architecture for this modernizing this new data center, and the architecture for the AI factory is the same.
36:49That's the nice thing. And you made this clear. You've got a trillion of old stuff, you've got to modernize, you at least have a trillion of new AI workloads coming on. You give or take, you'll do 125 billion in revenue this year. You know, there was one point somebody told you the company would never be worth more than a billion. As you sit here today, is there any reason, right? If you're only 125 billion out of a multi trillion, that you're not going to have 2X the revenue, 3X the revenue in the future that you have today. Is there any reason your revenue doesn't? No, yeah. As you know, it's not about, it's not about, everything is, you know, companies, companies are only limited by the size of the fish pond.
37:36You know, a gold fishing can only be so big. And so the question is, what is our fish pond? What is our pond? And that requires a little imagination. And this is the reason why market makers think about that future, without creating that new fish pond, it's hard to figure this out looking backwards and try to take share. You know, share takers can only be so big. Market makers can be quite large. For sure. Yeah. And so, you know, I think the good fortune that our company has is that since the very beginning of our company, we had to invent the market for us to go swim in. That mark, and people don't realize this back then, but anymore, but, you know, we were at the, at the ground zero of creating the 3D gaming PC market.
38:22Right. Right. We largely invented this market and all the ecosystem and all the graphics card ecosystem, we invented all that. And so the need to invent a new market to go serve it later is something that's very comfortable for us. Exactly. And speaking to somebody who's invented a new market, you know, let's shift gears a little bit to models and open AI, open AI raised as you know, six and a half billion dollars this week. At like $150 billion valuation, we both participated. Yeah, really happy for them, really happy that came together. Right. Yeah, they did a great stand and the team did a great job.
39:03Reports are that they'll do five billionish of revenue or run rate revenue this year, maybe going to 10 billion next year. If you look at the business today, it's about twice the revenue as Google was at the time of its IPO. They have 250 million. Yeah, 250 million weekly average users, which we estimate is twice the amount Google had at the time of its IPO. And if you look at the multiple of the business, if you believe 10 billion next year, it's about 15 times the forward revenue, which is about the multiple of Google and Met at the time of their IPO. When you think about a company that had zero revenue, zero weekly average users 22 months ago.
39:44Brad has an incredible command of history. When you think about that, talk to us about the importance of OpenAI as a partner to you and OpenAI as a force and driving forward kind of public awareness and usage around AI. Well, this is one of the most consequential companies of our time. The a pure play AI company pursuing the the vision of AGI and whatever its definition. I almost don't think it matters fully what the definition is. Nor do I, really believe that the timing matters. The one thing that I know is that that AI is going to have a road map of capabilities over time. And that road map of capabilities over time is going to be quite spectacular.
40:53And along the way, long before even gets to anybody's definition of AGI, we're going to put it to great use. All you have to do is right now as we speak, go talk to digital biologists, climatech, researchers, material researchers, physical sciences, astrophysicists, quantum chemists. You go ask video game designers, manufacturing engineers, roboticists, pick your favorite, or whatever industry you want to go pick. And you go deep in there and you talk to the people that matter and you ask them, has AI revolutionized the way you work. And you take those data points and you come back and you then get to ask yourself how skeptical do you want to be?
41:50Because they're not talking about AI as a conceptual benefit someday. They're talking about using AI right now. Right now, ag tech, material tech, climate tech, you pick your tech, you pick your field of science. They are advancing, AI is helping them advancing their work right now as we speak. Every single industry, every single company, every university, unbelievable, isn't that right? Right. It is absolutely going to somehow transform business. We know that. Right. I mean, we know it's so tangible you could. It's happening today. It's happening today. It's happening today. And so I think the awakening of AI, Chatchy PT, triggered, is completely incredible.
42:48And I love their velocity and their singular purpose of advancing this field. And so really, really consequential. And they build an economic engine that can finance the next frontier of models. Right. And I think there's a growing consensus in Silicon Valley that the whole model layers commoditizing. Lama is making it very cheap for lots of people to build models. And so early on here, we had a lot of model companies, you know, character and inflection and cohere and Mistral and go through the list. And a lot of people question whether or not those companies can build the escape velocity on the economic engine that can continue funding those next generation.
43:35My own sense is that there's going to be, that's why you're seeing the consolidation, right? It's open AI clearly has hit that escape velocity. They can fund their own future. It's not clear to me that many of these other companies can. Is that a fair kind of review of the state of things in the model layer that we're going to have this consolidation? Like we have in lots of other markets to market leaders who can afford, who have an economic engine and application that allows them to continue to invest?
44:07There's a first of all, there's a different fundamental difference between a model and artificial intelligence. Right. Yeah. A model is an essential ingredient for artificial intelligence. It's necessary but not sufficient. Correct. And so, and artificial intelligence is a capability but for what? Then what's the application? The artificial intelligence for software and cars is related to the artificial intelligence for human or robots but it's not the same. Which is related to the artificial intelligence for a chatbot but not the same. Correct. And so you have to understand the taxonomy of the stack.
44:46Yeah, of the stack. And at every layer of the stack, there will be opportunities but not infinite opportunities for everybody at every single layer of the stack. Now, I just said something. All you have to do is replace the word model with GPU. In fact, this was the great observation of our company 32 years ago that there's a fundamental difference between GPU, graphics chip or GPU versus accelerated computing. And accelerated computing is a different thing than the work that we do with AI infrastructure. It's related but it's not exactly the same. It's built on top of each other. It's not exactly the same.
45:26And each one of these layers of abstraction requires fundamental different skills. Somebody who's really really good at building GPUs have no clue how to be an accelerated computing company. I can, I can, there are a lot of people who build GPUs. And I don't know which one came, we invented the GPU but you know that we're not, we're not the only company that makes GPUs today. And so, they're GPUs everywhere. And but they're not accelerated computing companies. And there are a lot of people who, you know, they're accelerators, accelerators that does application acceleration. But that's different than an accelerated computing company.
46:06And so, for example, a very specialized AI application. Right. Could that, could be a very successful thing? Correct. Meta's MTIA. That's right. But it might not be the type of company that, that had broad reach and broad capabilities. And so, so you've got to, you've got to decide where you want to be. There's opportunities probably in all these different areas. But like building companies, you have to be mindful of the, the, the shifting of the ecosystem and what gets commoditized over time. Recognizing what's a feature versus a product versus a company. For sure. Okay. I just, I just went through, okay.
46:41And there's a lot of different ways you can think about this. Of course, there's one new entrant that has the money, the smarts, the ambition. That's an X .AI. Yeah. Right. And, um, well, there are reports out there that, that you and Larry and Elon had dinner. They talked you out of 100 ,000 H100s. They went to Memphis and built a large coherent supercluster in a matter of months. You know, so first, three, three points don't make a line. Yes, I had dinner with them.
47:16Cossality is it. What do you think about their ability to stand up that supercluster? And there's talk out there that they won another 100 ,000 H200s, right to expand the size of that supercluster. You know, first talk to us a little bit about X and their ambitions and what they've achieved. But also, are we are ready at the age of clusters of 200 and 300 ,000 GPUs? The answer is yes. And then the first, first of all, acknowledgement of achievement, where it's deserved. From the moment of concept to a data center that's ready for NVIDIA to have our gear, gear there to the moment that we powered it on, had it all hooked up, and it did its first training.
48:09Okay. So that first part, just building a massive factory, liquid cooled, energized, permitted, in the short time that was done. I mean, that is like superhuman. Yeah, there's, and as far as I know, there's only one person in the world who could do that. You know, I mean, Elon is singular in this understanding of engineering and construction and large systems and marshalling resources. Incredible. Yeah, just it's unbelievable. And then, and of course, then his engineering team is extraordinary. I mean, the software team is great, the network team is great, the infrastructure team is great. You know, Elon understands this deeply.
48:58And from the moment that we decided to get to go, the planning of with our engineering team, our networking team, our infrastructure computing team, the software team, all of the preparation and advance, then all of the infrastructure, all of the logistics and the amount of technology and equipment that came in on that day, and video and video infrastructure and computing infrastructure and all that technology, to training 19 days. Hang on, you just, you know what, do you don't want to do? Did anybody sleep 24 or 7? No question, then nobody slept. But, but first of all, yeah, some 19 days is incredible.
49:41But it's also kind of nice to just take a step back and just, do you know how many days 19 days this, this is a couple of weeks? Yeah, right. And the amount of technology, if you're ever to see it, is unbelievable. All of the wiring and the networking and, you know, networking and video gear is very different than networking, hyper scale data centers. Okay. The number of wires that go in one node, the back of a computer is all wires. I'm just getting this mountain of technology integrated and all the software, incredible. Yeah, so, so I think, I think what Elon and next team did, and I'm really appreciative that he acknowledges the engineering work that we did with him, and the planning work and all that stuff.
50:24But what they achieved is, is singular, never been done before. Just a put in perspective, 100 ,000 GPUs, that's, you know, easily the fastest supercomputer in the planet, that's one cluster. A supercomputer that you would build would take normally three years to plan. Right. And then they deliver the equipment and it takes one year to get it all working. Yes. We're talking about 19 days. Wow. What's the credit of the Nvidia platform, right? That it's the whole processes are hardened. That's right. Yeah. Everything's already working. Yeah. And of course, there's a whole bunch of, you know, ex algorithms and ex framework and ex stack and things like that.
51:08And we said we got a ton of integration we have to do. But the planning of it was extraordinary. Just pre -planning of it to, you know, end of one is right. Elon isn't end of one. But you answered that question by starting off saying, yes, 200 to 300 ,000 GPU clusters are here. Right. Does that scale to 500 ,000? Does it scale to a million? And does the demand for your products depend on it scaling to millions? That part, the last part is no. My sense is that distributed training will have to work. And my sense is that distributed computing will be invented. Right. And in some form of federated learning and distributed, you know, asynchronous distributed computing is going to be discovered.
52:06And I'm very enthusiastic and very optimistic about that. The, of course, of course, the thing to realize is that the scaling law used to be about pre -training. Now we've gone to multi -modality. We've gone to synthetic data generation. Post -training has now scaled up incredibly synthetic data generation, reward systems, reinforcement learning based. And then now inference scaling has gone through the roof. The idea that a model before it answers your answer had already done internal inference. Incredible. 10 ,000 times is probably not unreasonable. And it's probably done treaser, just probably done reinforcement learning on that.
52:52It's probably done some simulations, surely done a lot of reflection. It probably looked up some data. It looks some information, isn't it? So this context is probably fairly large. I mean, this type of intelligence is, well, that's what we do. That's what we do, isn't that right? And so the ability, this scaling, if you did that math, and you compounded it with, you compounded that with 4x per year, on model size and computing size. And then on the other hand, demand continues to grow in usage. Do we think that we need millions of GPUs? No doubt. Yeah, that is a for certainty now. And so the question is, how do we architect it from a data center perspective?
53:40And that has a lot to do with, are there data centers that are gigawatts at a time, or are there 250 megawatts at a time? And my sense is that you're going to get both. I think analysts always focus on the current architectural bet. But I think one of the biggest takeaways from this conversation is that you're thinking about the entire ecosystem, and many years out. So the idea that because Nvidia is just scaling up or scaling out, it's to meet the future. It's not such that you're only dependent on a world where there's a 500 ,000 or a million GPU cluster. It's by the time there's distributed training, you'll have written the software to enable that.
54:27That's right. Remember without Megatron, that we developed with some seven years ago now, the scaling of these large training jobs would have happened. And so we invented Megatron, we invented Nickel, GPU Direct, all of the work that we did with our DMA. That made it possible for easily to do pipeline parallelism. All the model parallelism that's being done, all the breaking of the distributed training, and all the batching, and all that. All of that stuff is because we did the early work. And now we're doing the early work for the future generation. So let's talk about strawberry and O1. I want to be respectful of your time.
55:12So we'll all the time in the world. Actually, well, you're very generous. Yeah, we've got all the time in the world. But first, I think it's cool that they named O1 after the O1 visa, which is about recruiting the world's best and brightest, and bringing them to the United States. It's something I know we're both deeply passionate about. So I love the idea that building a model that thinks and that takes us to the next level of scaling intelligence. Is an homage to the fact that it's these people who come to the United States by way of immigration that have made us what we are. Bring their collective intelligence to you.
55:51Surely an alien intelligence. Certainly. It was spearheaded by our friend, Noam Brown. Of course, he worked at Pluribus in Cicero when he was at Metta. How big of a deal is inference time reasoning? As a totally new vector of scaling intelligence, separate and distinct from, just building larger models. It's a huge deal. It's a huge deal. A lot of intelligence can't be done a priori. Right. You know. And a lot of computing, even a lot of computing can't be reordered. I mean, just, you know, out of order execution can't be done a priori. And so a lot of things can only be done in runtime. And so whether you think about it from a computer science perspective, or you think about it from intelligence perspective, too much of it requires context.
56:48The circumstance. Right. The quality, the type of answer you're looking for. Sometimes just a quick answer is good enough. It depends on the consequential impact of the answer. You know, it depends on the nature of the usage of that answer. So some answers, please take a night. Some answers take a week. Yes. Is that right? So I could totally imagine me sending off a prompt to my AI. And telling it, you know, think about it for night. Right. Think about it overnight. Don't tell me right away. Right. I want you to think about it all night. And then come back and tell me tomorrow what your best answer and reason about it for me.
57:30And so I think the, the, the, the quality, the segmentation of intelligence from now from a product perspective, there's going to be one shot versions of it. Right. Sure. Yeah. And then there'll be some that take five minutes, you know. And the intelligence layer that roots those questions to the right model. Yeah. For the right use case. I mean, we were using advanced voicemode and O1 preview last night. I was, I was coaching my son for his AP history test. And it was like having the world's best AP history teacher. Yeah, right. Sitting right next to you thinking about these questions, it was truly extraordinary.
58:09Again, my tutor is an AI today. Right. I'm sure of course they're here today, which comes back to this. You know, over 40 % of your revenue today is inference. But inference is about ready because of chain of reasoning. Yeah. Right. It's about to go up by a billion times. Right. By by by by a million X by a billion X. That's right. That's the part that most people have, you know, haven't completely internalized. This is that industry we were talking about, but this is the industrial revolution. That's the production of intelligence. That's right. Right. And yeah, it's going to go up a billion times.
58:47Right. And so, you know, everybody's so hyper focused on Nvidia as kind of like doing training on bigger models. Yeah. Right. Isn't it the case that your revenue if it's 50, 50 today, you're going to do way more inference in the future. Yeah. Right. Then I mean, training will always be important. But just the growth of inference is going to be way larger than we hope than training. We hope. It's almost impossible to conceive otherwise. Yeah, we hope. That's right. That's right. Right. Yeah, I mean, it's good to go to school. Yes. But the goal is so that you can be productive in society later. And so it's good that we train these models.
59:22But the goal is to inference them, you know. Are you already using chain of reason and, you know, tools like O1 in your own business to improve your own business? Yeah. Our cybersecurity system today can't run without without our own agents. We have agents help with design chips. Hopper wouldn't be possible. Blackwold may possible. Ruben, don't even think about it. We have digital. We have AI chip designers, AI software engineers, AI verification engineers. And we build them all inside because, you know, we have the ability and we rather we use the opportunity to explore the technology ourselves.
1:00:01You know, when I walked into the building today, somebody came up to me and said, you know, Ash Jensen about the culture. It's all about the culture. I look at the business, you know, we talk a lot about fitness and efficiency, flat organizations that can execute quickly, smaller teams. You know, Nvidia is in a league of its own, really, you know, about four million of revenue per employee, about two million of profits or free cash flow per employee. You've built a culture of efficiency that really has unleashed creativity and innovation and ownership and responsibility. You've broken the mold on kind of functional management.
1:00:39Everybody likes to talk about all of your direct reports. Is the leveraging of AI, the thing that's going to continue to allow you to be hyper creative while at the same time be inefficient? No question. I'm hoping that someday, Nvidia has 32 ,000 employees today. And we have four, we have 4 ,000 families in Israel. I hope they're well. Thinking of you guys. And I'm hoping that Nvidia someday will be a 50 ,000 employee company with 100 million, you know, AI assistants. And they're in every single group. We'll have a whole directory of AI's that are just generally good at doing things. We'll also have our inbox is going to fall of directories of AI's that we work.
1:01:35We know are really good specialized at our skill. So, AI's will recruit other AI's to solve problems. AI's will be in slack channels with each other. And with humans. And with humans. So, we'll just be one large employee base, if you will. Some of them are digital and AI, some of our biological. And I'm hoping some of them even make electronics. I think from a business perspective, it's something that's greatly misunderstood. You just described a company that's producing the output of a company with 150 ,000 people. But you're doing it with 50 ,000 people. That's right. Now, you didn't say I was going to get rid of all my employees.
1:02:17You're still growing the number of employees in the organization. But the output of that organization is going to be dramatically more. This is often misunderstood. AI will change every job. AI will have a seismic impact on how people think about work. Let's acknowledge that. AI has the potential to do incredible good. It has the potential to do harm. We have to build safe AI. Let's just make that foundational. The part that is overlooked is when companies become more productive using artificial intelligence, it is likely that it manifests itself into either better earnings or better growth or both.
1:03:08And when that happens, the next email from the CEO is likely not a layoff announcement. Of course. Because you're growing. Yeah. And the reason for that is because we have more ideas than we can explore. And we need people to help us think through it before we automate it. And so the automation part of it, AI can help us do. Obviously, it's going to help us think through it as well. But it's still going to require us to go figure out what problems do I want to solve. There are trillion things we can go solve. What problems does this company have to go solve? And select those ideas and figure out a way to automate and scale.
1:03:47And so as a result, we're going to hire more people as we become more productive. People forget that. And if you go back in time, obviously, we have more ideas today than 200 years ago. That's the reason why it's a piece of larger and more people are employed. And even though we're automating like crazy underneath. It's such an important point of this period that we're entering. One, almost all human productivity, almost all human prosperity is the byproduct of the automation and the technology of the last 200 years. You can look at from Adam Smith and Shumpeter's creative destruction, you can look at chart of GDP growth per person over the course of last 200 years.
1:04:31And it's just accelerated. Which leads me to this question. If you look at the 90s, our productivity growth in the United States was about 2 .5 to 3 % a year. And then in the 2000s, it slowed down to about 8%. And then the last 10 years has been the slowest productivity growth. So that's the amount of labor and capital or the amount of output we have for a fixed amount of labor and capital. The slowest we've had on record actually. And a lot of people have debated the reasoning for this. But if the world is as you just described and we're going to leverage and manufacture intelligence, then isn't it the case that we're on the verge of a dramatic expansion in terms of human productivity?
1:05:12That's our hope. That's our hope. And of course, we live in this world, so we have direct evidence of it. Right. We have direct evidence of it, either as isolated a case as an individual researcher for sure, who is able to with AI now explore science at such an extraordinary scale that is unimaginable. That's productivity. Right. 100 % Measure productivity. Or that we're designing chips that are so incredible at such a high pace. And the chip complexities and the computer complexities we're building are going up exponentially while the company's employee base is not measure of productivity. Correct.
1:05:58The software that we're developing better and better and better because we're using AI and supercomputers to help us, the number of employees is growing barely linearly. Okay. Okay. Another demonstration of productivity. I can go into, I can spot check it in a whole bunch of different industries. I could gut check it myself. Yes. You're right. That's right. And so I can, you know, and of course, you can't, we could be over -fed. But the artistry of course is to generalize what is it that we're observing and whether this could manifest in other industries. And there's no question that intelligence is the single most valuable commodity the worlds are known.
1:06:43Right. And now we're going to manufacture it at scale. And we, we, all of us, have to get good at, you know, what would happen if you're surrounded by these AIs. And they're doing things so incredibly well and so much better than you. Right. And when I reflect on that, that's my life. I have 60 direct reports. Right. The reason why they're on, the reason why they're on E -staff is because they're world -class of what they do and they do it better than I do. Much better than I do. Right. I have no trouble interacting with them. And I have no trouble prompt engineering them. Right. You guys got to start.
1:07:22I have no trouble programming them. Right. Right. And so, so I think that that's, that's the thing that that people are going to learn is that they're all going to be CEOs. Right. They're all going to be CEOs of AI agents. Right. And their, their ability to have the creativity, the will, the, the, the, the, the, the, the, and some knowledge and how to reason break problems down so that, so that you can program these AIs to help you achieve something like I do. Right. That's called running companies. Right. Now, it's, you mentioned something, this alignment, the Safe AI. Right. You mentioned the tragedy going on in the Middle East.
1:08:09You know, we have a lot of autonomy and a lot of AI that's being used in different parts of the world. So let's talk for a second about bad actors, about Safe AI, about coordination with Washington. How do you feel today? Are we on the right path? Do we have a sufficient level of coordination? You know, I think Mark Zuckerberg has said the way we beat the bad AIs is we make the good AIs better. Is how would you characterize your view of how we make sure that this is a positive net bit, benefit for humanity as opposed to, you know, leaving us in this dystopian world without purpose? The conversation about safety is really important and good.
1:08:53Yes. The abstracted view, this conceptual view of AI being a large, giant, neuro network, not so good. Right. Right. Okay. And the reason for that is because, because as we know, artificial intelligence and large language models are related, not the same. There are many things that are being done that I think are excellent. One, open sourcing models so that the entire community of researchers and every single industry and every single company can engage AI and go learn how to harness this capability for their application. Excellent. Number two, it is under celebrated the amount of technology that is dedicated to inventing AI to keep AI safe.
1:09:41Yes. AI is to carry data to carry information to train an AI. AI created to align AI synthetic data generation AI to expand the knowledge of AI, the cause of to hallucinate less. All of the AI's that are being created to, for vectorization or graphing or whatever it is, to inform an AI, guard railing AI, AI's to monitor other AI's, that the system of AI's to create safe AI is under celebrated. Right. Then we've already built. That we're building everybody all over the industry. The methodology is to red teaming the process, the model cards, the evaluation systems, the benchmarking systems, all of the harnesses that are being built at the velocity that's been built is incredible.
1:10:36I wonder if they are. Under celebrated, do you guys understand? Yes. You know, the world's the thing. And there's no, there's no government regulation saying you have to do this. Yeah. This is the actors in the space today who are building these AI's are taking seriously and coordinating around best practices with respect to these critical matters. That's right. Exactly. And so that's under celebrated, under understood. Yes. Somebody needs to to, well, everybody needs to start talking about AI as a system of AI's and system of of engineered systems, engineered systems that are that are well engineered, built from first principles, well tested, so on and so forth.
1:11:16Regulation. Remember, remember AI is a capability that can be applied.
1:11:27It's necessary to have regulation for important technologies. But it's also don't overreach to the point where some of the regulation ought to be done, most of the regulation ought to be done at the applications. The FAA, NITSA, FDA, you name it, right? All of the different, all of the different ecosystems that already regulate applications of technology. Now have to regulate the application of technology that is now infused with AI. And so I think, I think there's, don't misunderstand, don't overlook the overwhelming amount of regulation in the world that are going to have to be activated for AI.
1:12:14And don't rely on just one universal galactic, you know, AI council that's going to possibly be able to do this because there's a reason why all of these different agencies were created. There was, there was a reason why all these different regulatory bodies were created and go back to first principles again. I'd get in trouble by my partner Bill Gurley if I didn't go back to the open source point. You guys launched a very important, very large, very capable open source model. Yeah. You were trying recently. Obviously, meta is making significant contributions to open source. I find when I read Twitter, you know, you have this kind of open versus closed, a lot of, a lot of chatter about it.
1:12:59How do you feel about open source, your own open source models, ability to keep up with frontier? That would be the first question. The second question would be, is that, you know, having that open source model and also having closed source models, you know, that are powering commercial operations? Is that what you see into the future and do those two things? Does that create the healthy tension for safety? Open source versus closed source is related to safety, but not only about safety. You know, so for example, there's absolutely nothing wrong with having closed source models that are the engines of an economic model necessary to sustain innovation.
1:13:45Okay, I celebrate that wholeheartedly. Right. It is, I believe, wrong -minded to be closed versus open. It should be closed and - Last open. Yeah, right. Because open is necessary for many industries to be activated. Right now, we didn't have open source. How would all these different fields of science be able to activate be activated on AI? Right. Because they have to develop their own domain -specific AI's and and they have to develop their own using open source models create domain -specific AI's. They're related, again, not the same. Right. Just because you have an open source model, there's a mean of an AI.
1:14:25And so you have to have that open source model to enable the creation of AI's. So financial services, health, care, transportation, delusive industries, fields of science that has now been enabled as a result of open source, unbelievable. Are you seeing a lot of demand for your open source models? Our open source models, so first of all, Lama downloads, right? Obviously, yeah, Mark and the work that they've done, incredible. Off to charts. And it completely activated and engaged every single industry, every single field of science. Right. It's true. The reason why we didn't nimotron was for synthetic data generation.
1:15:05Intuitively, the idea that one AI would somehow sit there and loop and generate data to learn itself, it sounds brittle. And how many times you can go around that infinite loop, that loop, you know, questionable. However, it's kind of my mental image is kind of like, like a, you get super smart person, put him into a padded room, close the door for about a month. You know, what comes out is probably not a smarter person. So, but the idea that you could have two or three people sit around and we have different AIs, we have different distributions of knowledge, and we can go QA back and forth. All three of us can come out smarter.
1:15:51And so the idea that you can have AI models, exchanging, interacting, going back and forth, debating, reinforcement learning, synthetic data generation, for example, kind of intuitively suggests and makes sense. And so our model, nimotron 350B is, 340B is the best model in the world for reward systems. And so it is the best critique. Okay. Interesting. Yeah. And so, fantastic model for enhancing everybody else's model. Irrespective of how great somebody else's model is, I'd have recommend using nimotron 340B to enhance and make it better. And we've already seen Mates, Nama better, made all the other models better.
1:16:38Well, we're coming to the end. Thank goodness. As somebody who delivered DGX1 in 2016, it's really been an incredible journey. Your journey is unlikely and incredible at the same time. Thank you. You survived. Just surviving the early days was pretty extraordinary. You delivered the first DGX1 in 2016. We had this Cambrian moment in 2022. So I'm going to ask you the question, I often get answer, get asked, which is how long can you sustain what you're doing today? With 60 direct reports, you're everywhere. You're driving this revolution. Are you having fun? And is there something else that you would rather be doing?
1:17:39This is a question about the last hour and a half. The answer is, I great time. I great time. I couldn't imagine anything else I'd rather be doing. Let's see. I don't think it's right to leave the impression that our job is fun all the time. My job is fun all the time. Do I expect it to be fun all the time? Was that ever an expectation that it was fun all the time? I think it's important all the time. I don't take myself too seriously. I take the work very seriously. I take our responsibility very seriously. I take our contribution and our moment in time very seriously. Is that always fun? No.
1:18:28But do I always love it? Yes. Like all things. Whether as family friends, children, is it always fun? No. Do we always love it? Absolutely, deeply.
1:18:45How long can I do this? The real question is, how long can I be relevant? That piece of information, that question can only be answered with how am I going to continue to learn? I am a lot more optimistic today. I'm not saying this simply because of our topic today. I'm a lot more optimistic about my ability to say relevant and continue to learn because of AI. I use it, I don't know, but I'm sure you guys do. I use it literally every day. There's not one piece of research that I don't involve AI with. There's not one question that I even if I know the answer, I double check on it with AI. Surprisingly, the next two or three questions I ask it reveals something I didn't know.
1:19:40You pick your topic. You pick your topic. I think that AI is a tutor, AI is an assistant, AI as a partner to brainstorm with. Double check my work. You know, boy, you guys, this is completely revolutionary. I'm an information worker. My output is information. I think the contributions that all have on society is pretty extraordinary. If that's the case, if I could stay relevant like this and I can continue to make a contribution, I know that the work is important enough for me to want to continue to pursue it. My quality of life is incredible. I can't imagine you and I been at this for a few decades.
1:20:35I can't imagine missing this moment. The most consequential moment of earth's years. We're deeply grateful for the partnership. Don't miss the next 10 years. For the thought partnership, you make it smarter. I think you're really important as part of the leadership that's going to optimistically and safely lead this forward. Thank you. Really enjoyed it. Really. Thanks, Brad. Thanks, Clark. Good job.
1:21:09As a reminder to everybody, just our opinions, not investment advice.
From the publisher
Open Source bi-weekly convo w/ Bill Gurley and Brad Gerstner on all things tech, markets, investing & capitalism. This week, Jensen Huang, CEO of NVIDIA, makes a guest appearance. In Bill’s absence, Brad is joined by Clark Tang (Partner at Altimeter) as they discuss with Jensen scaling intelligence towards AGI, the acceleration of machine learning, NVIDIA's competitive advantages, the significance of inference alongside training, future market dynamics in the AI landscape, the impact of AI on various industries, the future of work, inference time reasoning, AI’s potential to enhance productivity, the balance between open source and closed source, Elon’s Memphis Supercluster, X.ai, OpenAI, the safe development of AI, & more. Enjoy another episode of BG2.
Chapters
(00:00) Introduction
(1:50) The Evolution of AGI and Personal Assistants
(06:03) NVIDIA's Competitive Moat
(15:51 ) The Future of Inference and Training in AI
(19:01) Building the AI Infrastructure
(31:35) Inventing a New Market in an AI Future
(38:40) The Impact of OpenAI
(43:25) The Future of AI Models
(46.44) X.ai and Memphis Supercluster
(51:21) Distributed Computing and Inference Scaling
(55:54) Inference Time Reasoning and Its Importance
(01:00:46) AI's Role in Growing Business and Improving Productivity
(01:08:00) Ensuring Safe AI Development
(01:12:31) The Balance of Open Source and Closed Source AI
#jensenhuang #nvidia #bradgerstner #billgurley #clarktang #xai #memphiscluster #elonmusk #noambrown #openai #gptstrawberry
