In short
Odd Lots Podcast Episode Notes
Episode Title
How to Build the Ultimate GPU Cloud to Power AI
Podcast Hosts
- Joe Weisenthal
- Tracy Alloway
Guest
- Brannin McBee - Co-founder and Chief Strategy Officer of CoreWeave
---
Episode Overview This episode focuses on the booming demand for GPU infrastructure in the artificial intelligence sector. The hosts and Brannin McBee discuss how CoreWeave provides specialized cloud computing services using GPUs, particularly in light of the growing need for AI capabilities.
Key Topics
- The current landscape of GPU demand driven by AI advancements.
- The complexities of building infrastructure to support AI workloads.
- Challenges in acquiring chips and data centers for AI applications.
- The evolution of CoreWeave from cryptocurrency mining to AI-focused cloud services.
---
Detailed Discussions
1. The Current AI Boom
- AI technologies, particularly generative AI, are getting significant investor attention.
- The rapid adoption of AI software is outpacing infrastructure developments.
- Significant demand for GPUs, particularly from companies training large AI models.
2. Challenges of Acquiring GPUs
- High demand leads to difficulties in obtaining NVIDIA chips.
- Acquisition processes are complicated due to allocation issues and overwhelming demand.
3. Infrastructure for AI Workloads
- The infrastructure for AI is fundamentally different from previous generations.
- Transitioning from CPU-based to GPU-based compute.
- Requirement for high-density power and cooling solutions.
- CoreWeave operates within tier 3 and tier 4 data centers to ensure high availability.
4. CoreWeave's Offerings
- CoreWeave specializes in GPU infrastructure for:
- AI
- Media and Entertainment
- Computational Chemistry
- The company claims to deliver services 40-60% more efficiently than traditional hyperscalers.
- The importance of leveraging NVIDIA’s ecosystem, particularly their CUDA software.
5. Market Dynamics
- The disparity in supply and demand for GPUs could persist for several quarters.
- The anticipated growth of the inference market is staggering, requiring millions of GPUs.
- Potential challenges as companies transition from AI training to delivering services.
6. The Evolution of CoreWeave
- Originally started as a cryptocurrency mining company, now pivoted to AI.
- Focused on creating a diversified offering to utilize hardware efficiently.
- Plans to maintain flexibility in workload allocation between AI and cryptocurrency as needed.
7. Competitive Landscape
- Discussion on how larger cloud providers (hyperscalers) may face challenges in adapting.
- CoreWeave’s unique positioning allows for a nimble response to market needs compared to larger organizations.
---
Key Takeaways
- AI Infrastructure Demand: There is a critical need for GPU capabilities that is not being met by current supply.
- Transition to AI: Companies that trained AI models require exponentially more GPUs for inference than was needed for training.
- Ecosystem Moat: NVIDIA's establishment of a developer-centric ecosystem (e.g., CUDA) creates a competitive advantage.
- Adaptation Challenges: Larger tech companies may struggle to pivot quickly to meet the demands for AI infrastructure.
- CoreWeave's Strategy: By focusing on GPU compute and specialized workloads, CoreWeave aims to dominate this niche market.
---
Closing Thoughts The episode highlights the urgent need for advanced GPU infrastructure in the fast-evolving AI landscape. Brannin McBee's insights into CoreWeave's operations shed light on the intersection of technology and market dynamics, illustrating not only the challenges but also the opportunities present in the AI and cloud computing sectors.
Follow-Up Listeners are encouraged to explore the implications of these developments for the broader market, including potential investment opportunities and the future of AI technology.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00You're being sold an AI future where you're obsolete or irrelevant. That vision is wrong. At Palantir, they're building AI that helps workers and unlocks their full potential. American workers are our nation's greatest strength. AI shouldn't eliminate them. It should elevate them. Palantir is here to tell their stories. From factories to hospitals, AI is freeing people from drudgery, letting them do what humans do best. Create. Solve. Build. Palantir, making Americans irreplaceable.
1:01Acrobat Studio. Learn more at adobe.com slash do that with Acrobat.
1:16Hello and welcome to another episode of the Odd Lots podcast. I'm Joe Weisenthal. And I'm Tracy Alloway. Tracy, have you looked at NVIDIA stock chart lately? And by lately, I don't mean like over the last two years. I mean like just like over the last like two weeks or two months. I don't need to look at it because everyone keeps talking about it. So I know what's happening. You know what I'm pretty happy about, can I just say? You know, we did that episode like two months ago. Yes. With Stacey Rasgen. And we were like, what's up with NVIDIA? Like, you know, I know it's at the center of the AI chips boom and whatever.
1:49And then like we did that episode and it came out. And then a week later, like they just like knocked it out of the park. The stock took off. Yeah, so, you know. We were early. We were at least like, you know, a good like two weeks early. Yeah. Hey, hey, two weeks. I'll take it. I'll take it. So clearly something that, you know, we and we talked about this with Stacey, like, you know, something that NVIDIA has is like everyone's trying to buy it. Everyone's trying to get it. But then it raises the next question of like, OK, but what is that market like? How do you buy a chip? Yeah. How do you buy a chip?
2:20And then I guess what do you actually do with it once you have it? Because my impression is that for a lot of these AI applications, the way you use the chips, the way you set up the data centers is very, very different to what we've seen in the past. And I think also what NVIDIA is doing now is kind of different. But maybe we can get into this with our guests. My impression is they're trying to create a sort of like holistic approach for customers where they provide not just the hardware, but also some services to go along with it. Yes, right. And like all the software and Stacey talked about that with the CUDA ecosystem.
2:57That was it. How dominant that is. But right. Like what do you do with it? Like how do you get one? If like what, you know, what would we do, Tracy, if a big pallet of NVIDIA chips wound up here? Would we even? Joe, you want to know a secret? Yeah. My basement is filled with H100 chips. Just got a pile of them. It came with the house. It was on that chip that was stuck on the Chesapeake. Instead of getting your couch, you got it. I just got a pallet of H100 chips. That would be, that will, we're manifesting that into reality. So anyway, I like how this world works. So essentially like the trading and dealing of these, like the hottest commodity in the world, right?
3:34Which is these advanced chips from AI and how that works and who can get one. I still think it's like a sort of mystery that we need to delve further into this question. I agree. And there is also, there's a lot of excitement around it right now for the obvious reasons of everyone's really into generative AI and NVIDIA stock is exploding, as we already talked about. But we're also seeing a lot of previous, I guess, consumers of chips, like the crypto miners, start to pivot into the space. And I'd be curious to see what they're doing in it as well and how much of that is just, you know, desperation versus a real business opportunity.
4:14And the video game market. Yeah, oh, totally. I forgot about video games. Which was like the other thing. it. It's like for years, I thought of NVIDIA as the video game company. Because they had their logo on Xbox. And how realistic is that pivot? What proportion of those types of chips can be used for AI now? Well, I'm very excited. We do have, I believe, the perfect guest. We are going to be speaking with Brandon McBee. He is the chief strategy officer and co-founder of CoreWeave, which is a specialized cloud services provider that's basically providing this sort of like high volume compute to AI type companies.
4:49They recently raised over$400 million. We've been in this space for a little while. So Brandon, thank you so much for coming on OddLots. Thanks for the opportunity, guys. Really excited to chat with you all today. So let me start with, if Tracy and I, I don't know why they would do this, but if some VC was like, we want you to do OddLots GPT. We want you to do a core base large language model off of all the work you've done. We want you to compete with OpenAI. And they gave us like, I don't know, some like, you know,$100 million raise. They said, go start, do your startup. Could I call NVIDIA and buy chips?
5:26Would I be able to like get in the door there? Gosh, I mean, you're, I think you and everyone else is asking that question and you're going to have a huge problem doing that right now. It's mostly just around how much in demand this infrastructure became, right? I mean, you could argue it's one of the most critical pieces of information technology resources on the planet right now. And suddenly everyone needs it. And, you know, I like to contextualize it in that, you know, the pace of software adoption for AI is like one of the fastest adoption curves we've ever seen, right? Like you're hitting these milestones faster than any other software platform previously.
6:06And now all of a sudden you're asking infrastructure build to keep up with that, A space that traditionally takes more time. And it's created this massive supply demand imbalance just on in-place infrastructure today. And not only infrastructure, it's available to purchase. And it's an issue that is going to be ongoing for a bit as well, we think. So can I ask the basic question, which is CoreWeave? What do you do exactly? Joe mentioned the capital raise, which I think has you valued at something like$2 billion. So congrats. But what exactly are you doing here? Yeah, thank you. So Quarive is a specialized cloud service provider that is focused on highly parallelizable workloads.
6:51So we build and operate the world's most performant GPU infrastructure at scale and predominantly serve three sectors. That's the artificial intelligence sector, the media and entertainment sector, and the computational chemistry sector. So we build, we specialize in building this infrastructure at super compute scale. It's like quite literally, you know, it's 16 ,000 GPU fabric and we can get into all the details and how complex that is. But we build that so that entities can come in and train these next generation foundation machine learning models on. And we found ourselves in a spot where we can do that better than literally anyone else in the market and do it on a timeline that's faster.
7:31or I think the only entity with H100 available to clients at scale globally today. So you have an actual basement full of H100 chips. Well, can you talk to us, you know, when you say infrastructure, we help clients build out the infrastructure. Help us conceptualize this. Yeah, visualize that. Yeah. What does the infrastructure for this type of AI actually look like? And how does it differ to infrastructure for other types of large-scale technology projects. Yeah, totally. So, you know, I think during the last NVIDIA quarterly earnings call, Jensen put this a really great way in the Q &A section.
8:12He said that we are at the first year of a decade-long modernization of the data center or like making the data center intelligent, right? You can kind of, you can suggest that the last generation or the 2010s data center was comprised of CPU compute, storage, and these things that didn't really work together that intelligently. And the way that NVIDIA has positioned itself is to make it a smart data center. That's like smart routing of data, packets of different pieces of infrastructure in there that's all focused on how do you expand the throughput and communicability of and in between pieces of infrastructure.
8:52Right. It's this amazingly different approach to data center deployments. And so the way that we're building it and we're working with NVIDIA infrastructure, we design everything to a DGX reference spec and a DGX is NVIDIA's like, how do you draw the most performance out of NVIDIA infrastructure as possible with all the ancillary components associated with it? So all this stuff is going into what's qualified as a tier three or a tier four data center. We co-locate within these things. We're not quite building in a basement. And in our past history, we certainly had time doing that. But this is within just amazing co-location sites that are operated by our partners, such as Switch.
9:35So a Tier 3, a Tier 4 site is something that's qualified based on its ability to serve workloads with an extremely high uptime. So we're talking like 99.999 % uptime rate. And that's guaranteed by its power redundancy, its internet redundancy, and its security. And then ultimately, like its connectivity to the internet backbone, right? So as it's like as a first step, you're housed within these data centers that are just critical parts of the internet infrastructure. And then from there, you start building out the servers within there. And I can go into that detail. So you mentioned, actually, I want to just get sort of define some terms.
10:20Can you just real quickly before we move on? Tier three, tier four, what do you mean by this? Yeah, so tier three, tier four, this all goes back to like the quality of the data center that you're in. It's all about the reliability and uptime that you should be able to achieve out of that data center. It's another way to qualify the services around it. It's like power, you get redundant power, right? Like multiple power services in case one goes offline, there's another one. you get redundant cooling, you get redundant internet connectivity. It's all these services that have extra fail-safes that allow for you to operate at the highest uptime and security level possible.
10:58Is higher tier better, like tier three, four? Is that better than tier one and tier two? That's correct. Okay. So quick follow-up question then, you know, we're interested in like, okay, where the rubber hits the road, the scarcity is here. Let's say Tracy miraculously opens her basement And there really is like, you know, all these pallets of these NVIDIA chips there. Is there capacity at the data centers right now? She's like, you know what? We want to co-locate with you. You guys have great power. You're pretty well connected to the internet. You have like good security guards. So that's operated 24-7.
11:31We want to set something up. Like, is there space there? Yeah, it's a fantastic question. It's an issue that didn't really pop up until really in the last eight weeks or so. It's really happening that fast? It's happening that fast, Joe. Eight weeks. Okay, so what's that? So the two-week lead time on NVIDIA was very important, Joe. That should be the takeaway. You're right. You're right. Wow. Wait, what happened? So wait, what happened 16, describe 16 weeks ago versus eight weeks ago. Sure. Even last year, right? So this is a space, the data center space, co-location space that's been fairly chronically underinvested in because the hyperscalers just built out their own data centers instead.
12:14But what's happened is the infrastructure changed. The type of compute that we're putting in these data centers, it's different than the last generation, right? So we're predominantly focused on GPU compute instead of CPU compute. And GPU compute, it's about four times more power dense than CPU compute. And that throws the data center planning into chaos, right? Because ultimately, let's say you have a 10 ,000 square foot room in the data center, right? And you have a certain amount of power. It was called 100 units of power that go into that 10 ,000 square feet. Well, because I'm four times more power dense, it means that now I take those 100 units of power, but I only require about 25 % of that data center footprint, or in other words, 2 ,500 square feet within that 10 ,000 square foot footprint.
13:02So that then leads to like, not only is the space in the data center being used inefficiently now, because you theoretically have to run more power into the data center to use that full 10 ,000 square feet due to the power density delta. But now you have cooling issues, right? Because you designed that footprint to be able to cool 10 ,000 square feet spread out across that entire area. But now you're dropping all the power into the core of the area. Sorry, I just want to back up because this is extremely interesting. So I just want to get this detail right. Just – sorry. Just to – and then move on.
13:37But the – let's say given an X amount of power at 100 units of power, what you're saying is that with this next generation of compute, it now only gets – that's now only sufficient for a quarter of the data center. In other words, to power that whole – Of the space. That space and that to then power the whole space, you really would need like 4X the power. That's accurate. Okay. Okay. But the complication really arises out of the cooling that's required from that, right? So if you imagine you could cool a 10 ,000 square foot space and you design for that, that's one thing. But now if you have to cool in a much more dense area, that's a different type of cooling requirement.
14:17And so that's led to this issue where there's only a certain subset of tier three and four data centers across the US that can are currently designed for or can quickly be designed and changed to be able to accommodate this new power density issue. So now not only like if you had all those H-100s in your basement, you might not have a place to plug them into. And that's become a pretty big problem for the industry very quickly and truly has only arisen in the last eight weeks or so. And it's going to persist for a few quarters. So you were describing the difference between CPU and GPU. you, how do you actually connect these newer types of or these different types of chips together?
15:02Because I imagine, you know, old data centers, I guess you just have a bunch of like Ethernet cables or something like that. But for this type of processing power, do you need something different? That's exactly correct, Tracy. So what we so the legacy, the generalized compute data centers are really what the hyperskillers look like. you know, Amazon, Google, Microsoft, Oracle, they predominantly use something that's called Ethernet to connect all the servers together. And the reason you use that was, you know, you don't really need to have high data throughput to connect all these servers together, right?
15:36They just need to be able to send some messages back and forth. They talk to each other about what they're working on. But they're not, you know, necessarily doing highly collaborative tasks that require moving lots of data in between each other. That's changed. So today, what people are focused on and need to build are these effectively supercomputers, right? And so we refer to the connectivity between them, the network between them as a fabric, right? It's called a network fabric. So if we're building something to help train like the next generation GPT model, Typically, clients are coming to us saying, hey, I need a 16 ,000 GPU fabric of H100.
16:17So there's about eight GPUs that go into each server. And then you have to run this connectivity between each one of those servers. But it's now done in a different way, to your point. So we're using a NVIDIA technology called InfiniBand, which has the highest data throughput to connect each of these devices together. And, you know, taking this 16 ,000 GPU cluster as an example, there's two crazy numbers in here. One is that there are 48 ,000 discrete connections that need to be made, right? Like plugging one thing in from one computer to another computer, but there's lots of switches and routers that are between there.
17:02But you need to do that 48 ,000 times. And it takes over 500 miles of fiber optic cabling to do that successfully across the 16 ,000 GPU cluster. And now again, you're doing that within a small space with a ton of power density with a ton of cooling. And it's just a completely different way to build this infrastructure. And it's just because the requirements have changed, right? We've moved into this area where we are designing next generation AI models, and it requires a completely different type of compute. And it's caught the whole sector by surprise. So much so that it's really challenging to go procure it at the hyperscalers today because they didn't specialize in building it.
17:45And that's where CoreWeave comes in is we only focus on building this type of compute for clients. It's our specialty. We hire all of our engineering around it. All of our research goes into it, and it's been a fantastic spot to be. But our goal at the end of the day is just to be able to get this infrastructure into the hands of end consumers so that they can build the amazing AI companies that everyone's looking forward to using and incorporating into enterprises and software companies.
18:26Silicon Valley is selling you a future where you're obsolete, or worse, identical. At Palantir, they're witnessing something different and revolutionary. From re-industrializing the nation's defense base, to shipyard workers building faster, and frontline workers boosting productivity, AI is transforming work across the nation. AI is not replacing American workers or flattening them into conformity. It's unleashing what makes each one irreplaceable, their judgment, their craft, their creativity. When American workers become more powerfully themselves, they own the future. Palantir, making Americans irreplaceable.
19:08Support for the show comes from Public.com. You're thoughtful about where your money goes. You've got your core holdings, some recurring crypto buys, maybe even a few strategic option plays on the side. The point is, you're engaged with your investments, and Public gets that. That's why they built an investing platform for those who take it seriously. On Public, you can put together a multi-asset portfolio for the long haul. Stocks, bonds, options, crypto, it's all there. Plus an industry-leading 3.6 % APY, high-yield cash account. Switch to the platform built for those who take investing seriously.
19:41Go to public.com slash market and earn an uncapped 1 % bonus when you transfer your portfolio. That's public.com slash market. Paid for by Public Investing. All investing involves the risk of loss, including loss of principal. Brokered services for U.S.-listed registered securities, options, and bonds in a self-directed account are offered by Public Investing, Inc., member FINRA, and SIPC. Crypto trading provided by XeroHash. Complete disclosures available at public.com slash disclosures. You know, you mentioned these special or purpose-built connections that NVIDIA is making, And this kind of leads nicely into my next question, which is what exactly is your relationship with NVIDIA?
20:21And in order to provide this type of service, you know, vast amounts of processing power that is well suited to a particular type of technology, in this case, AI. Do you have to have a really good relationship with NVIDIA to make that work? Like, do you have to have special access to H100s and other chips? It's a great question, and I'll try to offer it from NVIDIA's perspective. And it goes a little bit back to the answer I just provided as well, in that I would think from NVIDIA's seat, what's most important is empowering end users of their compute to be able to access their compute in the most performant variant possible at scale, and to be able to access it quickly.
21:08Like a new generation comes out, they want to be able to get their hands on it. And we've built CoreWeave around hitting every single one of those checkboxes. We build it at DGX reference spec, we build it at scale, and we bring it online on a timeline that's within months of a next generation chipset launch, as opposed to the more traditional legacy hyperscalers that take quarters at a time. So us being in a position to do that has enabled us fantastic access within NVIDIA. And we have a history of consistently executing on exactly what we say we'll do, right? We under-promise and over-deliver as a business.
21:48And I think that's just put us in this place where NVIDIA has the confidence in allocating infrastructure to us because they know it's going to come online. They know it's going to get to consumers faster than anyone else in the market. And they know it's going to be delivered in its most performant configuration that exists. You know, I was thinking as I listened to some of these answers, I keep having like these like imagines like, oh, you know, there's probably like some random industrial company that's like traded like, you know, on the like S &P 400 that makes some cooling fluid whose like sales are going to be up 10x.
22:24So I'm like Googling while we're talking like, oh, what is a company that makes cooling fluid? Or like, who is some company that's like really good at making these like Infinibands? Invest in HVAC. Right. Yeah. Like what are the anyway? Right. But like, right? Like, you know, there's going to be some charts that are like these like, yeah, or tertiary plays that are like 30X up. But, you know, I want to get a sense from you of, so it's really changed a lot. And I, you know, in the last several months, could we see it from NVIDIA results and what you're describing? Like, how big is the market getting?
22:55And the way I think, you know, I know like with AI, there's training and they sort of build the model and then there's inference and the inference is how they spit out the results. Can you talk a little bit about what you're seeing in terms of the growth of both of those aspects of AI? Which is bigger and which is growing faster? And how do they compare to like the size of the installed compute base that already exists? Oh, absolutely. So this is one of my favorite topics because it's just mind-blowing, the scale that's going to be needed to support AI and the scale of this infrastructure. So, okay.
23:28So today, most of the funding that's going into the AI space is for funding to train next generation foundation models, right? So when a company is raising a bunch of money, at the end of the day, most of that money is going into cloud compute to go train this next generation found model to build that intellectual property so they have this model, they can go bring it into the inference market. And what I would say is we're having a supply-demand issue, like a chip access crunch in the training phase, where in reality, the scale of the inference market is where all the demand truly is going to sit.
24:09So what I'd offer to help contextualize that is let's take, you know, there's some well-known models in the market today. Let's say there's an in-market trained model and it took about, let's say, 10 ,000 A100 or so to train. A100 is the last generation GPU, but it still applies in terms of relative scale here. So that company that used 10 ,800 to train their model, our understanding is they're going to need about a million GPUs within one to two years of launch to support the entire inference demand. So you can train the model on 10 ,000 of these chips, 10 ,000 of these chips, whatever there.
24:53And then if they're actually going to be in the market and sell something or provide some service to make it worthwhile, they're going to need a million? A million. And I think that's just within first two years of launch, Joe. Like we're talking about something that's going to continue growing afterwards. And so what does a million GPUs mean? Obviously, right? So, you know, a couple, I think it was like end of last year, all the hyperscalers combined, right? Amazon, Google, Microsoft, Oracle, you can throw a core even there. There was about, you know, 500 ,000 GPUs globally, right? Available across those platforms.
25:29I'd say at the end of this year, it'll be closer to a million or so. But that's suggesting then that one AI company with one model could consume the entire global footprint of GPUs. And now you start to think, wait, aren't there a bunch of other companies training these models in market right now? And I would say, yes, there are. So it can imply that there are in the short term the demand of several million GPUs just to support the inference market. And there's just nowhere near enough globally of this infrastructure. And it's going to be a big challenge for the market as we exit this training phase and move into the productization or really just the commercialization of these models.
Read the full transcript
26:15Like, how do you generate revenue off them? And it's something that I don't think many people truly understand, just the amount of scale and construction that needs to take place. And now you put that in the same framework of the data centers that we were talking about, right? So there's this lack of data center space. There's lack of chipset supply. Like it's going to be an issue for years that we see. So when it comes to scale, you know, you keep mentioning the hyperscalers, which is a great term. But people like Amazon, Google, I guess, Microsoft, IBM, et cetera. How quickly or what is your impression of how quickly they are able to ramp up in this space?
26:57How fast could they react to some of the trends that you've been outlining? Yeah, so I can offer what I'm seeing today. You know, the H100s started to be distributed globally to all of us, right? Like all the entities that have these, you know, kind of upper tier relationships with NVIDIA back in March, right? So we started getting them this infrastructure online in April, really scaling in May. And, you know, we have builds going on at 10 data centers across the U.S. right now. And we're delivering it to clients. The guidance that we're seeing from the hyperscalers is that they're not going to begin delivering scale access to the H100 chipset until late Q3, maybe mid Q4.
27:44And some of them are even beginning to guide into Q1. And it's all driven by the fact that this is just a different type of compute that they're building relative to last generation, right? You're no longer just running Ethernet to your point between all these devices. You're not just plugging in CPU blades. You're having to deal with like totally different data center power density and cooling requirements. You're having to build supercomputers instead with 500 miles of fiber and all these connections. It's just it's a completely different way to build the cloud. And it's it's taking them some time to catch up because you have to retrain entire organizations to do this.
28:19So, you know, as of now, I'd say the direct answer is three quarters after a chipset launch. But it's seeming it might take longer. And I think that's all going to contribute to this just kind of slower ability to to scale infrastructure than than what's being dictated by the adoption rate of AI software. And it's going to lead to this supply-demand imbalance that will just last for a while. You know, you keep mentioning, or we both keep mentioning, the H100 for obvious reasons. But do you look at other chips? Or what would happen to, you know, your own business if, for instance, a new chip was developed that could do the same thing or better than an NVIDIA H100?
29:03Like, for instance, I hear a lot of excitement about some of the stuff that AMD is developing. And I'm not a chips expert, except maybe when it comes to Fritos or Lays. But like, how big a difference would that make to you if we suddenly got a different chip manufacturer gained prominence in AI? Sure. So I'd offer kind of two broad responses. One, typically when you train a model, you're going to use the same chips for inference on that model as well, right? So GPT-4, for example, I was trained on A100s. They're predominantly going to use A100s going forward. You might fit in some kind of newer generation, hyper-efficient chips into there, but it's not like you need a, quote, a GP with more VRAM on it, right?
29:56Like you're going to need your 40 gig or your 80 gig RAM chip because that's the size of the model that you trained, right? You're not going to need like next multiple generations. You're not going to like really be able to adopt them to change the efficiency of serving that model. So what we view is that a chip's lifespan is like its first two to three years is spent training models. And then its next four to five years is spent doing inference for those models that it trained. And then within there as well, you do this thing called fine tuning, which is updating the model with new information, right?
30:33Like how do you keep a model like up to date with what's happened on Twitter or what's happened on in the media, right? You have to keep retraining it, right? And you'll use those same chips to do that. But it's your question on other chip sets. And this is something that we have a particularly interesting view into because we have like, you know, call it 650 AI clients, right? And we're having conversations with them daily to ensure that we're meeting their scaling demands. So it gives us a look into six to 12 months into the future, what type of infrastructure they expect to need. And it's overwhelmingly people still want access to NVIDIA chips.
31:11And the reason for this is something that dates back, I think it's nearly 15 years, when NVIDIA and Jensen made the decision to open source CUDA and to make this software set accessible to the machine learning community. And, you know, today, if you go to GitHub and you search a machine learning project, they just all reference CUDA drivers. And he's established this utter dominance of ecosystem around his compute within the ML space, really similar to like the x86 instruction set for CPU versus ARM, right? Like x86 is used predominantly. ARM has been trying to find its way into the space for a while now, and it's just really struggled because all the engineers and developers are used to x86, similar to how all the engineers and developers in the AI space are used to using CUDA.
32:05So it's something that, like, obviously AMD is highly incentivized to find a way into the sector, but they just don't have the ecosystem. And it's a huge moat to deal with. And kudos to NVIDIA for establishing themselves and having the patience to stick with it and to continue to support that community over the last 15 years. And it's really paying off for them in spades today. If the demand comes for that infrastructure at some point, we can run other pieces of infrastructure within our data center. But I also find that NVIDIA has such an advantage on the competition with not only its GPUs, but all of its components that support the GPUs, like the InfiniBand fabric, that it's going to be a really difficult company to displace from the market in terms of the best standard for AI infrastructure.
32:58sure. Can I ask you a question? And I'm going to, I want to ask this politely because it's not intended to be accusatory or anything like that. So I do want you to, you know, hear this. But like, when you're like talking about like hyperscalers and you're like, you know, Amazon, Google, Microsoft, and you know, kind of core wave. And it's like, okay, those are trillion dollar companies and you're a$2 billion company. Like why? Like, I don't, still don't think I like wrap my head around like, and I know like they're all like in, they're all talking about AI, et cetera. Can you still just explain to me a little bit, why aren't they just going to, frankly, steamroll you or be able to, let's put it this way, be able to, okay, maybe it'll take a few quarters to reevaluate things.
33:40But eventually, this just becomes this sort of de facto offering from these big companies that have these huge cloud budgets that must be orders of magnitude larger than yours. Yeah, yeah. I would really love to be able to have access to their cost of capital. That's for sure. So the way, look, it's the way I talk about this is we don't have a silver bullet necessarily. Right. I can't point to like a super secret piece of technology that we put inside of our servers or anything along those lines. But the way I like to broadly contextualize it is is referencing another sector. And it's that like Ford should be able to produce a Model Y.
34:23Right. Like they have the budget. They have the people. They have the decades of expertise. But in order to ask them to produce a Model Y, you would have to ask them to foundationally change the way that they produce a vehicle, all the way from research to servicing. And that entire mechanism, like it's a giant organization. Now you have to go ask that huge organization of people to change the way that they go about producing things. And I get that. But just to push back a little bit, and this is like a theme that comes up in various flavors on Odd Lots a lot, which is that like companies have internal, it's really hard to replicate sort of like tacit knowledge within a corporation.
35:05And we see that with companies that make semiconductor equipment. We see that with companies that make airplanes. We see that with real estate developers that know how to turn an office building into a condo. And so I think this is like a deep point. But, you know, they are offering AI stuff. Like I can look at Google right now, like there's cloud AI, like, and there's Azure AI and they all have their announcements. So I'm still trying to understand, like, what is it that you're offering that all the hyperscalers, they all have, they all say they have AI offerings. What is the difference between sort of like what you have and what they say is like their, you know, AI compute platforms?
35:40Absolutely. And this will really depend on how much technical detail you'd like for me to get into. But broadly, through infrastructure differentiation, like literally using different components to build our cloud, and through software differentiation, we use different pieces of software to operate and optimize our cloud, we're able to deliver a product that's about 40 to 60 % more efficient on a workload-adjusted basis than what you find across any of the hyperscalers. So in other words, if you were to take the same workload or like go do the same process at a hyperscaler on the exact same GPU compute versus CoreWeave, we're going to be 40 to 60 percent more efficient at doing that because of the way that we've configured everything relative to the hyperscalers.
36:26And it comes back to this analogy between like why Ford can't produce a Model Y. Again, like they can't. These are trillion dollar companies we're talking about. To your point, they have the budget, they have the personnel, and they certainly have the motivation to do so. But, you know, it's not just one singular thing they have to change. It's a completely different way to building their business that they would have to orchestrate. And it's what's the analogies, however many miles it takes to turn an aircraft carrier, right? Like it's going to take them a while to do that. And I think if they do get there at some point, which, you know, I don't disagree with you, they're certainly motivated to.
37:03It's going to have taken them some time, literally years to get there. And they're going to look really similar to us. And meanwhile, I've dominated market share and I've really established my product and market. And I continue and I'll continue to differentiate myself on the software side of this as well.
37:24Thank you.
37:54multi-asset portfolio for the long haul. Stocks, bonds, options, crypto, it's all there, plus an industry-leading 3.6 % APY, high-yield cash account. Switch to the platform built for those who take investing seriously. Go to public.com slash market and earn an uncapped 1 % bonus when you transfer your portfolio. That's public.com slash market. Paid for by public investing. All investing involves the risk of loss, including loss of principal. Brokerage Services for U.S.-listed registered securities, options, and bonds in a self-directed account are offered by Public Investing, Inc., member FINRA, and SIPC.
38:28Crypto trading provided by ZeroHash. Complete disclosures available at public.com slash disclosures. Introducing the all-new Adobe Acrobat Studio, now with AI-powered PDF spaces. Do more with PDFs than you ever thought possible. Need AI to turn 100 pages of market research into five insights with a click? Do that with Acrobat. Need templates for a sales proposal that'll close that deal? Do that with Acrobat. Need an AI specialist to tailor the tone of your market report to sound real smart in real time? Do that with the all-new Adobe Acrobat Studio. Learn more at adobe.com slash do that with Acrobat.
39:06Since we're on the topic of adaptation, can I ask about, you know, your own evolution as a company? Because I think I read that you started out in Ethereum mining. And at one point, I'm pretty sure crypto mining was a substantial, if not the biggest portion of your business. But you have clearly adapted or pivoted into this AI space. So what has that been like? And can you maybe describe some of the trends that you've seen over your history? Yes, absolutely. And you're right. We did start within the cryptocurrency space back in 2017 or so, and that was spawned out of just, frankly, curiosity from a group of former commodity traders.
39:51So myself, my two co-founders, we ran hedge funds, we ran family offices. So we traded in these energy markets. We were always attracted to supply-demand mechanics. But what attracted us within cryptocurrency was there's this arbitrage opportunity that was a permissionless revenue stream. I knew the cost of power. I knew what the hardware could generate in terms of revenue with using a power input. Plus, it's effectively an arbitrage. So we explored that. We had some of that infrastructure operating literally in our basements, as you said. Then that quickly turned into scaling across warehouses.
40:29And at some point in 2018, maybe late 2018, we were the largest Ethereum miner in North America. We were operating over 50 ,000 GPUs. We represented over 1 % of the Ethereum network. But during that whole time, we just kept coming back to the idea that there's no moat. There's no advantage that we could create for ourselves relative to our competitors. Right. Like, sure, you can maybe focus on power price and just kind of chase the cheapest power. But that that just felt like chasing to the bottom of the bucket. Right. You know, I think an area we could have gone into is producing your own chips.
41:09Right. Because if you produce your own chips and you run the mining equipment before anyone else has access to it, then you have an advantage for that period. But, you know, we weren't going to go design and fab our own chips. So what we kept coming back to was this GPU compute. Man, what if it could do other things? What if we could develop uncorrelated optionality into multiple high-growth markets? And those markets are where we predominantly sit today within artificial intelligence, media and entertainment, and computational chemistry. And the original thesis was, well, whenever our compute isn't being allocated into those sectors, we'll just have it mining cryptocurrency.
41:50And we'll build out this fantastic company that has 100 % utilization rate across the infrastructure because it could switch immediately from being released from an AI workload into going back into the Ethereum network. And we did get a brief glimpse of being able to operate that way in 2021 as we had our cloud live and we had AI clients in place. But Ethereum mining effectively ended during the merge in Q3 of 2022. But I'd say the other thing that we never appreciated was the utter complexity of running a CSP, forgetting about the software side of the business, which in and of itself, we spent about four years developing the software to build a modern cloud, to do infrastructure orchestration and actually be a cloud service provider.
42:41The components themselves that the sector broadly used for crypto mining were these retail grade GPUs, right? The kind of things that you plug in your desktop to go play. Right, the video game. They were like selling them on StockX. Yes, yes. It was crazy during that period to get your hands on that infrastructure for crypto mining. And all the video gamers hated the crypto people, right? Because they're like, I want to play this game. And they would line up, what is it? GameStop and the Geek Wire Shop and all that or whatever it is. And they couldn't get it because you got it. Not you, but the crypto people were getting access to the chips first and getting more value out of them so that you could bit them up.
43:22We were certainly part of the problem. That's absolutely correct. But what we found ultimately is those chips, that's not what you run enterprise-grade workloads on. That's not what's supporting the largest AI companies in the world. And starting in 2019, we stopped buying any of those chips and only focused on purchasing enterprise-grade GPU chip sets that NVIDIA has probably about 12 different SKUs that they offer, including A100 and H100 chips. and really oriented our business around it. So I don't expect to see much repurposing of this kind of older retail grade GPU equipment that was used for crypto mining because in crypto mining, you want to buy the cheapest chip that can do the thing for it, right?
44:11That can participate in crypto mining. But there's a huge difference in price between a retail plug it into your computer so you can play video games chip and an enterprise grade. You can run it 24-7. There's not going to be downtime. You're going to have a low failure rate. There's a large technology difference and there's a large pricing difference between those. And the crypto miners, you only needed the retail grade chip because if it went down for 2%, 5 % of the time for a failure rate, that's not a big deal. But the tolerance, the uptime tolerance for these enterprise grade workloads is measured on the thousandths of a percent.
44:47And it's a different type of infrastructure. So we don't expect to see the components really being reused, if at all. And then the other variable, going back to the very beginning of our conversation, are the data centers in which these are housed. So, Joe, to your point, earlier, we sit within tier three, tier four data centers, and that's basically the broad industry standard for being able to serve these kind of workloads. The crypto miners sat within tier zero, tier one data centers. And these things are like highly interruptible. They do like really interesting things like helping load balance the power markets in places like ERCOT, right?
45:27Like they'll shut down when power prices go too high and it load balances the grid. But enterprise AI workloads don't have a tolerance for that. Their tolerance, again, is measured on the thousandths of a percentage in terms of uptime. So not only does the infrastructure not work from crypto mining, but the data centers that they built within don't work either the way that they're currently configured. Now, they could potentially convert their sites into tier three and tier four data centers. I'll tell you that in and of itself, that is an extremely challenging task and it takes a lot of proprietary knowledge and industry expertise to do so.
46:07It's not just throwing a few fans in a room and a few air conditioning units. It's a – it's honestly, it feels like walking to a spaceship. Tracy, this is an episode. I don't know about you, Tracy. There's like six follow-on episodes. It's like how do you – no, seriously, like the whole data center market and the coolant and all the electricity. There's so many different rabbit holes you could go down just like with the infrastructure you're talking about. For sure. And I think the estimates that I've seen on repurposing crypto GPUs, I think I've seen like 5 to 15 percent. So to Brandon's point, but I'm sure there will be people out there who try.
46:49You've got to try, right? Because what if it works? Right. If you can make that work, that's amazing. But we're just coming as an entity that was an extremely large operator of that infrastructure and has built one of the largest cloud service providers for AI workloads. I can tell you it's going to be really, really hard to do it because we've had exposure in both of those places. And at the end of the day, they're just very, very different businesses, both from the type of engineering and developers that you employ to the infrastructure to the data centers that you sit within. So can I just go back, you know, just sort of like big picture.
47:26And I guess it sort of goes back to like who gets access to what? Who gets access to chips? And I imagine that, you know, not only do you need a lot of money to like build a relationship with like NVIDIA, you also probably need like a, you know, expectation you're going to be back the next year and back the next year and back the next year and that you actually like have a relationship and so forth. But I have to imagine like planning is really tough. And when like, you know, you have this sort of like AI machine language, whatever, like industry, and then something like JetGPT comes out and like suddenly everyone is like, oh, I need to like have AI access.
48:02Talk to us about like this sort of like challenge of just sort of like planning the build when it can move that fast and like everyone is just sort of guessing how big this market is going to be in two to three years. Oh, my gosh. It's been utterly insane, right? Like, you know, back to last year, you know, the supply chain and ability to get your hands on components, you know, you would call your OEM. The OEM is the original equipment manufacturer. Like, those are the super micros, the gigabytes of the world who actually, you know, build the nodes, build the servers. And you're buying through them.
48:35And then they buy the GPUs from NVIDIA and build all the components together, right? So if you called them and said, hey, I need this many nodes to be delivered, they'll say, great. You know, we'll start assembling. It takes us a week to two weeks to get the parts in assembling, and then it's another week for them to ship them to you. And then it takes us two to three weeks to plug them in, put them online, get them going. Now, that's completely changed. As you know, all the supply chain has gotten thrown off, so much so that NVIDIA is fully allocated. They've fully sold out their infrastructure through the end of the year.
49:13You can't call the OEM and just say you need more compute chips. Like that's not possible. It's so much so that, you know, when clients are coming to us today and they're asking for, you know, like a 4000 GPU cluster to be built for them, we're telling them Q1. And increasingly it's moving towards Q2 at this point because Q1 is starting to get booked up right now. So it's something that a lot of time has been added to it. And then there's other supply chain variables within there as well. We had a client earlier this year that we were in negotiations with him on the contract, and we really wanted to perform well on timing for it.
49:51So we knew because of our orientation within the supply chain that there were some critical components that needed to be ordered ahead of time so that it would reduce our time to bring in the infrastructure online. And at that point, it was the power supply units and the fans for the nodes that the OEMs were putting together. And if we hadn't have done that, it would have been another, I think, eight weeks on top of the build process, just because not all the components would have been there at the same time. So you're navigating this within other kind of global supply chain disruptions and inflation and all these other things that are going on right now.
50:31And it's just an insanely complex task that I think, you know, the generation of software developers and founders that we're working with today were used to being able to go to a cloud service provider and just getting whatever infrastructure they needed. Right. You go to your hyperscalers and say, all right, I need this. And it was just there and available. And that just doesn't exist today because of the pace of demand growth that we've been on and just the lack of this infrastructure's availability. And it's just caught everyone by surprise. Again, you're asking infrastructure to keep pace with the fastest adoption of a new piece of software that's ever occurred.
51:15Brandon McBee, Corleave thank you so much, that was a great conversation like I said, I always sort of measure the quality of a conversation do I get 7 ideas that is a pretty good proxy for a good conversation, do you get like 8 ideas for future episodes, we got a bunch there so thank you so much for coming on the podcast always happy to chat with you guys and thank you for the invite thanks Brandon
51:50Tracy, I want to find that company that makes the coolant for the data. No, seriously, for the data centers that allows them to pack more compute and more energy into this space. Because it feels like they're probably going to make a fortune in the next few years. Joe, I think you just want to talk to an HVAC contractor that's installing air conditioners. Can we talk to an—it's just some random—I love the idea. Maybe it would have been such a funny thought. like these like really advanced data centers like oh do we have like a local air conditioning guy who can like come in i imagine actually that would have been a good question for brandon wouldn't it like the the labor constraints yeah in in building and adapting some of those data centers but there was so much in there one of the things one of the things that i was thinking about was the point about how well okay if you train a model on one type of chip you're going to keep using that type And I guess it's kind of obvious, but it does suggest that there's some stickiness there.
52:45Like if you start out using an NVIDIA H100, you're going to keep using them. And in fact, you're going to consume even more because the processing power required, the compute required for the inference is higher than for the actual initial training. I knew that that was the case because Stacey said so as well, but I did not realize quite the scale of how much more – like, OK, if you train a model and then we try to take it to market – productize it, as a business person might say. If we try to productize – like how much more computing power we would need for the inference aspect. And meanwhile, we have to keep training it all the time to keep it up with fresh data and stuff like that.
53:23Yeah, totally. And the other thing that I was thinking about, and again, Stacey mentioned this in our discussion with him as well, but this idea of NVIDIA building a kind of large ecosystem around the hardware. So you have the open source software, CUDA, which we talked about a little bit. And then you have these sort of high touch partnerships with companies like CoreWeave, where they're trying to make it as easy as possible for you to use their chips and set them up in a way that works for you. It feels almost like what Bitmain used to do. Do you remember that? No. Maybe they're still doing it.
54:05Anyway, but it does feel like they're trying to build this ecosystem moat around the chip technology. Yeah, no, it's absolutely true. And I really do take that point that Brandon made about every company has a sort of knowledge that cannot be written down on a piece of paper, which is a Dan Wong point that we've been talking about for years. And so it's like, to your point, you know, like you have to like use different types of connectors and different types of power and all these stuff, like the ease with which any sort of traditional cloud provider or data center provider can, you know, sort of switch to, it's like, you know, it's not trivial even with lots of money.
54:43No, but I'm coming away from that conversation thinking like the big question here is how quickly can those other hyperscalers adapt? Yeah. And like how big a moat can NVIDIA build around this business? And then, I mean, the other question I have is like, what if none of these companies make any money building AI models? Like I still don't think like that's been proven. And so you have this like huge boom and like, hey, we got to build an AI model. We're going to build like, you know, outlawed GPT for like data stuff and whatever. But it all is somewhat predicated on these companies being successful and making a lot of money.
55:18And if they're not, and if it turns out that the monetization of AI products is trickier than expected, then that also raises a question about how long this boom lasts. I'm sorry, Joe. So you're saying that tech companies should make money. Is that it? Are you sure? That's right. It's real post-Zerp thinking of me. I know. All right. Shall we leave it there? Let's leave it there. This has been another episode of the Odd Lots podcast. I'm Tracy Allaway. You can follow me on Twitter at Tracy Allaway. And I'm Joe Weisenthal. You can follow me on Twitter at The Stalwart. Follow our guest, Brandon McBee.
55:52He's at Brandon McBee. Follow our producers, Carmen Rodriguez at Carmen Armin and Dashiell Bennett at Dashbot. And check out all of the Bloomberg podcasts under the handle at podcasts. And for more Odd Lots content, go to Bloomberg.com slash Odd Lots, where we have transcripts, a blog and a newsletter that comes out each Friday. And check out our Discord. We have an AI channel and a semiconductor channel in there. So people talk about these topics 24-7. Maybe they'll be talking about them in both of those rooms when this comes out. Discord.gg. And if you enjoy OddLots, if you appreciate conversations like the one we just had with Brandon McBee, then please leave us a positive review on your favorite podcast platform.
56:36Thanks for listening. Thank you.
57:37We'll see you next time. BlackRock Investments, LLC. With the B2B card payment landscape evolving, large corporations face pressure as buyers increasingly demand to pay invoices by virtual card. For merchant acquiring businesses like yours, this is a high growth opportunity waiting to be unlocked. With MasterCard's adaptive approach to B2B acceptance, you can enhance your infrastructure for high value payments and meet your customers' unique needs. MasterCard offers solutions and support for every step of the supplier lifecycle, helping you deepen merchant relationships. Start fast, grow strategically, and scale at your pace with a modular toolkit you can flexibly deploy.
58:16Discover how at MasterCard.com slash commercial acceptance.
From the publisher
Artificial Intelligence is all the rage right now and most of the investor excitement has so far been focused on the companies providing the hardware and computing power to actually run this new technology. So how does it all work and what does it actually take to run these complex models? On this episode, we speak with Brannin McBee, co-founder of CoreWeave, which provides cloud computing services based on GPUs, the type of chips pioneered by Nvidia and which have now become immensely popular for generative AI. He walks us through the infrastructure involved in powering AI, how difficult it is to get chips right now, who has them, and how the landscape might change in the future.
See omnystudio.com/listener for privacy information.
