CoreWeave’s Brannin McBee on the future of AI infrastructure, GPU economics, & data centers | E1925

4 Apr 2024 · 56 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: This Week in Startups - CoreWeave’s Brannin McBee on AI Infrastructure

Episode Information

  • Title: CoreWeave’s Brannin McBee on the future of AI infrastructure, GPU economics, & data centers
  • Host: Jason Calacanis
  • Guest: Brannin McBee, CDO and Co-founder of CoreWeave
  • Episode Number: E1925
  • Air Date: [Insert Date Here]

Episode Overview In this episode, Jason Calacanis interviews Brannin McBee, discussing the evolution of CoreWeave from a cryptocurrency-focused company to a major player in AI infrastructure. The conversation covers GPU economics, energy dynamics, innovations in data center cooling, and the future impact of AI infrastructure on various industries.

Key Highlights Transition from Cryptocurrency to AI

  • CoreWeave initially operated in the cryptocurrency space, renting GPUs for Ethereum mining.
  • The company pivoted to focus on AI infrastructure around 2019, anticipating the rapid growth of AI workloads.
  • Brannin emphasizes the importance of building a cloud infrastructure specifically designed for AI, which differs significantly from traditional cloud services.

GPU Economics and Energy Dynamics

  • GPUs are increasingly critical for AI training and inference workloads.
  • Energy consumption is a major factor in costs, with approximately 10% of operational costs attributed to power.
  • Brannin discusses the future demand for data center capacity and the challenge of meeting growing energy needs.

Innovations in Data Center Cooling

  • A shift towards liquid cooling technologies is anticipated to improve efficiency in managing heat produced by GPUs.
  • Direct-to-chip liquid cooling is preferred for operational efficiency compared to immersion cooling.

Market Demand for AI Infrastructure

  • Brannin notes an unyielding demand for GPU capacity, with revenue at CoreWeave projected to increase significantly.
  • Inference—the process of applying AI to real-time scenarios—will drive future demand for infrastructure as user bases grow.

Future Challenges

  • The conversation touches on potential bottlenecks in sourcing enough data center capacity and power for the burgeoning AI market.
  • The need for a new approach to infrastructure is highlighted, as existing cloud solutions are not designed for the parallel workloads associated with AI.

Discussion Points The Current State of AI Infrastructure

  • Current cloud infrastructure was built for serial workloads, necessitating a complete overhaul to cater to AI demands.
  • CoreWeave’s unique engineering solutions set it apart from traditional hyperscalers.

Competitive Landscape

  • NVIDIA holds a dominant position in the GPU market, but there is speculation about future competitors like AMD.
  • The interview explores the implications of NVIDIA's software ecosystem on market dynamics and adoption.

Potential Use Cases for AI

  • The integration of AI into existing products and services will facilitate rapid adoption.
  • Brannin predicts AI’s impact on sectors like advertising, where personalized ads could revolutionize marketing strategies.

Key Takeaways

  • AI Infrastructure Demand: The demand for AI infrastructure is set to grow exponentially, driven by increasing user engagement and the need for real-time data processing.
  • GPU Economics: Understanding GPU cost dynamics and energy usage will be critical for startups and companies leveraging AI.
  • Innovation in Cooling: Liquid cooling technology is poised to enhance efficiency in data centers, addressing heat management challenges associated with high-performance GPUs.
  • Market Outlook: A substantial investment in infrastructure will be necessary to keep pace with AI advancements, with projections suggesting a supply-demand balance may not be achieved until the end of the decade.

Conclusion The conversation with Brannin McBee provides valuable insights into the evolving landscape of AI infrastructure, highlighting the challenges and opportunities that lie ahead in this rapidly advancing field. With a focus on innovative solutions and energy efficiency, CoreWeave is well-positioned to be a key player in the future of AI technology.

Additional Resources

  • CoreWeave Website: [CoreWeave](https://coreweave.com)
  • Follow Brannin McBee on Social Media:
  • [Twitter](https://twitter.com/branninmcbee)
  • [LinkedIn](https://www.linkedin.com/in/branninmcbee)
  • Follow Jason Calacanis:
  • [Twitter](https://twitter.com/Jason)
  • [LinkedIn](https://www.linkedin.com/in/jasoncalacanis)

Sponsors

  • OpenPhone: [Get 20% off your first six months](http://www.openphone.com/twist)
  • Gusto: [Get three months free](http://gusto.com/twist)
  • Northwest Registered Agent: [Form your business for $39](https://www.northwestregisteredagent.com/twist)

---

This markdown summary encapsulates the essence of the podcast episode, providing an organized structure for readers to grasp the key discussions and insights shared by Brannin McBee and Jason Calacanis.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00You look at the existing cloud infrastructure that was built over the last decade, it was built for serializable workloads. It wasn't for parallelizable workloads. And it's like you're having to rebuild the cloud, so to say, and you're having to rebuild physical infrastructure at the pace of AI software adoption. It's a mind-blowing concept, right? Because AI software is being adopted at the most rapid scale of any technology that we've ever observed. I mean, we're building at, I think it's 28 data centers this year across North America. We're one of the largest operators of this infrastructure in the world.

0:34And we are unable to keep up with demand. And we really don't see that subsiding for years to come. This Week in Startups is brought to you by OpenPhone. Create business phone numbers for you and your team that work through an app on your smartphone or desktop. Twist listeners can get an extra 20 % off any plan for your first six months at openphone.com slash twist. Gusto is easy online payroll benefits and HR built for modern small businesses. Get three months free when you run your first payroll at gusto.com slash twist and Northwest Registered Agent. When starting your business, it's important to use a service that will actually help you.

1:24Northwest Registered Agent is that service. They'll form your company fast, give you the documents you need to open a business bank account and even provide you with mail scanning and a business address to keep your personal privacy intact. Visit northwestregisteredagent.com slash twist to get a 60 % discount on your next LLC. All right, everybody, welcome back to this week in startups. We've got a great guest for you today. You may have been wondering who's buying all of these nvidia h100s and how are people getting access to all of this hardware well there's a couple of companies that got to the hosting of ai and gpus early one of those companies is core weave they started in the space um i believe doing a lot of crypto where miners were renting gpus from their cluster and uh fortune favors the bold they were in the catbird seat when the ai revolution happened and everybody decided well they got to train their own models going to need a bunch of nvidia's hardware and other people's hardware and we'll talk about that today and they've since grown uh to a massive scale um for those of you who don't know nvidia's market cap has increased more than 7x from 300 billion to 2.3 trillion if you've been living under a buck and you haven't been watching this it's because people want access to these chips welcome to the program brannon mcb who is uh the cdo and co-founder what does cdo stand for chief development officer um so so my role is raising capital for the business i interface with um equity and debt participants for the company and help fuel the growth of the business And this is a very capital intensive business.

3:15You spent a lot of money on GPUs and setting up infrastructure. The company's been around for just under a decade. Am I correct? Yes. We founded the company in 2018. Got it. Am I also correct that you were supplying GPUs largely to crypto and Bitcoin miners and this cohort of individuals? Or that was the beachhead market? It was absolutely the beachhead market. So I can take a couple of steps back on our founding story. So it's myself, my two co-founders. We're from the institutional commodity trading sector. So we're risk managers by background. We're from hedge funds. Finance guys is probably the best way to look at it.

3:57But we were finance guys who were heavily data oriented. We worked in this commodity sector that you can actually solve for price. You can figure out supply demand. And there is a dollar per barrel of oil, so to say, that that solves market dynamics. So we've always worked with a lot of compute. We've worked with a lot of software. And the crypto space was interesting because it was a arbitrage opportunity. Right. There's there's a very discrete input price cost of power. And you could model the revenue very efficiently because there was no customers. Right. You were just participating in this this network.

4:37And thus, if you were just to sell the revenue from the crypto mining proceeds every day, you could effectively qualify it as an arbitrage opportunity. And that was interesting to us. But it wasn't as compelling as a large business because at the end of the day, all you're going to do is chase the price of power lower. right like that's that's the only advantage you can really extract unless you expand into other markets right so when we're looking at the cryptocurrency space it wasn't bitcoin mining that we were interested in it was ethereum mining and gpu oriented mining because a bitcoin miner basic uh things that are produced by entities like like that man they can only do that one thing They can only participate in Bitcoin, and they're very good at it.

5:31But a GPU, well, it can do lots of things, including running AI workloads. So we started in the crypto space, but it was always with this idea. And we had no idea how complicated an idea was at the time, but started with the idea that, well, you could do crypto and other things. Right. Started there. The other use of these is, I guess, running video games in the cloud. Is that correct? yes cloud video games become a real market or do people who are into video games just buy themselves an alienware dell whatever and be done with it i think it's more the latter um that they they use it more for that we certainly don't see that demand for video game streaming um i think a few of the hyperscalers tried to launch into that market and i don't believe that there's been a substantial demand.

6:24Got it. And so when we look at crypto, that market kind of fizzled right as AI was starting to boom. So you were able to sort of just navigate that? Or is the crypto still going on and people are still using your services for Ethereum and being part of that network? Or is that just too hard of an arbitrage now because people in China have stolen electricity, we hear, or their friend runs the hydro dam, so they run an extension cord, so to speak, over to their warehouse with a bunch of servers in it. And you're up against people getting zero cost of input electricity. It's a great question. And yes, I'm extremely excited that we're not involved in that cryptocurrency market anymore.

7:06We haven't been involved for a number of years at this point. We actually started making this transition into the cloud infrastructure market in 2019. We hired Peter Selenke, who was recently elevated to our chief technology officer position in 2019 to build a cloud for us. And it's not just plugging in GPUs and having users come access it. It's this really complicated software stack that runs the cloud. Or in other words, it's an orchestration environment that enables users to access and use our infrastructure. And it's one that it's very different than the way that the hyperscalers built it because they built for hosting websites and storing data lakes.

7:53And we built our cloud from a no compromises engineering solution for running AI workloads and highly parallelizable workloads. And there's engineering decisions you make in doing that that you wouldn't make for hosting websites. and that's allowed for us to, I would say, overperform in the market with a product that really doesn't have competition. We started in 2019. Yeah, the software layer to provision these H100s, A100s, whatever people are using, that's a key part of the puzzle that you have to build, you have to master. And AWSs, Google's cloud, everybody's Azure, they're all slightly different in using their own provisioning software or is there some open source standard there for doing all that?

8:41That's exactly correct. We use and contribute to open source as much as we can, but we have a proprietary orchestration solution that looks different than the hyperscalers do. My favorite analogy for this actually comes out of the automobile sector, where at the end of the day, everyone produces vehicles the same way, right from research design scaling servicing it's the same sort of product with different badges and different colors on it right and it's been that way for 60 plus years and then in the 2000s a company came along and said well what if we started with a blank slate and designed this process today and you know ultimately uh ford might have to produce vehicles like tesla does But I think we can all appreciate the foundational difference in the way that those vehicles have brought to market and the challenges that Ford will have to go through to get there.

9:42And I think it's a lot of the same for the hyperscalers, right? I'm not going to tell you a trillion dollar company with tens of thousands of engineers can't do what we do. But I will highlight the innovators dilemma that sits there because there's an existing product. You have to change everything underneath to run infrastructure like we do, and it's going to be hard to get there. And if we go through the economics of, let's just say, NH100, this is NVIDIA's state-of-the-art. It's not just a GPU, it's a rack, essentially. It's a platform to put many GPUs on. I'm not sure exactly how many the H100 holds, but it holds a number of GPUs.

10:25they go for like 30 40 grand each is my understanding so it's a um it's a server or node as we call it within a network fabric and a server has typically eight gpus within it um and then you put those into a cabinet and you put those into a data center and you bring power into the data center and you connect it to the internet right and then yes those things um that that price range was accurate for each one of those gpus in a server so a server can cost upwards of a quarter million dollars and so you rent out one of those h100 gpus an individual gpu in a server like that a cluster a node for four bucks an hour something to that effect yeah yeah yeah that that's right um where we specialize though is doing it at scale yeah right like we don't have many clients who just use one at a time.

11:22Our clients will use 10 ,000 at a time in a single contiguous fabric, which makes it a supercomputer. And it's interesting, this actually has become where we operate some of the world's largest supercomputers at this point. I think several of the top 10 now sit on our platform because of how large these fabrics are and how performant these GPUs are at these specific tasks. tasks so somebody fires up 10 000 of those i'm assuming they get some kind of volume discount so if it was two or three dollars an hour they're spending twenty thirty thousand dollars an hour on a job at one of those correct something in that range yes but i i will correct the the discounts it actually works inverse right because it's uh it's extremely surprising but it's because building a a single fabric of this size is so engineering intensive that not many people are able to do it.

12:21There's not a template. Not a lot of companies have gone out and done it. There's actually maybe three or four on the planet who are actually building fabrics of this scale. So you actually make a more scarce resource through scaling. You're kind of like decommoditizing the market through scale. It's not one GPU or 10 ,000 GPUs. It's, oh, 10 ,000 GPUs. That's a totally different engineering solution. Juggling multiple devices and apps to run your business is a mess. Open Phone is here to make it simple by simplifying your business communications with one easy to use app. Open Phone has rethought every detail of what a modern business phone should be.

13:04And here's the magic. It works through a beautiful, elegant app on your phone, or you can just use it on your desktop, making it super easy to get a business phone number for your entire team. And And you know how brilliant OpenPhone is? My teams use it every single day. My sales team loves it. My ops team, they use it all day long. And here's the features that we love. You can create a shared phone number, like customer support, with multiple employees fielding all the calls and all the texts to that one number. At my investment firm launch, we pride ourselves on replying to every single call or email instantly.

13:37And OpenPhone is the number one rated business phone on G2 for customer satisfaction. So here's your call to action. Super easy. OpenPhone. is already affordable. Starts at just 13 bucks a month, but Twist listeners get an extra 20 % off any plan for the first six months at openphone.com slash twist. And if you have existing numbers with other services, no problem. Open Phone is going to port them over easy peasy, lemon squeezy, no extra cost. Head over to openphone.com slash twist to start your free trial and get 20 % off. And how much of that cost, when we look at a$4 an hour cost, would you say energy is?

14:10Because these things are extremely energy reliant, right? They consume a lot of energy. So I'm curious how much of all this is energy. And then where do you put your, this will take us down the energy rabbit hole, but where do you put your data centers? And then where are data centers going to be in the future? Because these GPUs are taking a multiple of what CPUs would use. Yeah. So maybe you could explain that to us. Yeah. Yeah. It actually causes a pretty substantial bottleneck that exists in the market now is data center capacity. Right. And it's not square footage. of data centers. It's data centers that have enough power brought into them and adding more power to those data centers.

14:51And arguably, that's where the next bottleneck in this cloud infrastructure or GPU cloud infrastructure market sits, is how do we access enough data center space to accommodate the volume of demand that's coming in? So power, it's roughly about 10 % of our cost to deliver this infrastructure. The infrastructure itself is actually where most of the cost sits from a depreciation perspective. Depreciate over a six-year life on the infrastructure. The power side is... That's my background. It was in power markets and trading different electricity markets. So it consumes a lot of power. It has an immense amount of efficiency over CPU infrastructure.

15:38as well for what it's doing there. To run the same workloads on GPU versus CPU, it's actually more power efficient to run it on a GPU, right? If you're trying to achieve the same outcome, because you'd have to use so many CPU cores to get to the same solution. So yes, they consume more power on a density basis, but on a workload basis, they're more efficient. And this is, I mean, it's staggering. I was looking at one study that said, Like one of these GPUs at 60, 70 % capacity year is like the average American households, energy consumption. And that's just one of them. So this would be the equivalent of like, if somebody is using 10 ,000 of these, or I think Zuckerberg's going to be using low millions of these.

16:24It's like putting a million households online or something to that effect. Yeah. So. Yes, it's an immense amount, but it's also a, you know, it's a transformational technology. methodology i'd say that we're looking at and it's and its ability to unlock value from data is something that we've never observed before um it going to yeah so that speaks to justifiable like yeah i mean i'm not even i'm not even looking at this judgmentally like is it worth the amount of energy it's consuming i was looking pragmatically where let's assume it is where it's going to come from yeah let's say it's going to cure cancer it's it's going to find uh solutions for renewables or fusion that like we didn't even conceive of or the gains from it will be so extraordinary.

17:09It will obviously pay for itself and create an energy independent future. But what's happening in the industry today as people are buying these and looking for places to store them, you're looking to build up your infrastructure. Are we just out of energy? And where are people where are the nooks and crannies where people are looking to locate these facilities? I going to become like a place where people put these uh uh the plan is to put a nuclear power plant and these gpu data centers next to each other yeah is there any truth to that yeah yeah look i i believe that was microsoft or amazon is is effectively taking that nuclear plant and citing a data center next to it power i think you know beyond data center space right it is a national concern, so to say.

17:58There's been an immense amount building out renewable capacity over the past decade, which is fantastic, but it's also not necessarily the right kind of capacity, what you need for consistent demand growth. As you know, solar works when the sun's out, wind works when the wind blows. Neither of those things work for a data center or even necessarily for electric vehicles, for all these kind of demand areas. We need more baseload power. That's traditionally come from coal and natural gas. Fortunately, it's been more so from natural gas over the last decade because coal is quite dirty from an emissions perspective.

18:40And my personal hope is that it's more nuclear going forward, but it takes time to build nuclear sites. I think it's a decade for siting and build. In the United States, and we haven't built one in a long time. I mean, I think the last one broke ground in the late 60s or early 70s, and we haven't had one since. So this would lead one to believe that somebody with a lot of nuclear power and a lot of GPUs would have a massive advantage. It certainly helps. We've contracted a substantial amount of capacity for looking to ensure the growth profile of our business, but it is going to be a bottleneck for all other participants in the market.

19:26what about heat uh these things throw off a lot of heat um and you know some areas in the country are warmer than others is it are people moving these data centers north in order to get the cold air to just you know we've we've seen pictures of you know data centers that have open sides where cold air just blows right in or you know open doors essentially um because in other places if you were to put these gpus in texas i think you're going to be air conditioning them which seems doubly inefficient. So maybe talk a little bit about the heat these things generate today and if there's any hope of cooling them down without air conditioning.

20:04Yeah, so it's a great question. It's funny, like it takes me vividly and visually back to my crypto mining days where we did run those warehouses with the open sides and the giant fans and we were up north, could never run this infrastructure in those environments, right? From a security perspective, from a reliability perspective, it is mandatory to run this infrastructure in what's called a tier four data center environment, or sometimes even a tier five data center. And that's the highest classification in terms of reliability, redundancy, security, and environmental handling. So these are sites that you would see like Amazon or Google or Microsoft running within, Like true data centers that are meant for cloud infrastructure.

20:54And the way to think about the heat output, the other variable in there, which is wild, is actually the sound. They're extremely loud in these environments, upwards of 100 decibels, which is a direct derivative of the heat they're consuming. And then to move the air. But the way you look at it is the critical load around the infrastructure. So if it takes one unit of energy to run the infrastructure, it takes another 0.2 to 0.3 units of energy to cool the infrastructure and run the networking and everything else around it. The way that the world's leading sites handle this is just through forced air.

21:35Just move tens of thousands of cubic feet per minute of air through these highly contained pods. So you have all the hot air in a really small area and you're just jamming air through it. Right. And sometimes it's conditioned. Sometimes it's just air. But eventually it's going to be liquid. Right. We're going to move to an environment where you have direct to chip liquid cooling instead. And that efficiency ratio, call it 1.3, will drop to about 1.1 instead. So your GPU infrastructure will inherently become more energy efficient as we move to a liquid cooled environment. And we're working with leading data center operators, such as switch to, to facilitate and implement that movement for these upcoming generations of GPUs.

22:28when people say liquid cooled um most people have not actually physically seen that yeah unless maybe you're a gamer and you've seen your chip in a a tube run to the chip and there's literally liquid on the top of the chip that's cooling it um how do these uh when you say liquid cooled what could people envision uh of how these solutions are going to work are they going to just be like a bunch of racks in a in a in an olympic-sized swimming pool or is it just like little contained amounts of water on top of the gpus yeah so you're qualifying that correct there's two broad categories of liquid cooling there's immersion cooling which is the olympic swimming pool method downsized obviously and then there's direct-to-chip liquid cooling which is running the pipes to the chips um we will sit on the direct-to-chip liquid cooling side because it's uh operationally more efficient for us um that that's where we think that the sector is broadly going to go.

23:27If you think of immersion cooling, you're literally dunking a server into a vat of liquid. That liquid has its own problems with it as well. But let's say you had to go service that server. There's a node or a component that was wrong with it. Well, you got to lift it out of the liquid. And what happens then? Well, you got to wait probably an hour for all the liquid to drain out now before the tech can even get into it so you're extending these service times yeah materially and response times versus direct the chip liquid cooling uh you know pop it out and you don't have water containment issues you know things splashing across data center like it's it's a mess these are highly sterile and contained environments that even let us bring cardboard inside of the data center area because it's combustible right and you can have little particulates that float around the site and can accrete into the nodes um bad you don't want fires and data centers i mean when you are talking about that much air being pushed around that means any particulates in the air are going to get pushed around and so if you just had some very small amount of particulates floating around a room now imagine that room is changing the air every x amount of time the number of particulates is going to grow and then you're going to have a small fire on a chip which is just absolutely crazy do you think the demand is going to keep up are you seeing any signs of demand people saying okay we we built our 10 000 gpus we're making more efficient software making more efficient use of the chips okay yeah we're getting to a steady state we've bought enough so are you starting to see that with your customers saying you know what we've got enough gpus we got enough infrastructure right now or are they still in the begging, pleading, and doubling?

25:19What are you seeing from top customers? Are they doubling their capacity every year? Are they tripling? What's the field report? Yeah, I'll qualify it a couple of ways. So one way, we will increase our revenue by about tenfold this year. And we're already sold out of all of our capacity through the end of the year. So I have a build schedule. We have about 500 employees today. I'll be closer to 800 by the end of this year. That build schedule is fully booked this year already. We see that broadly across the sector. There's just an immovable wall of demand for this compute. A lot of it is being driven from this move from training the models to inference.

26:04And inference is actually bringing the commercial value out of training. So you want to go train a foundation model that takes compute to be built in the configuration that we build it in these 10 ,000, 30 ,000 GPU clusters. And then you got to go make it actionable, drive revenue off it and bring a product. And what we're observing is it might take 10 ,000 GPUs to train a model, but inference is linked to the number of users. If you go into chat GPT, for example, and query, that's spinning up a GPU. And now there's a million of you doing it, 5 million, 10 million. That informs the size of inference.

26:48So inference will really, truly be linked to the growth of this market. And we're seeing users who are using 10 ,000 GPUs for training need hundreds of thousands for their early stage inference products. So we don't see demand going anywhere but up to the right for this infrastructure. So while they may not need exponential use of, you know, GPUs for training, they'll get more and more efficient at that. And what they will need is those inference when people ask the query, that's inference, not training the model, but asking a question of the model. That is massively compute intensive. and in h100 if we were to look at that unit in an hour at full capacity how many queries you know i know it depends is always the answer but an average query like these things were costing a couple of pennies per query is that correct ballpark yes yeah that that's correct and i think that's the right right way to qualify it is is uh cents per query or dollars per query So you're getting in hundreds of queries within that period.

28:04And as you said, that will become more efficient over time as well. But it's just an unbelievable volume of demand. And when you step back and think about it, you look at the existing cloud infrastructure that was built over the last decade. It wasn't built for this use case. It was built for serializable workloads. It wasn't built for parallelizable workloads. And it's like you're having to rebuild the cloud, so to say, and you're having to rebuild it. You're having to rebuild physical infrastructure at the pace of AI software adoption. It's a mind-blowing concept, right? Because AI software is being adopted at the most rapid scale of any technology that we've ever observed.

28:46And you're asking people to build. I mean, we're building at, I think it's 28 data centers this year across North America. We're one of the largest operators of this infrastructure in the world. And we are unable to keep up with demand. And we really don't see that subsiding for years to come. Listen, as a founder, there are things I love doing, like building products or meeting with partners. hanging out with my team and dreaming up new ideas. And then there are chores that I don't want to do. I don't want to do HR. I don't want to do payroll. I don't want to deal with all that. So I use Gusto.

29:21Gusto is the best for payroll, for HR services, and for running a small business. It makes everything so much easier. Even a midsize business, man. I get a lot of portfolio companies that are pretty sizable using Gusto because it is designed for you, the small business owner. And payroll is something you definitely do not want to mess up. You got to get it right. And Gusto is going to make it perfect for you by calculating paychecks perfectly. Also payroll taxes, you got to get your taxes right. You can't make mistakes there. And you want to set up open enrollment, you want to be good to your people.

29:52Gusto handles onboarding, health insurance, 401k, time tracking, commuter benefits, awful letters, and they even give you access to HR experts. So Gusto takes all of this off your hand and lets you focus on important stuff, your product and your customers it's super easy to set up and get started and if you're moving from another provider gusto will transfer all your data for you here's your call to action because you're a twist listener and you're part of the family you're going to get three months free incredibly generous totally unnecessary thank you so much to our friends at gusto.com slash twist you must go to gusto again gusto.com slash t-w-i-s-t to get three months free thank you gusto team so then this would lead us to um lpus uh obviously using a gpu very expensive right um but grok my friend chamat's company um has this uh uh inference engine and these lpus are you starting to see those and that hardware stack emerge these language processing um units and do you think that'll have a good effect on the industry in terms of lowering cost and having purpose-built hardware for the inference moment yeah so so as opposed to the lord of the rings right where there's one ring to rule them all i don't think that there's going to be one gpu lpu one accelerator to rule them all nor do i think there's gonna be one model to rule them all either i think there's going to be lots of different models with different objectives right like models that do different things, whether it's helping drive a car or cure cancer or be an AI character.

31:33Models will do different things. And then there will be infrastructure that is most efficient for each different type of model. And I think that's why you're seeing entities like Microsoft, Meta, et cetera, who are focused on building their own silicon. They're not trying to replace the GPU. They're just trying to solve for different models that they're running internally. So I think the Groks of the world will absolutely have a place somewhere, but I also think that you'll see GPUs have this place. And what we're observing their place is at foundation models, at latest generation models, the most demanding and complex workloads will continue to sit on GPUs.

32:18And NVIDIA just has this unbelievable solution for iterating continually better generations of GPUs. And we think that those models will continue to accrete to NVIDIA's platform. So they're going to win the day, no doubt. uh nvidia when it comes to training the models inference you might see other folks carve a niche for themselves is how you would bet this emerges yeah yeah yeah i think inference will will have various levels of infrastructure that provide solutions for it um i will say you know if a model is trained on a100s it'll probably run inference on a100s as well like like kind of tough to make that architecture shift.

33:03And it's tough because of the software that NVIDIA has, their driver solution, CUDA. NVIDIA very thoughtfully open-sourced that driver solution in the early 2010s to support this sector and the engineers who wanted to work on these products. and it has become effectively a default solution across the market, right? It's similar to drivers for CPU, right? Everything was x86 for decades, right? And it didn't really matter if something was better or not than it. It's just what people use, right? Because there's an efficiency loss if you say, well, I'm going to go learn this other thing and just hope other people will use it or you could just use the thing that everyone else uses.

Read the full transcript

33:54And that's what dominates the market. And NVIDIA has an amazing moat that they've developed out of the superiority of their software solution for their infrastructure. And I think that's going to keep people using their platform for a long time to come. Now, are people using CUDA yet to address other GPUs? because it's open source and it's obviously being used for parallel computing here when you've got a supercomputer you need to you know send a job across many different gpus are people for cuda or have they adapted cuda in order to you know have it send a job to some intel server some nvidia ones and is that opening up possibilities i think for a more open source future and then i'm curious what you think of open source chips and chip architecture and if you think that is ever going to have some sort of an impact here on the space sure so this will get a little bit you know beyond my domain expertise but yes there has been forks so to say and software that enables kudo to run on different infrastructure but it comes at a hefty cost right it comes at performance loss it comes with configurability loss, so much so that none of our clients are requesting that.

35:22We're talking 30%, 60%, 80 % performance loss. So the most natural thing for the largest consumers of this compute is to stick on NVIDIA infrastructure with NVIDIA software. And that goes to your second question, which is around open source. It's tough for me to say, but I would highlight the behemoth that's driving the research and the path forward on NVIDIA GPUs. They just have so much capital they're putting to work to ensure that they have the most performant piece of infrastructure in the market. that, you know, sure, there might be some use cases for that open source infrastructure to be applied, similar to how Grok can be there or other custom silicon chips.

36:16I think that the vast majority of workloads are going to accrete and stay with GPU infrastructure of which... Who's number two or three in the space? Does anybody have a chance of closing the gap? And because obviously people are watching NVIDIA print money, you know, and obviously that's, I don't know what percentage of your infrastructure that you provide is NVIDIA, but I'm guessing it's 90 % plus. But is there a number two or three in this space? And do they have a chance of gaining market share? Or do you think this is fait accompli? We're going to live in an NVIDIA world for the next decade.

36:50I think we're in NVIDIA world for a while. You have AMD out there, but AMD doesn't have a performant training fabric. That's something that's proprietary with an NVIDIA is in Fit and Band. So you can't build this comparatively performant training fabric with AMD infrastructure. So you can only use it then for inference. And it's sort of, well, if you've already trained your model on NVIDIA, it's a tough leap to want to move your software, move your infrastructure over to AMD compliant. So it's certainly a market I would expect that AMD is allocating their time to, but we're not seeing the customer demand for it at scale.

37:38And we really serve as scale consumers of compute. Certainly, there's your guys who want ones or tens of GPUs out there who will say, oh, I'd love to work with MI300s. but it's those entities want tens of thousands of gpus that are sticking with nvidia and we haven't really seen any deviation from that and for folks who don't know infiniband is kind of a contemporary or a competitor to ethernet or fiber in a data center if you had a bunch of storage in one location or even in a gpu between gpus passing data between them there has to be some way to move data from one cluster to the other if they were you know passing uh you know training data or something the speed at which the training data can get on the gpu to be processed that is a bottleneck and uh infiniband is the solution to moving large amounts of data am i correct in my description that's exactly right that's exactly right it's um it's it's uh infrastructure that nvidia acquired um and have integrated into their solution called the dgx you know solution and it is the most performant fabric solution other words a network solution for this infrastructure for data throughput hey startups you're a new company and you're looking to form your business but navigating through a maze of hidden fees and legal jargon it's complicated it's going to eat up all your time well northwest registered agent will form your business quickly and easily and it only takes 10 clicks and 10 minutes they provide you with a full business identity setup that means they'll give you everything you need to start and to maintain your business when you hire a registered agent to form your company they take care of everything you get a registered agent service a business address their corporate guide service a phone line mail scanning a free domain a website and hosting northwest registered agent makes the whole process transparent, quick, and enjoyable.

39:40Whether you're setting up an LLC, a corporation, or a nonprofit, they've got you covered. Here's your call to action. For just$39 plus state fees, Northwest Registered Agent will form your company and launch your business in minutes. Visit northwestregisteredagent.com slash twist today. That's northwestregisteredagent.com slash twist today. Is this still one of the key challenges in terms of training large language models is the throughput of the InfiniBand or Ethernet solutions to just move the data around? This is the bottleneck over GPUs in many of these jobs? Yes. It's critical to build with a non-blocking InfiniBand fabric.

40:20So non-blocking means that every component can operate at the same performance and efficiency as everything else. There's nothing blocking that performance. No bottlenecks. Yes, no bottlenecks. And it's really interesting because it's a physical engineering problem. So a 16 ,000 GPU fabric, which is about 2 ,000 nodes or individual servers with eight GPUs per server, it has 48 ,000 discrete connections that have to be made across the fabric. So you plug in InfiniBand into each GPU in the server, then that goes out to a switch and you're part of this fabric. Every connection has to be made correctly.

41:01And you're doing this with 500 miles of fiber optic cabling within that 16 ,000 GPU fabric. And we run a number of those of larger and smaller size. So we built a lot of these things, run a lot of fiber in our days. But it's a complex physical problem that no one's really been presented with before. This wasn't a problem when you're running with Ethernet or hosting websites and storing data lakes. You didn't have to build fabric this way. It'd be one connection per server, not eight connections per server. And to be doing it in this contiguous, non-blocking fabric in a single footprint. Right.

41:44And so it's just lots of new things that are happening at the same time with an immense amount of capital at risk and immense amount of capital is being consumed in the fastest paced technology environment that we've we've ever been in. And it's creating problems all over the market. And where we've found ourselves is having a software solution and a company that's only focused on these types of workloads. and accordingly we we accrete clients into our platform for having that best engineering solution and actually being able to deliver it to end consumers yeah and it's never we've never seen at scale companies like a microsoft like a meta like a google these companies are at scale they have massive amounts of capital which they can't deploy in m &a anymore right we have a framework in the West where you're not allowed to buy companies, and I made this point on All In a couple of months ago, instead of like, if you were Apple or you're Googling, you're sitting on tens of billions, hundreds of billions of dollars in cash.

42:49You can't buy Uber, Airbnb, you can't buy Coinbase. You're not allowed to buy even Figma for 20 billion. You can't even make a small purchase like that without getting blocked. But what's the next best thing you can do with that capital? You can build infrastructure as a weapon. you know and now you've got this massive infrastructure will you have jobs for it i'm sure there'll be some will those jobs turn into commercial products some will some won't but it's a better use than sitting on the cash or it's a better bet it's a better you know uh use of capital rather than trying to make a couple of points on it and you know we're buying back your shares it feels like gosh if you have this infrastructure you could have induced jobs which is to say some crazy person on the meta team is going to be like, what if we did X and having that infrastructure allows somebody with a crazy idea to then go give it a shot and spend a million dollars running a job across this infrastructure, whatever the pro rata version of it is.

43:47And maybe they find something really interesting. Who knows? Yeah. What, what people are going to do with this infrastructure? You do, you know, you're watching them. What's, what is the interesting jobs you're starting to see and use cases? I mean, some of it's public and some is private. So I'm obviously don't want you to betray anybody's trust here, but just what are people doing with this infrastructure that you find interesting when they come to you and they say, Hey, we need a solution for this, or here's what we're building. What are some of the things that you think are most promising or certain verticals, sectors that are most promising?

44:18Sure. So I think the areas where AI will be adopted first and fastest and do it at scale, right? Because you can always find like five users to do something, right? But how do you get 5 million users? Yeah. It's going to be within products that the user doesn't have to learn something new. It might not even be a new but. It just comes naturally to them. It feels organic. It doesn't require a new app to be somewhere. It's integrated into existing products. And I think that's largely going to be co-pilot. Various co-pilot solutions, not the name to one product, but just the idea that you're integrating AI into apps to assist a user with a pre-existing process.

45:02That's something that we're seeing scale right now. And the ability for those products to scale are limited by the amount of cloud infrastructure that's able to handle those users. Again, remember, each time you come in and query that co-pilot product, it's using a GPU. So cloud infrastructure inherently limits the pace at which those products can grow. And I think you've seen some products delayed even because there wasn't enough cloud infrastructure available to power their launch even. Yeah. If you look at search engines like Bing, Bing kind of was doing the custom answers. You'd have to click a second button to get it, right?

45:45Go to another experience. Whereas some search engines powered by AI were doing it automatically because they didn't have a large flow of it. if every single google search resulted in a query to a gpu they would actually bankrupt google right now because they have so many queries and at three or four cents extra per query there's not enough infrastructure in the world to convert all of those queries today look it brings up a question of will ai be a tax or a margin expander for software products i think some of them it will be a tax right it'll become mandatory and they might not be able to drive incremental direct revenue off those products.

46:23But the other outcome, if you didn't integrate that AI at that tax, could be you lose users and you lose market share to someone else. Right. If you look at search, that would be the perfect example. If Bing offers this to their 4 % or 5 % market share, they can lose money on it because they're building that business. Whereas Google, it's their core business. If they put it on all 90 % and they start losing money, they could just flip their business upside down right yes that's right and you know the other interesting point to that is you know google might have the option to integrate ai if it doesn't have the infrastructure available at the volume that's required and i think that's why you're seeing some companies bet up uh uh microsoft throw so much capex into ensuring they have the volume of infrastructure necessary because it having to compute at scale go back to my point earlier It decommoditizes compute.

47:20That in and of itself is a strategic advantage. So I'd say the other area that I personally think will accrue AI rapidly is in the advertising sector. Oh, really? I thought you were going to say healthcare or biology or something. I agree. I think that's the second half of this decade thing that we're extremely excited about. I mean, I can't wait to have infrastructure that directly supports the advancement of healthcare solutions. But advertising, I mean, think of the way that ads work, right? Like you throw an ad, you hope it reaches an audience, and then a subset of that audience will actually identify with it.

47:58It's probably a pretty small sliver of it. Yeah. Instead, if you could use generative AI to create on-demand, always-on ads for people that are 100 % specified to the metadata associated with that user, those are going to be much more highly effective. Yeah. Here's the example. You live in Utah, you have a green kayak, and you're searching for a new Kia. It's a blue Kia. And you've been looking at it for a few days. And instead of just receiving the general Kia ad of a gray Kia somewhere, you now get an ad that is a blue Kia with a green kayak on top driving through the desert in Utah to a river.

48:43And you can get several different iterations of that until you go buy that Kia. That's an area where the user doesn't know that it's generative AI, but it'll be so accretive and disruptive to the advertising sector that it'll just be mandatory for them to use it. Because the ad's effectiveness will increase that much. It's fascinating. you know you you look at what happened with meta there was this idea that when they lost access to smartphone data when apple anonymized it you know they would have a really hard time doing targeted advertising it actually kicked them in the ass and made them implement ai and they have now recovered and gone further i think in terms of personalization you're exactly right if it knows you have two kids it's going to put you know in that jeep wrangler you know with your kayak on the top or whichever car it is two car seats and it's going to show kids in it and the message will have something about how great it is for toddlers or young kids and here are some you know here's the um media center that puts the tvs on the back so they can watch netflix it's going to have so much information to customize the ad that you get the gap between the aspiration of the ad and the reality of your life is going to close right because ads are aspirational so it's my it's like minority report if you remember and everything goes back to minority report you know the customization of the ads will be absolutely phenomenal to a level that yeah it's like beyond creepy it's just like mind reading ads yeah and you've never had that before right and what's important there is it it took x amount of resources to generate that that one kind of mass media ad previously well now each time you have that iterative always on ad that that's querying infrastructure, right?

50:29So the infrastructure demand for this new type of advertising will be voluminous. Yeah. Right. It'll be more effective. And I think it will actually be better dollars spent in advertising, but it'll be an immense amount of infrastructure demand behind it. So I think co-pilot's there today and scaling, but the next big thing in there to really scale will be within the advertising space. Yeah. It makes a lot of sense. It's a huge business and yeah you know in anywhere there's a lot of data and a frequent transaction i mean that's just a great place for gpus and this ai revolution to take part of because it's frequent and there's a transaction this is where like amazon and how amazon sells you stuff and walmart and target my lord e-commerce in the last mile it's already been impacted in a way i mean if you look at the amount of advertising revenue for uber and instacart which are you know for uber eats essentially you know that's like being in line at the checkout counter and then they're doing a billion each i think a year roughly and then amazon might be doing 30 or 40 billion in advertising now like they're those three businesses which are seemingly transaction-based businesses right shopping for groceries food delivery mobile transportation then amazon buy anything they're all becoming advertising businesses it's like pure profit for them uh it's going to be wild the ai impact on those businesses and all roads lead back to generative ai right i guess it's just all converging here right and it all needs this type of infrastructure and that goes back to my point of like how much demand is there we just don't see the path to to resolve the amount of infrastructure that needs to be built for the demand that there is within you know at minimum the next few years.

52:14Yeah. Right. There's just so much needs to be built because last generation's clouds aren't designed for this. And it's not like you're swapping out a UI and saying like, oh, you like tweak some software here and there and all of a sudden it works, right? No, it's the foundational difference. It's the Tesla versus Ford manufacturing process. And it's - CPUs are just never going to take these workloads. CPUs will be for serving up images or light work. It's not going to - ever compete with this level of it's completely different um any worries about overbuilding this infrastructure at this point if we were going to start talking about a slowdown or and a certain amount of infrastructure is there somewhere on the chart that you start thinking yeah this is gonna we'll fill the demand um five years out ten years out where do you think we have enough capacity enough you know and the supply supply demand becomes normalized right now it's abnormal obviously when do you think this normalizes when do we catch up so between the infrastructure demand and the data center demand right so it's multiple components in here and then you know it's all the the infrastructure pieces that go into a data center yep then it's the power that goes into it like it's this really complicated physical stack yeah uh to serve it it honestly could be the end of this decade until you see this rebalancing of supply and demand.

53:39And not to say that that's overbuilding, right? That's just still on a heavy growth trajectory. That's just when infrastructure may have had an ability to catch up to where demand is. And I geek out over this stuff because that's my background and my co-founders. We're all from this commodity trading sector where all we did was assess supply and demand and understand physical disruption of commoditized markets. And that's exactly what we're looking at here. Yeah. But these aren't commodities yet. You know, like H100s trade at a price, and I guess they're commodities, but they feel like a very resource-contrained commodity right now.

54:17So I guess they are commodities. Yeah. They're sort of, you know, if you think about like cloud infrastructure for hosting websites, right? It was fungible, right? It didn't really matter if you were on AWS, JCP Azure to host your website, it all felt like the same thing. It was the same product to go host your website. What's changing is that lack of fungibility. An H100 hosted at AWS is very different than an H100 hosted at CoreWeave because of the way that we run that infrastructure differently from a software perspective and the way we build it from a physical perspective. right so like that that's the the commoditization that did exist that's now being decommoditized through software and infrastructure disruption right yeah it's amazing what a moment what a time to be alive well it's so much fun it was absolutely fascinating to talk to you for an hour and i'll let you get back to uh racking and stacking i'm sure you've got tons of h100s and A100s to unbox.

55:20I mean, just unboxing and racking stuff. I mean, you have hundreds of people doing that at this very moment? Hundreds. And semi-trucks arriving to our 28 data centers across the US. It's an operational feat. I think we're hiring 20 people a week right now. Yeah. And these are like system operations people. These are high-level people to come in and configure this infrastructure. Well, massive success. and thanks for building out the infrastructure. Let's solve some huge problems and we'll see you all next time in this week in startups. Bye-bye.

From the publisher

This Week in Startups is brought to you by…

OpenPhone. Create business phone numbers for you and your team that work through an app on your smartphone or desktop. TWiST listeners can get an extra 20% off any plan for your first 6 months at http://www.openphone.com/twist

Gusto is easy online payroll, benefits, and HR built for modern small businesses. Get three months free when you run your first payroll at Gusto.com/twist.

*Gusto pricing shown in ad is based on pricing prior to March 2025

Northwest Registered Agent. Northwest Registered Agent will form your business quickly and easily. For just $39 plus state fees, Northwest will handle your complete business identity. Visit https://www.northwestregisteredagent.com/twist⁠ today.

*

Todays show:

Brannin McBee joins Jason to discuss CoreWeave’s transition from cryptocurrency to AI (3:36), the energy dynamics of GPUs (14:04), innovations in data center cooling (19:28), and the future challenges and impacts of AI infrastructure on industries like search, advertising, and e-commerce (45:36).

*

Timestamps:

(00:00) CoreWeave’s Brannin McBee joins Jason

(3:36) The founding story of CoreWeave and the transition from the cryptocurrency market to AI

(12:49) OpenPhone - Get 20% off your first six months at http://www.openphone.com/twist

(14:04) Energy reliance of GPUs, the future of data centers, and the cost and efficiency of GPU cloud infrastructure

(19:28) Direct-to-chip liquid cooling and other cooling solutions for GPUs

(24:52) Demand trends for GPU capacity and the growth of inference linked to user growth

(29:06) Gusto - Get three months free when you run your first payroll at http://gusto.com/twist

(30:28) LPUs Vs. GPUs

(34:13) Dominance of Nvidia in training models, implications of open source chips and chip architecture

(39:01) Northwest Registered Agent - For just $39 plus state fees, Northwest will handle your complete business identity. Visit ⁠https://www.northwestregisteredagent.com/twist⁠ today.

(40:00) Challenges in training large language models and the use of infrastructure as a weapon by big tech

(45:36) The potential impact of AI on search engines, advertising sector, and ecommerce

(52:46) Timeline for supply-demand balance in AI infrastructure and the operational feat of running multiple data centers

*

Subscribe to This Week in Startups on Apple: https://rb.gy/v19fcp

*

Follow Brannin:

X: https://twitter.com/branninmcbee

LinkedIn: https://www.linkedin.com/in/branninmcbee

*

Follow Jason:

X: https://twitter.com/Jason

LinkedIn: https://www.linkedin.com/in/jasoncalacanis

*

Thank you to our partners:

(12:49) OpenPhone - Get 20% off your first six months at http://www.openphone.com/twist

(29:06) Gusto - Get three months free when you run your first payroll at http://gusto.com/twist

*Gusto pricing shown in ad is based on pricing prior to March 2025

(39:01) Northwest Registered Agent - For just $39 plus state fees, Northwest will handle your complete business identity. Visit ⁠⁠https://www.northwestregisteredagent.com/twist⁠⁠ today.|

*

Great 2023 interviews: Steve Huffman, Brian Chesky, Aaron Levie, Sophia Amoruso, Reid Hoffman, Frank Slootman, Billy McFarland

*

Check out Jason’s suite of newsletters: https://substack.com/@calacanis

*

Follow TWiST:

Substack: https://twistartups.substack.com

Twitter: https://twitter.com/TWiStartups

YouTube: https://www.youtube.com/thisweekin

Instagram: https://www.instagram.com/thisweekinstartups

TikTok: https://www.tiktok.com/@thisweekinstartups

*

Subscribe to the Founder University Podcast: https://www.founder.university/podcast

More from This Week in Startups

All 653 episodes
CoreWeave’s Brannin McBee on the future of AI infrastructure, GPU economics, & data centersThis Week in Startups · 56 min
Listen in VO