The marketplace for AI compute with Jared Quincy Davis from Foundry

22 Aug 2024 · 43 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: No Priors - The Marketplace for AI Compute with Jared Quincy Davis from Foundry

Episode Overview In this episode of *No Priors*, co-hosts Sarah Guo and Elad Gil interview Jared Quincy Davis, the founder and CEO of Foundry. The conversation revolves around the state of AI cloud computing, GPU utilization, and innovations aimed at improving cloud economics for AI workloads.

---

Key Participants

  • Sarah Guo: Co-host, startup investor, founder of Conviction.
  • Elad Gil: Co-host, serial entrepreneur, and startup investor.
  • Jared Quincy Davis: Guest, former DeepMind researcher, founder and CEO of Foundry.

---

Episode Highlights

Introduction

  • The podcast introduces Jared Quincy Davis, discussing his background and the genesis of Foundry, which aims to provide a public cloud specifically for AI workloads.

Foundry's Mission

  • Genesis of Foundry: Inspired by breakthroughs like AlphaFold 2 and ChatGPT.
  • Objective: Make high-quality computational resources accessible beyond the scope of organizations like OpenAI and DeepMind.
  • Vision: Reduce the cost of advanced computational tasks significantly, thereby increasing the frequency of groundbreaking innovations.

Foundry's Product Offerings

  • AI Cloud Services: Described as a public cloud built specifically for AI workloads, with a focus on improving economic efficiency.
  • Infrastructure as a Service: Providing economically viable access to state-of-the-art AI systems.

GPU Utilization

  • Current State: High-level discussion on the underutilization of GPU resources across various environments (hyperscalers like AWS and private clusters).
  • Utilization Rates: Even sophisticated organizations experience sub-80% utilization, often due to hardware failures and the need for healing buffers.

Cloud Economics and Market Dynamics

  • Current Challenges: The existing AI cloud market primarily functions as co-location rather than a true cloud model, complicating cost efficiency.
  • Innovations: Introduction of new models and mechanisms to improve the economics of AI cloud computing, including dynamic resource allocation strategies.

The Future of GPU and AI Workloads

  • Predictions: Jared discusses the evolving GPU market dynamics and how Foundry is positioned to leverage these changes.
  • Compound AI Systems: A new paradigm where AI workloads are designed to be more efficient, involving the integration of smaller models trained on high-quality data.

Recent Innovations

  • New Releases: Jared discusses recent products from Foundry aimed at improving cloud economics and utilization, including mechanisms for better resource sharing and management.

Closing Thoughts

  • Future Trends: A push towards batch inference and synthetic data generation as primary workloads that require less interconnect and can be executed more efficiently.
  • Paper Discussion: Jared highlights a recent paper on compound AI systems design, emphasizing the importance of verifiability in AI tasks and its implications for future AI model performance.

---

Key Takeaways

  • Foundry aims to democratize access to advanced AI compute resources.
  • GPU utilization is currently low across the board, highlighting inefficiencies in the market.
  • The evolution of AI workloads is shifting towards more efficient, compound systems that use smaller models and synthetic data.
  • Innovations in cloud economics could lead to a more sustainable and scalable future for AI development.

---

Further Reading & Resources

  • Find transcripts and more episodes at [No Priors](https://no-priors.com).
  • Follow the podcast on Twitter: [@NoPriorsPod](https://twitter.com/NoPriorsPod).

Feedback

  • Email feedback to: show@no-priors.com

This podcast episode provides rich insights into the current and future landscape of AI cloud computing, and the innovative strategies being employed to improve efficiency and accessibility.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:05Hi, listeners. Welcome to KnowPriors. Today, we're talking to Jared Quincy Davis, the founder and CEO of Foundry. Jared worked at DeepMind and was doing his PhD with Matei Zahari at Stanford before he began his mission to orchestrate compute with Foundry. We're excited to have him on to talk about GPUs and the future of the cloud. Welcome, Jared. Thanks, Sarah, and great to see you. Thanks a lot as well. Yeah, great seeing you. The mission at Foundry is directly related to some problems that you had seen in research and at DeepMind. Can you talk a little bit about the genesis? A couple of the most inspiring events I've witnessed in my career so far were the release of AlphaFold 2 and also Chachupiti.

0:44I think that one of the things that was so remarkable to me about AlphaFold 2 is initially it was a really small team, you know, three and then later 18 people or so. And they solved what was kind of a 50-year grand challenge in biology, which is a pretty remarkable fact that, you know, every university, every pharma company hadn't solved. And similarly with Chai TPT, a pretty small team, opening up with 400 people at the time, released a system that really shook up the entire global business landscape. That's a pretty remarkable thing. And I think it's kind of intriguing to think about what would need to happen for those types of events to be a lot more common in the world.

1:18And although those events are really amazing because of the small numbers of people working on them, I think it's not quite the David and Goliath story. Neither are quite the David and Goliath story that there appear to be when you double click. In OpenAI's case, you know, they were only 400 people, but had$13 billion worth of compute, you know, which is quite a bit of computational scale there. And in DeepMind's case, it was a small team, you know, but obviously they were standing on the shoulders of giants in some sense with Google, right? And the leverage that they had via Google. And so one thing I think that we thought about is, you know, what can we do to make the type of computational leverage in tools that are currently exclusively the domain of OpenAI and DeepMind kind of available to a much broader class of people?

1:57And so that's a lot of what we worked on with founders saying, can we build a public cloud, built specifically for AI workloads, where we reimagine a lot of the components that constitute the cloud end to end from first principles? And in doing that, can we make things that currently cost a billion dollars, cost 100 million, then 10 million over time? And that'd be a pretty massive contribution. I think it would increase the frequency of events like AlphaFold 2 by 10x, 100x, or maybe even more super linearly. And we're already starting to see the early signs of that. But quite a lot of room left to push this agenda.

2:32So really exciting. So that's kind of maybe an initial introduction preamble to how we thought about it. And I can trace that line of reasoning a bit more, but that's kind of part of what we've done. Jared, for anybody who hasn't heard of Foundry yet, what is the product offering? Yeah, so Foundry, we're essentially a public cloud built specifically for AI. And what we've tried to do is really reimagine all of the systems undergirding what we call the cloud end to end from first principles for AI workloads. And we've tried to do this in a bit of a new way. I think the AI offerings from the existing major public clouds and kind of some new GPU clouds haven't really re-envisioned things.

3:13And by thinking about a lot of these things a bit anew, we've been able to improve the economics by, in many cases, 12 to 20x over lower-tech GPU clouds and existing public clouds. And, you know, partially based on some of these products that we'll talk about today that we're releasing and a lot of new things that we're working on, we think we can push that quite a bit further as well. And so, you know, our primary products are essentially infrastructure as a service, so our customers come to us for elastic and really economically viable access to state-of-the-art systems. And also a lot of tools to make leveraging those systems really seamless and easy.

3:49And we've invested quite a bit in things like reliability, security, elasticity, and just the core price performance. How underutilized are most GPU clouds today? And I think there's almost three versions of that. There's things on hyperscalers like AWS or Azure. There's large clusters or clouds that people who are doing large-scale model training or inference run for themselves. And then there's more just like everything else. It could be a hobbyist. It could be a research lab. It could be somebody with just, you know, some GPUs that they're messing around with. I'm sort of curious for each one of those types of, or categories of users, like what is the likely utilization rate and how much more do you think it could be optimized?

4:23Is it 10 %? Is it 50 %? Like, I'm just very curious. One of the most, I'd say, positive cases with the highest utilization, which is the case where you're running kind of an end-to-end pre-training job, right? And so that's a case where you've done a lot of work up front. You've designated a time that you're going to run this pre-training workload for, and you're really trying to get the most utilization out of it. And for a lot of companies, utilization, even during this phase, is sub-80%. So why? One reason is actually that these GPUs, particularly the newer ones, actually do fail a lot, as block practitioners would know.

5:00And so one of the consequences of that is that it's very common now to hold aside 10 % to 20 % minimum of the GPUs that a team has as buffer, as healing buffer in case of a failure so you can slot something else in to keep the training workload running. And so even for a lot of the more sophisticated orgs running large pre-trainings at scale, the utilization is sub 80%, sometimes less than 50 % actually, depending on how bad of a batch they have and the frequency of failure in the cluster. And so even in that case, now there are often also large gaps and intermissions between training workloads.

5:33Even if you have the GPUs are dedicated to a specific entity. And so even in those most conservative cases, which I'll come back to the less conservative extreme cases, utilization can be really quite a bit lower than people would imagine. So we can pull on that case a bit more because I think it's actually quite counterintuitive and really interesting. I think there's a really fundamental disconnect between people's mental image of what GPUs are today and what they actually are. I think that in most people's minds, GPUs are chips, right? And we talked about them as chips, but actually the H100 systems are truly systems.

6:09You know, there are 70 to 80 pounds, 35 ,000 plus individual components. They're really kind of monstrosities in some sense. And the remarkable thing that I think Jensen and NVIDIA have done, one of many, is they've basically taken an entire data center's worth of infrastructure and compressed it down into a single box. And so when you look at them from that perspective, the fact that it's 70 pounds isn't quite as alarming, but it is, these are really gnarly systems. and when you end up composing these individual systems, these DGXs or HTXs into large supercomputers, what you're often doing is you're interconnecting thousands, tens of thousands, hundreds of thousands of them.

6:46And so the failure probability kind of multiplies. And so because you have millions perhaps of individual components in this supercomputer, the probability that it will run for weeks on end, and this is basically a verbatim quote from Jensen's keynote, is basically zero. That's a little bit of a challenge and funnily enough, I think the AI infrastructure world is still somewhat incoet and so doesn't have perfect tooling, broadly speaking, to deal with these types of things. And so one of the compensatory measures I think people take is reserving this healing buffer, for example. I think that disconnect maybe helps explain why these things fail.

7:21And actually, funnily enough, the newer, more advanced systems anecdotally fail a lot more than historical systems that were worse in some ways. And do you think that's just like a quality control issue for those systems? Or do you think it's just some form of complexity with some failure rate per component that's rising? What do you think is sort of the driver of that? I think it's more of the complexity has grown, right? And we're in a different regime now. You know, I think that it's fair to say, so maybe stepping back again to definitions, we throw the term large around a lot in the ecosystem.

7:54I guess one question is, what does large mean? and one useful definition of large that I think roughly corresponds to what people mean when they invoke the term is that a large language model you enter the large regime when the essentially the amount of compute necessary to contain even just the model weights starts to exceed the capacity of even a state-of-the-art single GPU or single node I think it's fair to say you're in the large regime when you need multiple state-of-the-art servers from NVIDIA or from someone else to even just contain the model, you know, just basically run the training or definitely even just to contain the model.

8:35That's definitely the large regime. And so the key characteristic of the large regime is that you have to somehow orchestrate a cluster of GPUs to perform a single synchronized calculation, right? And so it becomes a bit of a distributed systems problem. I think that's one way of characterizing the large regime. Now, a consequence of that is that you have many components that are all kind of collaborating to perform a single calculation. And so any one of these components failing can actually potentially, you know, lead to some degradation or challenge down the street and stop the entire workload, right?

9:10You know, you've probably heard, people have talked a lot about InfiniBand and the fact that NVIDIA, part of NVIDIA's advantage comes from the fact that they do build systems that are also save the art from a networking perspective, right? And their acquisition of Mellanox was one of the better of all time, arguably. from a market cap creation perspective. And the reason why they did this is because they realized that it would be really valuable to connect many, many machines into a single, almost contiguous supercomputer that almost acts as one unit. Yeah, and the challenge of that, though, is that there now are many, many more components and many more points of failure.

9:42And these things kind of, you know, the point of failure kind of multiplies, so to speak. You implied that, like, GPUs, well, you described GPUs as this unique asset that is more CapEx than OpEx and that the hyperscalers like an Amazon are making a certain assumption about what that depreciation cycle is. Like, where do you think the assumptions for Foundry versus those hyperscalers or versus, let's say, like a core weave are different? This opens up a pretty interesting conversation around like, what is cloud? I think we've kind of forgotten in this current moment what cloud was originally supposed to be and it was value proposition was intended to be.

10:22I think current AI cloud is not cloud in the originally intended sense by any means. So we should pull on that thread. But I'd say right now it's basically co-location. Yeah, it's basically co-location, right? It's not really cloud. Maybe, yeah, it's definitely worth pulling on that thread a little bit. Yeah, do you want to break that down for sort of our listeners in terms of what you view as the differences? First, I guess, as a little bit of context for people, the cloud as we currently know it is arguably one of the most important business categories in the world. That's, I think, pretty clear.

10:49The biggest companies in the world are either Clouds, Azure, AWS, GCP, you know, core component of the biggest companies in the world, or NVIDIA, who sells to the clouds, obviously. So it's clearly an important category. AWS is arguably a trillion dollars plus of market cap if you broke it out from Amazon. So, you know, at one point, relatively recently, it was all of Amazon's profit and more. So it's an important category, to say the least. Cloud, as we know today, really started in 2003, which is when Amazon, you know, they cated 50 people initially, and I believe they started working on AWS.

11:19And they worked on this for three years with an Amazon, launched in 2006 in March with S3. Later that year, in September, they launched EC2. And that kind of was the beginning of the cloud. It took quite a while for this model to catch on. And people were not quite clear on why this would be useful, even until quite recently. And so in 2009, 2010, you know, Matei, Sarah, who you mentioned, wrote this paper called Above the Clouds with some collaborators at Berkeley, Berkeley-Baron Cloud Computing. And they talked about why cloud would be a big deal. and I think it was not, to say the least, not appreciated that the points they were making were valid at the time.

11:53I think it was kind of not clear to people, to say the least. Fast forward a few years, Bezos in 2015 said AWS was, for all intents and purposes, market size unconstrained. And he was kind of laughed at by many. That was seen as a really ludicrous statement to make. Fast forward, no one's laughing. In 2019, I remember people saying, it's not clear if cloud would be a big deal. Snowflake hadn't yet scaled. you know, Databricks had a new scale. I wonder if it's worth making a distinction on these things because, you know, I agree with parts of what you're saying, but, you know, I was building startups in that era and at the time, there was a pretty strong belief that at least for startups, these clouds were incredibly useful because it used to be you'd spend a lot of money and effort on setting up your server racks and wiring everything up and all the rest.

12:41And then, you know, the emergence of things like AWS and nobody expected it to be Amazon. It just didn't fit what people thought of as basically an e-commerce marketplace company. They were suddenly building infrastructure and providing it. It felt very natural for somebody like Google to provide that. But I think for startups, because I was doing a startup in 2007, we thought that it was magic. Because suddenly, and not all the services were there yet and everything else, but suddenly you didn't have to deal with all the infra. And there's a number of companies that started before that, like Twitter and others, who kind of ended up having to continue to build and maintain their own clouds.

13:13And it was really brutal. And then to your point, I think the big transition was degree to which enterprises, particularly in regulated services or financial industries, kind of fought it initially. And then they started adopting it. So I agree with the law here saying I wouldn't say that it was one of those things that nobody believed in, though. You know, like I actually felt there was a lot of buy in and a lot of belief and, you know, adoption. Definitely not nobody. I think, you know, pricent people and a lot of builders really saw the value really early, particularly startups in the VC community.

13:40You know, I think Berkeley and some of the researchers there really got it early, like 2009 and earlier. You know, by 2015, I think it was already around$3 billion a run rate. So it was not small. It was very far from what it is today, but it was not small. They didn't break it out yet, but it was, you know, not small at all. It was a meaningful thing. I think that people didn't recognize it would become anything like it is today. And it still wasn't clear to people what the value proposition was. You may recall that there was, you know, Dropbox, for example, famously exited the cloud and went back on-prem.

14:06and that story they published and why they did that and the economics of it. And that caught a lot of attention and people were saying, I'm not sure if cloud actually makes sense anymore, et cetera. And I think people kind of lost track of that. Yeah, I think one of the insightful things that some of the early cloud systems companies like Databricks and Snowflake recognized was one of the value propositions of the cloud that was really special. I'll actually use one of the quotes from well, Snowflake Saunders to kind of illustrate this was the thing, they were trying to think about what was unique about the cloud, what you could do in the cloud that you couldn't do anywhere else.

14:37And the killer idea that they converged on, fundamentally, the cloud made fast free. Quote, unquote, fast was free in the cloud, is how they put it explicitly. And the idea was, if you have a workload that was designed to run for 10 days on 10 machines, in the cloud, you could theoretically run it on 100 machines for one day, or 10 ,000 machines for 15 minutes. And that would cost the exact same. And so you could run it 1 ,000 times faster for the same cost, theoretically. And it's like, well, that's kind of a big deal, right? You can run something 1 ,000 times faster for the same cost and then give the compute back.

15:08Now, what that would require, though, a number of things to make that actually work. One being that you'd have to be able to kind of reshape a workload that was designed to run on 10 machines, to run on 10 ,000 machines, not trivial. Another is that you actually have to have the 10 ,000 machines worth of capacity in the cloud and make sure that was utilized to make the economics kind of work. But yeah, if you could do those two things, it'd be a really, really big deal. That has to be free in the cloud. And so one of the key ideas was this elasticity, right? And I think that's one of the things that's really absent in AI cloud today.

15:37In AI cloud today, you're kind of forced to get really long-term reservations, often three years for a fixed amount of capacity. No one really wants 64 GPUs for three years or 1 ,000 GPUs for three years. You want 10 ,000 for a couple of months, maybe nothing for a little while. Maybe you don't know how much you need because you're launching a new product and you're not sure how much demand there will be for it and how much inference capacity you'll need, et cetera. It's very, very challenging if you have to reserve the total amount that you may need up front for a long duration. I want to go back to this idea of like, it's actually hard for a lot of engineering teams, especially younger ones to picture like a pre-cloud world because their experience with it is like, I have virtual machines on Amazon, they're limitless, like maybe it's serverless.

16:29And like, if you go all the way back to what Alad was talking about. You gave us a great historical view, but if you look at like functionally, like I had, you know, I had my servers in my closet on-prem, right? And then I had co-location, which is like, I have my servers. They're still my servers. I control them and I manage them, but I, they physically live in somebody's data center where they're offering me like real estate, like cooling and power. Right. And then you had hosting, which is like, I'm buying a machine in a data center, like reservations for a long time, essentially, or for the life cycle of the machine.

17:06And then you had virtualization and containerization and all of these services that came out of this, like, you know, separation into cloud services where you have higher level functions with scheduling orchestration. And, you know, serverless is like, I'm just going to write like the logic, you deal with it and place that workload. And we haven't obviously gone to an endpoint in non-AI computing, but I feel like the engineering world is so used to being over here, whereas the AI hardware resource world, it's like we're somewhere between colo and hosting still. That's right. Yeah. And that's challenging.

17:42That means that you have to raise all the capital you may need up front. You know, means that, you know, that's a challenging model. And perpetuity means also that, you know, you can't quite grow the product elastically as demand and interest in it grows. So you're kind of bottlenecked by the supply chains in a way that I think developers haven't experienced in quite a while. So it's a pretty challenging state, I think, of affairs. And it's also a very challenging from a risk management perspective for these companies because they're making these big commitments. that are potentially, if they don't work out, are pretty catastrophic for them on this hardware, paying all up front, paying for these long-duration contracts, etc.

18:20It's a pretty challenging thing. And there's no analog yet. The markets aren't mature enough that there's any analog. So what we have in other domains, like in commodities markets like wheat, oil, etc., where you can buy options and futures and hedge and sell back and things like that. It's still pretty in-co-ed in terms of where the market's at. And yeah, I think that's leading to a challenged state of affairs that's going to continue to bring a lot of pain for people. And so there are several things I think we've employed to do this better. Some are mixed as kind of business model innovations and technical innovations at the same time.

18:52But I think we're making a pretty substantial dent in this, but also in a way that's really viable economically for us. And does it involve buying all of GPUs and taking undue risks on them per se? And so that's kind of a lot of what we've tried to do is can we do something a lot more efficient? And, you know, can we find some leverage, some points of leverage to address this problem? You guys have some new releases as of, I think, a day or two ago. Can you describe what's just come out from Foundry? So I think right now, the AI cloud business is very much like a parking lot business. And that sounds really funny because cloud is supposed to be high tech.

19:31And you can hardly conceive of a less sophisticated business, at least on the surface, than parking lots. And what do I mean by that? Well, there's fundamentally, there are two models in the parking lot business. One is pay as you go. For pay as you go, the race of your curious, you know, you may or may not find a space. I'm sure many of us have the experience of driving through SF and seeing a lot full sign, you know, for lot after lot as we drive around trying to park. And if you do get a spot, you might pay, you know, the$12 an hour, you know, rate or something like that. I'm choosing that rate because it's the rate of AWS, you know, for on demand.

20:04On the other hand, if you want to guarantee that you'll have a spot and also have a better rate, you can basically buy a spot reserved. So you can have your own reserve parking spot in your building. Maybe you pay$4 an hour, so you're getting a massive discount, but it's$3K a month, effectively, which is actually pretty substantial. And if you're only using it 40 hours a week when you're in the office as a typical worker, it actually might be effectively$16 an hour as opposed to$12. So it's actually worse. And so I think one kind of funny analogy for one thing that we want to do with a couple of these products is kind of create the, enable the equivalent of allowing people to park pay as you go in someone else's reserve spot.

20:46And that sounds kind of funny, but you can imagine that, okay, that'd be actually an interesting thing. And if you could do that, then depending on how much, how the percentage of the spots that are typically reserved, you might have 10x the effective capacity in a lot. You know, and then also it can kind of be a win-win. Instead of the pay-as-you-go person paying 12, they can pay something much lower. In this case, I'll say seven, but they actually be a lot lower. The person who owns the spot, instead of paying four, can make five equivalent. And then the lot can also make a couple of bucks. And so it was kind of a win-win-win for everyone.

21:15And you're kind of double-picking the lot and it's really, really efficient. Now, that sounds great, but it wouldn't quite work if you showed up to your reserve spot and there was someone parked there. That might be a little bit aggravating. It also might not work if you're forced to call two hours in advance and say, hey, I'm coming. And then the person who was parked in your spot had to leave their dinner reservation to move their car. That wouldn't be a fun model. And so I think one thing that we had to do was kind of create the analog of a system to make this all really convenient. And so maybe the V1 of this system was you came into the lot and a sensor was triggered saying, you're here to go to your reserve spot.

21:51And then some valet ran to the car that was parked there and moved it. And they maybe moved it from the second floor to the 10th floor. And then the person who had parked there previously now comes to the counter, asks the valet where their car is, and gets some ticket saying it's on the 10th floor. But there are no stairs, so they have to walk up the stairs and get it. I'm stretching this analogy, but you get the idea of it to be kind of inconvenient. And so part of what we did was added more and more convenience features, which we broadly call spot usability. And these are a lot of things that we're going to need to add.

22:16And so the scenario that we're now at is basically you letting someone else use your spot show up. The sensor kind of knows, okay, you're here. And then the car in your spot is automatically moved. via conveyor. To another spot, we're managing the spaces to ensure we can move it somewhere. And then when the person comes to pick up their car, it's kind of brought to them. Now, yeah, so it's all really convenient. It's kind of seamless for everyone involved. It creates a ton more effective space, allows us to get much better economics out of the machines, and is really, really helpful for companies.

22:42That's one really powerful, I think, thing that we've done. And this is kind of offered in the context of a spot product. I think people are somewhat familiar with spot usage in the typical cloud context. It's a lot more challenging to do with GPUs for a few different reasons, which is why there are not very many GPUs available on Spot, definitely not at scale and definitely not with Interconnect. And so one of the things we want to do is enable this, and it makes a lot of other things possible that are pretty neat, and so it's a mechanism we're employing in a few different ways. So that's, I think, one thing, and I'll give some more analogies to explain why that's powerful.

Read the full transcript

23:12Well, we found that companies are using this type of mechanism quite a bit for everything from training, which is classically seen as a workload that's difficult for Spot, but also especially for things like inference, batch inference especially. And that actually opens up another interesting conversation about the different classes of workloads and what each workload needs and cares about and how this might evolve over time. Actually, that ties a little bit to the compound AI systems concept. And also see a Lama 3.1 release in a funny tangential way. But yeah, that's kind of one analogy for the product that we launched on the Foundry Cloud Platform around Spot.

23:45Yeah, I think that Spot usability increasingly deep and flexible and automated is like a really powerful primitive. Changing tacks a little bit to just something I think the entire industry, like the tech industry is very interested in. We did a little bit of work together a while back, just understanding like where, you know, where is the GPU capacity in the world today, right? Of all of the different types, how much is it? Like how consolidated is it? And obviously this is near and dear to your business. Can you just describe a little bit like where you think we are and then like what caused the shortage sort of last year?

24:26One kind of funny bit of trivia that I've posed to a few people that I think, you know, reveals how off base a lot of our priors are is kind of what percentage of the world's GPU petaflock capacity or exa-flock capacity is kind of owned by the major public clouds. and I've asked this many people and I typically have gotten guesses in the high tens of percents and the only time I got a lower guess was from Saatchi at Microsoft who guessed basis points which actually is correct it's a very small amount and maybe as one anecdote to maybe illustrate how this looked at least a couple of years ago it's an evolving thing but how it's looked, the example of GPT-3 and its training is kind of an interesting one it's a bit dated but I'll use it just because the numbers of the types of machines and numbers there are public, as is not the case for some of these other systems.

25:17And so GPT-3 was trained on 10 ,000 V100 GPUs in an interconnected cluster in Azure for about 14.6 days. To put that in perspective, it was a state-of-the-art system. It was kind of, by many estimates, eight figures for a single run at the time. So it's a pretty substantial investment by OpenApp. And that tells you that 10 ,000 V100s running continuously for 14.6 days is quite a bit of compute. I think one interesting kind of maybe trivia question then is how many equivalent GPUs normalized in terms of the number of flops. You know, they weren't fully interconnected, but just interesting processing measure anyway.

25:53Were there in the Ethereum network at the peak of Ethereum? And so I kind of asked people this question and it's fun to see people guess. Can I solicit a guess actually a lot? Can you make a guess there? You might know. So we've talked about this. We have this conversation, so I'm not allowed to guess anymore. Yeah, we have this conversation. So do you know by chance? For Ethereum? Yeah, for Ethereum. And don't think too hard. Just make a guess based on priors. When? So when it first launched or? The very peak, the tippy top of Ethereum. How many V100 equivalents were there, given there were 10 ,000 for two weeks for GPT-3?

26:23How many were there in Ethereum? Noting, by the way, that these were running 24-7 in Ethereum. So you can modulate your guess based on that. I would guess a few hundred thousand, a few million. That's an aggressive guess. Yeah, that's an aggressive guess, and you're actually very correct. It was about 10 to 20 million, which is quite a substantial scale. And you can, by the way, syscheck this really easily by looking at basically the hash power in Ethereum at the peak, which was around 900 terahash per second, I believe, about a petahash per second, so quite a bit of peak power. And a typical V100 will give you between, I believe, 45 to 120 megahashes per second if you really know what you're doing.

26:59So that's kind of one guess. There are tens of millions that you put. Yeah, because it's funny, I remember Bitcoin even years ago, all the CPU dedicated to it at the time was like larger than all of Google's data centers. Yeah, Bitcoin, though, used a lot of ASICs in particular. Ethereum actually had a higher relative ratio of GPUs. And so the larger GPU, sorry, Ethereum mining providers like Hive, for example, had less than 1 % of global hash power. you know and so you can actually start to extrapolate you know in you know they had had quite a few gpus like tens of thousands of nvidia gpus yeah it's tough to give you a number of the the scale and then a lot of other you know mining my equipment for way less than one percent of the total hash power about 0.1 percent um so quite a bit of hash power in ethereum and i think that's kind of one proxy but i have to give you one more anecdotal on that line actually an iphone 15 pro now is actually stronger than a v100 as a funny as a funny example has about 35 teraflops in a 16, I believe, where V100 is around 30.

28:00And so there's actually quite a bit of compute in the world, broadly speaking. That's, I guess, the point I'm making. Now, it's not all useful. It's not all interconnected. It's not all accessible. It's not all secure. But this is one point to make, that there's a lot of compute in the world. And even for the high-end GPUs, there's a lot more than people think. By many measures, utilization of these even H100 systems are kind of state-of-the-art, the most viable, the most precious, et cetera, is in many cases 20%, 25 % or lower, according to some pretty high-quality data I've seen from some great sources here.

28:27Yeah, so quite slow. As I mentioned, even during these pre-training runs, it's often 80 % lower because of the healing buffer partially. That's actually ties to another product that we've launched, which is this product that we built actually largely for ourselves called Mars. It's kind of a funny name, but it's Monitoring, Alerting, Resiliency, and Security. It's basically a suite of tools that we've invested a lot of IP in to really boost and magnify the availability and uptime of GPUs for our own platform. It was actually something that we planned to make available to other people just to use in their own clusters as well.

28:57Actually, one of the reasons why we invested in Spot is because we reserve very aggressively healing buffer ourselves so that if there's a GPU failure, we can automatically swap in another GPU and a user won't perceive a disruption. And so we actually, we maintain buffer for that reason. And so actually being able to pack that buffer with preemptible nodes is actually a really useful thing. But now we're allowing other people to do this, including third-party partners who want to, for example, make their healing buffer available to others through Foundry is really offsetting their economics and the cost of the cluster for them.

29:30So it's a really, really powerful thing. And so between Mars and Spot, you can kind of see how these things are really interconnected and in a nice way. But yeah, the number of GPUs available, there's quite a few, particularly if you look at it more broadly in terms of total AI compute capacity, the percentage that's accessible, useful, and used is a pretty diminished fraction of the total. How do the GPU market dynamics and your prediction of them factor into foundry strategy going forward? Right. Because there are especially for anybody doing large scale training jobs, there is definitely a, you know, a significant effort to be at the leading edge.

30:07right? Access to B100 and beyond is at a premium and then access and, you know, the largest possible interconnected cluster with sufficient power is also a fight now. It sounds like you, you know, see the opportunity differently or you feel like there are resources that can be used that don't require just building new data centers. I think it's a little bit of all of the above, to be clear. I think two things would be true at once, and that's definitely the case here. I think there'll be many workloads and use cases for which having state-of-the-art, extremely large interconnected clusters is a really valuable thing.

30:39Part of something, though, that we're noticing and also trying to promulgate further are basically paradigms that don't require this, though, as well. And so here's why I'd say two things can be true at once. And so I think there's a massive shortage of both power, space, and interconnect for the largest of clusters. It's actually very hard to come by and to construct or to find a really large interconnected cluster. It starts to vanish the larger the cluster gets. Like there are a lot more 1K clusters than 10K clusters and 20K clusters and you can keep going. Now, I think one thing that, and it'll get harder and harder to keep the scaling going from there.

31:16You know, I think there's one question, which is how will we continue to push the scaling laws? You know, one fact about the scaling law curves is that they're all plotted on logarithmic or house and, you know, things get better predictably, but it requires a continued 2Xing or 10Xing to get that next bump in performance. but it's quite a bit harder to get the next to next thing. And so it starts to become kind of intractable pretty quickly. And so it's already prompted, I think, a lot of innovation. So Google, for example, has been doing a lot with training across facilities, for example, across data centers, interconnecting them for these models like Palm 2, something that previously would have seemed to be inconceivable or things like Deepaco or Deloco, these models they release that have trained across facilities.

31:57And that's one innovation, but I think actually an even slightly more radical thing is we're starting to see a shift towards a pretty different paradigm. I think myself and a number of my collaborators and Matei and others have kind of termed this compound AI systems. I think you actually see it with these most recent models like alpha geometry, alpha code, and Lama 3. And so I think this actually points the way towards what the AI infrastructure future might look like. And I think it looks a lot less like everything requiring these big clusters. And it's a little bit more interesting. And so maybe I'll use 5.3 as an example.

32:33With 5.3, they take a little bit of a different approach where they trained a really high, Microsoft trained a really high quality small model on high quality data, you know. And the small model did not need the kind of big interconnected cluster. You can train it on a pretty small cluster. However, it was still, you know, a non-trivial endeavor because they had to curate and get and obtain this kind of high quality data. And so one of the things I think you're seeing is for these models, like some of the, like the LAMA 3.1, 8B and 70B variants. Those models are really small, but they're extremely smart.

33:04They're smarter than, you know, much, much larger systems, you know, like the prior generation for OpenAI. And the way that they trained LAMA 8B and 70B looks a little bit different. So what they did was they did, they generated a ton of synthetic data. It seems with LAMA 3.1, 405B. And they distilled that larger model into the 70B and 8B variant. And so they got a very, very high quality small variant. Another example is with AlphaCode 2 was able to achieve extremely high code proficiency and win competitions with really, really small language models. But what they did was they called the model a million times for every question.

33:44That's an embarrassingly parallelizable workload. You can scale it horizontally infinitely. They called it a million times per query and then had a nice kind of pretty elegant regime to filter down to the top 10 responses, which they didn't try one by one. And so this is basically what they did. They generated a million candidate responses and basically filtered down to the best one as a way to solve coding, so to speak. So that's pretty interesting. And I think you're seeing that type of regime a bit more and more. You know, same with alpha geometry. Really, a powerful system was just announced recently.

34:15You know, one, you know, the silver medal level in the IMO and broadly, not just geometry for a broader class of problems. and you know these are kind of compound systems with major kind of synthetic data generation pieces and so i think you're seeing people kind of move computation around interpolate between training and inference for example to make the best use of the infer they have and there's actually a kind of a funny i think reframing of the scaling laws that we hold near and dear as an ecosystem i think that you know the chinchilla paper scaling laws that deep mine uncovered have fueled a lot of the scaling effort but actually one kind of funny way of looking at those results is they show that if you want to make a monolithic model as smart as possible, there is an ideal way to distribute parameters, compute, and basically train iterations.

34:58One, the funny thing I think some people have done, like Mistral, is actually choose to maybe inefficiently train a small model to be smarter than it should be, wasting money. But then that small model is actually really cheap to inference because it's small and it's way smarter than it should be for its size. And so I think people are getting more sophisticated at thinking about cost in a more of a lifecycle way. And that's actually leading to the workload shifting from large pre-training more and more towards things like batch inference, which is actually a really, really horizontally scalable workflow that you can paralyze.

35:34You don't need interconnect in the same way. You don't even need state-of-the-art systems in the same way. And I think you're seeing that type of workload maybe grow in prominence. And so just to give one, just to unpack one more statement on that, one thing you can do that people are doing sometimes is you unroll the current state of the art model many many times basically doing chain of thought on the current state of the art model and then you take what required six steps with the previous state of the art and that becomes your training data for the next day of the art right that type of approach to generating data that then you'll that and then filtering that down to high quality examples and then training models on it looks very different than just throwing more poor quality data into a massive supercomputer to get the next generation.

36:17Yeah, it seems like that bootstrap up is really sort of under-discussed relative to as you hit a certain threshold of model, the rate at which you can increase for the next model just kind of accelerates. One other thing that'd be great to cover, I know we only have a couple minutes left of your time, is the recent paper that you authored, which I thought was super interesting around compound AI system design and sort of related topics to that. So would you mind telling us a little bit about that paper and sort of what you all showed? So I think it's kind of in this regime that we were just talking about where more and more often to go beyond the capabilities on Frontier accessible to today's state-of-the-art models and kind of get GPT-5 or GPT-6 early, practitioners are starting to do these things oftentimes implicitly where they'll call the current state-of-the-art model many, many times.

37:08There are many scenarios where maybe you're willing to expend a bit of a higher budget. Maybe it's code or something. And if I said that I can give you a 10 % better model for code, many developers might pay 10x for that, access to that. Instead of$20 a month, they might be very willing to pay$200 a month, for obvious reasons. And so there's a question of what do you do in that setting. And so people are, you know, if you're willing to call the model many times, you can compose those many calls into almost a network of network calls. right and you know I guess one of the questions is how then should you compose these networks of networks what principles should guide their architecture we kind of know how to construct neural networks but we haven't yet elucidated the principles for how to construct networks of networks so to speak these compound AI systems so to speak where you have many many calls maybe external components and so one principle that we start to explore was maybe one thing you can probe to figure out how to compose these calls or whether composing many calls will help you is you can look at how verifiable the problem is.

38:03And if it's verifiable, you can actually bootstrap your way to really high performance. So what does this mean? Well, verifiable means that it's kind of easier to check an answer than it is to generate an answer. And there are a lot of cases where this is true. Most software engineering computing tasks kind of classically have this property. We looked at things like prime factorization or a lot of math tasks. Classically, it can take someone years of suffering to write a proof and you can read the proof in a couple of hours. I think we've all had that experience with some training. And so there are many examples like this.

38:33And so one thing you can do is you can have models, you can embarrassingly parallel, you know, you can horizontally scale out and generate many, many candidate responses and then relatively cheaply check those candidate responses and kind of do a best of K type of approach. and turns out the judge model or the verifier choosing the best candidate response might actually have a lot higher accuracy at selecting the best candidate response from the set. Because you can kind of repeat this as a procedure to actually bootstrap your way to really, really high performance in many cases. And so, you know, we did kind of some preliminary investigations here and we were able to, in one case, the prime factorization, you know, kind of 10x the performance, go from 3.7 % to 36.6%.

39:16on prime factorization, which is pretty hard, kind of factor, you know, taking a number that's a composite of two primes, two three-digit primes and factorizing it, factoring it into the constituent primes. It's kind of a classic problem that pops up a lot in cryptography. And then also we looked at subjects in the MMLU and found that for kind of the subjects you would expect, math, physics, electrical engineering, this type of approach was really helpful. Now, we use language models. It doesn't have to be a language model. It could be a simulator. These could be unit tests as your verifier, et cetera.

39:45But we think this type of approach kind of points towards maybe a very different paradigm for getting better performance than just kind of scaling the models and doing a whole new pre-training from scratch. The MMLU performance bump was about 3%. And to put that in perspective, the gap between some of the previous best models is often less than 1%. Between, for example, Gemini 1.5 and Llama 3.1 and things like that. So 2.8 % or 3 % is actually a pretty major gap on MMLU. So pretty intriguing. And I think a lot of practitioners are hopefully going to explore this setting a lot more. It's a super cool paper.

40:18Do you have any creative ideas about how you could apply some of the ideas here to improve performance on more open-ended tasks? In many ways, we're not originating these. I think some of these are baked largely into systems like alpha code, alpha geometry already. I was pretty inspired to see the alpha geometry results recently as well. Yeah, I think that what we'll see people doing is kind of composing, this sounds funny, but massive networks where maybe each stage in the network will basically be maybe some best of K component with many, many calls to different language models. You know, Claude, Gemini, GPT-4, each with their own spikes in terms of capabilities.

40:59You know, kind of throw multiple of them at questions in many cases and then kind of choose the best response. You might also ensemble that with other components like classical heuristic-based systems and simulators, et cetera, and kind of compose large networks that may make millions of calls to answer a question. And I think that type of approach, it sounds kind of farcical right now, but I think it'll seem common sense pretty soon. We think it's a fairly interesting approach and we've seen a lot of interesting evidence that we'll speak more about pretty soon for things like code generation and agentic tasks in the code regime for things like design, chip design, for things like actual neural network design or network of network design even, funny enough, in recursive ways.

41:43So it's actually a really good, it turns out a lot of these problems that we care about have that property where it's verifiable and you can compose these systems and bootstrap your way to much higher performance than people might have imagined. So it seems pretty applicable downstream, but there's a lot of open questions, a lot of work to do further. And I think part of our hope is that the community will explore this more and that these types of workloads that are a bit more paralyzable will become more and more common. There'll be a lot more batch inference, a lot more synthetic data generation.

42:10And you won't necessarily need the big interconnected cluster that maybe only OpenAI can afford to do kind of cutting edge work in the future. Yeah, really cool set of ideas. And overall, a great conversation. Thanks so much for doing this, Jared. No, thank you, Sarah. And thank you a lot. Great to see you. Yeah, great to see you too. Find us on Twitter at NoPriorsPod. Subscribe to our YouTube channel if you want to see our faces. Follow the show on Apple Podcasts, Spotify, or wherever you listen. That way you get a new episode every week. And sign up for emails or find transcripts for every episode at no-briors.com.

From the publisher

In this episode of No Priors, hosts Sarah and Elad are joined by Jared Quincy Davis, former DeepMind researcher and the Founder and CEO of Foundry, a new AI cloud computing service provider. They discuss the research problems that led him to starting Foundry, the current state of GPU cloud utilization, and Foundry's approach to improving cloud economics for AI workloads. Jared also touches on his predictions for the GPU market and the thinking behind his recent paper on designing compound AI systems.

Sign up for new podcasts every week. Email feedback to show@no-priors.com
Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @jaredq_

Show Notes: 
(00:00) Introduction 
(02:42) Foundry background
(03:57) GPU utilization for large models
(07:29) Systems to run a large model
(09:54) Historical value proposition of the cloud
(14:45) Sharing cloud compute to increase efficiency 
(19:17) Foundry’s new releases
(23:54) The current state of GPU capacity
(29:50) GPU market dynamics
(36:28) Compound systems design
(40:27) Improving open-ended tasks

More from No Priors: Artificial Intelligence | Technology | Startups

All 169 episodes
The marketplace for AI compute with Jared Quincy Davis from FoundryNo Priors: Artificial Intelligence | Technology | Startups · 43 min
Listen in VO