In short
a16z Podcast Episode Notes: Chasing Silicon: The Race for GPUs
Episode Summary In the second part of a three-part series, this episode of the a16z Podcast delves into the challenges faced by founders in the AI sector, particularly focusing on the hardware shortages critical for AI applications. The discussion covers the disparity between the supply and demand of AI hardware, strategic decisions regarding ownership versus renting of hardware, and the role of open source in building AI companies. The episode features insights from Guido Appenzeller, a special advisor at a16z and an infrastructure expert.
Key Topics Covered
- Supply and Demand Dynamics
- Rapid growth of AI has led to a significant increase in demand for GPUs and other AI hardware.
- Industry sources suggest that demand may outstrip supply by a factor of 10.
- The current market is unable to accommodate the surge in demand due to:
- Bottlenecks in chip manufacturing.
- Limitations in building AI infrastructure.
- Access to AI Hardware
- Founders face challenges in accessing necessary compute resources.
- Options for obtaining hardware include:
- Pre-reserving capacity from cloud providers, often requiring long-term commitments.
- Negotiating exclusive deals with cloud providers to secure a stable supply.
- The question of who gets access is influenced by financial capacity and negotiation skills.
- Renting vs. Owning Infrastructure
- The decision to rent or own compute resources depends on the scale of needs:
- Small scale: Renting from cloud providers may be sufficient.
- Large scale: Owning infrastructure may become necessary for consistent access.
- Renting allows flexibility but can become costly at scale.
- Open Source and Competitive Moats
- Open source projects are emerging as viable alternatives to proprietary models.
- Access to unique training data can be a competitive advantage.
- Smaller models trained efficiently can rival larger ones due to lower overhead costs.
- Future of AI Infrastructure
- Trends suggest a move towards more decentralized computing as model optimization increases.
- Local running capabilities (e.g., on phones) may improve as technology advances, shifting some compute needs away from cloud services.
- The potential for a new stack of applications driven by neural networks opens up opportunities for innovation.
Insights from Guido Appenzeller
- The importance of understanding hardware needs and how they align with application goals.
- The need for founders to assess the right fit for their AI infrastructure based on operational demands.
- The challenges and benefits of navigating the competitive landscape of AI compute.
Looking Ahead
- The series will continue with a third part focusing on the costs associated with AI compute, covering:
- The financial implications of training large models like GPT-3.
- The distinction in costs between training and inference and how those costs evolve over time.
Conclusion This episode highlights the critical considerations for founders in AI regarding hardware access amidst a rapidly evolving landscape. It emphasizes the need for strategic decision-making in navigating supply shortages and cost implications, while also recognizing the potential for innovation within the constraints of current technology.
---
Additional Resources
- Find Guido Appenzeller:
- [LinkedIn](https://www.linkedin.com/in/appenz/)
- [Twitter](https://twitter.com/appenz)
- Follow a16z on social media:
- [Twitter](https://twitter.com/a16z)
- [LinkedIn](https://www.linkedin.com/company/a16z)
- Listen to the a16z Podcast:
- [Spotify](https://open.spotify.com/show/5bC65RDvs3oxnLyqqvkUYX?si=3E8B3qT9TyiwAHJ7JnaKbg)
- [Apple Podcasts](https://podcasts.apple.com/us/podcast/a16z-podcast/id842818711)
---
Disclaimer The content of this podcast is for informational purposes only and should not be taken as legal, business, tax, or investment advice.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00We currently don't have as many AI chips or servers as we'd like to have. How do I get access to the compute that I need? Who decides this? You're looking at some very large investment projects that take some time to adjust. We're rebuilding a stack. You can look at AI just as a new application, but honestly, I think it's probably better way to look at a different type of computer. With software becoming more important than ever, hardware is following suit. And with the world constantly generating more data, unlocking the full potential of AI means a constant need for faster and more resilient hardware.
0:38That is exactly why we've created this mini -series on AI hardware. In part one, we took you through the emerging architecture powering LLMs, from GPU to GPU, including having work, who's creating them, and also whether we can expect more law to continue. But part two is for the founders trying to build AI companies. And here we dive into the delta between supply and demand. Why we can't just print our way out of a shortage, how founders can get access to inventory, whether they should think about renting or owning, or moats can be found, and even where open source comes into play. You should also look out for part three coming very soon where we break down exactly how much all of this costs from training to inference.
1:24And today we're joined again by A16Z, Special Advisor, Guido Appenzeller, someone who is truly uniquely suited for this deep dive as a storied infrastructure expert with experience like. CTO4 Intel's data center group dealing a lot with hardware and the low level components. It's given me so if I think a good insight how large data centers work, what the basic components are that make all of this AI boom possible today. Despite working with infrastructure for quite some time, here's Gido commenting on how the momentum of the recent AI wave is shifting supply and demand dynamics. The biggest thing that is trading that is just the crazy exponential growth of AI at the moment.
2:04AI has been booming since middle last year. I think nobody expected how quickly it would move. That is just traded in demand, which at the moment the market can't fill. As a reminder, the content here is for informational purposes only. Should not be taken as legal, business, tax, or investment advice, or be used to evaluate any investment or security, and is not directed at any investors or potential investors in any A16Z fund. Please note that A16Z and its affiliates may also maintain investments in the company's disgust in this podcast. For more details including a link to our investments, please see A16Z .com slash Disclosures.
2:44In a recent article, Gido even stated that some reputable sources indicate that demand for AI hardware outstrips supply by a factor of 10. Here's some commenting on how that dynamic is impacting competition. We currently don't have as many AI chips or servers as we'd like to have. So for some of our portfolio companies, finding the compute capacity that they need to run the applications is actually a real challenge, right? There's a whole value chain behind that. It's a combination of many things. We have some bottlenecks on the chip manufacturing side. We have some bottlenecks on building the actual cards.
3:20These development cycles take some time. So it's a combination of factors. But probably the biggest thing that is treating that is just the crazy exponential growth of AI at the moment. Maybe this is a silly question, but what really is stopping companies like Intel, like NVIDIA from going and 10Xing their production? Like is that on the roadmap where we're just gonna see a lot more chips and we won't see this discrepancy between supply and demand or is there something more complex at play? It's a bit more complex because if you want to make a chip, the way you do it is you make it in the foundry, which are extremely large, extremely complex, Intel makes trips on their own foundries but most companies manufacture with Taiwan semiconductor, TSMC and they are capacity constraints.
4:01You often have to reserve capacity along in advance. There's different processes So you know, it might be for a certain process which you don't want to use this capacity But for another one that you do want to use they don't have the capacity and you could just say like well in that case There's just build more fabs, but building a fab takes you a couple of years and probably a couple of billion or 10 billion of an Investment so you're looking at some very large investment projects that take some time to adjust and that's sort of what prevents us from reacting more quickly While some countries are making major multi -billion dollar investments in new semiconductor production plants, aka fabs, these will take time to scale and there are also no promises, given that expertise is concentrated in a few companies.
4:44So with demand not subsiding, what does this mean for who gets access to the supply available? It doesn't sound like the demand is going to subside, especially because we see this really what seems like intrinsic relationship between the power of these models and then the compute that's thrown at them. And so if we do expect demand to continue, I guess the question that arises is, how is this demand allocated? So how does a company, let's say if I'm a founder today, how do I get access to the compute that I need? Who decides this? Is it just who's willing to pay the most or how is that supply being distributed?
5:23Yeah, there's some of that, right? The moment capacity is expensive, wherever you go, I try just to run some personal experiments, try to reserve an instance of the cloud service providers a few days ago, and they just didn't have any. It's like, no, not available. What we're seeing is that often in order to get access to the newer cards and new chips, right? If you wanted at scale, you have to pre -reserve capacity. So often these are negotiations between a company and a large cloud where you say, okay, I need this many chips, this amount of time. Well, they'll often ask for as they ask for a certain time commitment.
5:53So I'll be like, okay, we can give you this many chips, but we want you to sign basically that you get them exclusively for two years and you pay for that amount. I think OpenAI wasn't the news with that, right? Where you have investment deals, where for example, a cloud provider comes in and invests in a company and as a result, the company gets capacity. So we're seeing all kinds of deals being struck as with any skies resource, right? There's a lot of deal -making going on. It's not just a matter of getting access to compute. It's about ensuring you get access to the kind of compute tailored to your needs and cost is not the only factor here.
6:26What would you say in terms of the considerations that they should be keeping in mind? Really, how much should founders know about hardware? And again, selecting which hardware to use? I think the first question honestly I would ask is do you really need to consume the hardware directly or do you really just want to consume something that runs on top of the hardware, right? Let's take an example. If I want to generate images with a stable diffusion, for example, a mobile phone app or something like that. It might be easier to go to a SaaS company, like replicate, for example, essentially will host the model for you where you just pay for access to the model and they send you back the generated images and they will manage all the provision of computer infrastructure and they will find the GPUs for you.
7:10If you do want to run your own model, I think my number one advice will be to shop around, right? There's a fair number of providers, the large clouds. In my experience, I'm not always the best option, right? a few price it out, we've seen that the startups typically are more likely to go to specialized clouds like Corwee or Lambda, right? That is specialized in providing AI infrastructure to startups. Shop around, look at the different offers compare prices. And when you're shopping around, in addition to price, which I feel like is a major motivating factor, what other factors are there in terms of these other companies who maybe aren't the big clouds?
7:45How are they differentiating relative to one other? How are they standing out in that market. There's a whole sort of decision tree there, but the first thing is one thing often drives the decisions how much memory do I need in my cards, right? If I have a small image model, right, I might be able to work with a more consumer grade card, which is much cheaper, right, per hour, if I reserve it in a cloud versus if I, for example, train a full -hash language model, and not only need a card with the most memory I can find, but I probably want to have as many cards as possible in one server, because communication between them matters, and I may even care about with the networking fabric behind it, right?
8:19Some of the very large models you actually network constraints in terms of how quickly you can train them. So it really becomes a question of what's your objective? Is inference as a training? If it's training, how big is your model, right? And based on that, you figure out what the card is, what kind of server you need, what kind of fabric you need between those servers, and then you sort of can decide what the right fit is for your application. Even prior to this AI wave, compute was a major line item for many software companies. and the calculus of leaning on the easily accessible cloud versus bringing infrastructure in house was becoming an increasingly important consideration.
8:53Here is Gido touching further on that very calculus in today's era and where scale comes into play. Compute is expensive. It's a major line item for many companies and this is even before the AI revolution, you could say. Even more so today, yeah. So how do you think about, again, how that impacts different companies' bottom lines and whether they really factor that in to having their own allocated GPUs versus using something more like replicate. You really have to figure out what is the right fit for you and it probably depends a lot on the scale of which you need them right. If you need a lot, you'll frankly you have to preserve them.
9:28You have to have your own. There's just no way around that. You need a smaller quantity. You may be able to reserve them on a more short -term basis or you have various models where you can consume only while your application runs but at a higher price. So this really comes down to what kind of load do you have? I will be typically seeing if somebody is training, they're more likely to do a long -term reservation for a GPU because you want to make sure you have access to it. If somebody has more continuous workloads where availability is important, if I just do inference, but I want to make kind of a sense, I'm sure that if a request comes in, I can service it.
10:00I can never be down. They're probably going to reserve capacity as well. On the other hand, if I have more batch jobs, where it's like, this job runs an hour later, that's not the end of the world, then you probably think it will vary able capacity and just preserve it at heart. But it's really a conversation of what is a usage pattern, what is your demand pattern from that comes the best pick for the part of the work with. We've seen that companies even prior to AI have benefited from building their own infrastructure by basically bringing that in -house because before that they were renting and they were paying a lot to rent that compute.
10:34Do you think that will be a differentiator for companies is moving forward or how should founders be thinking about that relationship between owning the infrastructure and renting it? Owning the infrastructure comes with cost as well, because you need to know how people that run it, or you need to get money for the cat -x and so on. So my guess is that most early stage founders and probably even most mid -stage, late -stage founders are better off by renting capacity, renting a cloud or using consumer SaaS service. Right. There's a couple of exceptions. If you have really, really specialized needs, right, you may just not find anybody who has exactly the kind of hardware that you need.
11:13There might be some cases where you have geopolitical concerns, your data is just too sensitive, you need to run your own data center, and there's probably certain scale where it makes sense for you to run your own data center. But it's a pretty large scale. If you're spending $10 million a year, you're probably still under -quietly followed. You're spending $100 million a year on infrastructure that may be the reason to look into options for on data. But if everyone is competing for the same compute, are there other ways to stand out? Where's the moat here? You could say a moat is getting access to different training data, but that actually doesn't necessarily have to do with compute or money being thrown out the problem.
11:51It's getting access to differentiated data. If you have access to differentiated data, that could be a moat. I mean, it's a bit more subtle because, look, if you had an area where there's just not much public training data. That's probably right. There might be areas like in finance or so where that's the case. But for a large language model, it turns out that just making a larger model in training on more data has more benefits than just absorbing more knowledge. It also means that it's better in reasoning and understanding abstract context and answering really complex multi -stage questions and so on.
12:23So probably five to guess. I think the future will be that we'll still train on all the data we can find, right? And then maybe you find you, meaning you yourself do some additional training on a particular problem to remain with your private data. That makes sense. Right. So your first go to elementary school to learn reading and writing and then you go to your vocational training for the specialized job that you have to do the future. Another important question worth addressing is who can realistically compete? If compute is expensive, will all the largest most heavily capitalized companies win since they can build the largest models with the most data?
12:56or what role does open source play? As one of many emerging examples, Bacunia was created by fine -tuning Meta's Lama -1 model for chat. The cost of fine -tuning added only an additional $300, but the result is competitive with much larger models like Chatchy -B -T or Bart. So what might this example and a growing number of open source projects tell us about the future of open -elps. So first of all, in general, larger models, if they're everything else being equal, perform better. And so the really small open source models that we're seeing out there today, they're not yet at the level of the GPT 3 .5 or GPT 4.
13:41And there's actually a website that runs sort of regular bake -offs where they basically ask you to prepare answers. And it's used to be pretty clear that the large ones are still a little bit ahead. That said, we're making big advances there. And we're figuring out a couple of things. So one thing we've learned is there's something called the Chinchilla scaling loss that basically gives us an idea how does data correspond to a model size. And if we over train, so don't train as efficiently as we could, we can actually get potentially a smaller and better model. And so you can match the performance of a large model with a smaller model if you train it more.
14:12So that's interesting. That reduces model sizes. And the trend at the moment is, it makes slightly smaller models and train them more to get equal performance. The second thing is that when we talk about models for slightly different purposes, right, you have the base -large language models, all they're trained in, practically speaking, is completing text, right, literally how you train them as you give them text and say, guess the next letter and then you tell them, nope, that was wrong, or yes, that was right, I didn't, and thus we propagate the, how they predict. And they're really good at that completing text, right?
14:41That's not quite the same that you want from a chatbot or from a model that you can tell to do something. So there's usually another step afterwards which is called fine tuning for instruction following or for chat specifically. Where basically I tell a model lock if somebody asks you to come up with a list of to do like a list of steps how to make pizza right this is roughly what I expect you to answer right these models are very good in learning these things so whether you first train them just complete text and then you train them how to react to human requests and instructions So it's for the destruction fine tuning.
15:14And so Lama, for example, that was a Facebook model where they published the weights for researchers. And then some people took that. And they fine -tuned it, meaning they took a bunch of instruction for things to turn it into Alpaca, or Vikunya, which is a much, much nicer model in terms of interacting with it, right? For humans, much, much more useful. And so the biggest challenge at Mone, we have in the open source side, is there's currently no large open source LLM out there, right? GPT 3 was 175 billion parameters. There's currently nothing in that weight class. that's open source and that people could use to find you or to play without modifying.
15:46It is worth noting that since this recording, several more open models have been released, including Lama 2 with 70 billion parameters and an open license, unlike its predecessor Lama 1. Another 40 billion parameter open source model Falcon was released as well. Both of these are still dwarfed in parameters compared to closed models like OpenAI's GPT -3 at 175 billion parameters, or GPD4 at an estimated 1 .8 trillion parameters, although the latter is speculated to be a collection of multiple smaller models. However, parameter count is not the only driver of performance. For example, while LAMM2 has fewer models than GPD3, its performance is actually much better due to being trained on more data.
16:35In fact, LAMATU is currently comparable to GBT3's successor, GBT3 .5, the current default of Chat GBT. And as many of these models continue to get larger, we may see some models compress, becoming more efficient and enabling inference on your device. You already mentioned stable diffusion can run on your computer's GPU. Do we expect to see more of that because right now they are all hosted by these companies right, they're trained by these companies on their dedicated servers and then even if you interface with chat GPT, it's running that inference for you. Do they expect to see that change at all as compute becomes cheaper, maybe more decentralized, or how would you think about that?
17:20That's a really good question and we're speculating a little bit here, but my guess is we will, right? And we're seeing some of these smaller models getting pretty good, they run on your laptop or even your phone. We're starting to see stable diffusion implementations that run well on phones, which I would have never thought. They take a couple of 10 seconds to create an image, which is comparatively slow, but there's certain applications that it's acceptable. So my guess is as both the devices get faster and the models get more optimized, this will be a trend that we see more and more. And the future might just be part of the operating system to have a basic, large language model, a messy machine -race model.
17:56Maybe I'm off base here, but we've talked about how expensive compute can be and how ultimately that can be a major line item for companies. And I guess probably the model training will remain with those companies and not necessarily on folks devices. But in terms of the inference, I assume that's still a pretty significant cost. And in a way, if someone is able to run that locally, doesn't that disjoint the company from having to pay for that compute because it's running on, let's say someone's MacBook GPU? Oh, yeah, totally. I mean, look, if I can generate an image from my phone directly, all takes us some battery power because a little warm right on that's it right so that's a huge advantage at the same time there's probably going to be a little bit by -focation there on quality and parameters right you can run things locally but you can probably run them a lot better in the cloud right because you have a much bigger server there so it probably depends a little what you want to do right if I just want to have a better spell checker that checks my email or maybe just some simple completion that's perfectly fine I can run that on my phone on the other hand if I want something that is more right -of -good speech or summarize a complex text.
19:00They might be like, oh, that I'm going to run the cloud because it takes so many more operations. Hopefully this is getting your wheels spinning in terms of what can be built. And here is Gido speaking to how this presents a fundamentally new stack and what that means in terms of opportunity. It feels like this really is like this massive way of this renaissance of innovation. It's full of opportunity so I mean we're rebuilding a stack. You can look at AI just as a new application, but honestly, I think it's probably better way to look at a different type of compute. We traditionally built software by composing algorithms in a way that we understand well.
19:36And where the end result was programmed or so, bottoms up constructed. Now, we have a second type compute where we're just trying to launch a neural network. And the big advantages, we don't actually need to know how to solve a problem as long as the network can figure it out, right? The neural network can figure it out, we're fine. And that opens up a bunch of new applications, but it also means you need a completely different stack in terms of all the different pieces, right? You probably want vector DBS to retrieve context. You want different types of hosting providers that are good in hosting these models and providing them to us service.
20:06It's a whole like Cambrian explosion of creativity as a whole new ecosystem forming. I think it's a ton of opportunities to build back companies. I think that paints a pretty incredible picture of opportunity across the stack. And as many of these trends continue to progress, like supply and demand, the calculus of renting versus owning compute, clothes versus open source models. We look to part three of the series to answer a very important question. How much does all of this cost? We'll explore all this in -depth, including how much startups are really spending on AI compute and whether that's sustainable, how much it really costs to train a model like GPT -3, the difference in cost between training and inference, and how all of this will change with time.
20:53We'll see you there.
20:57Thank you so much for listening to Part 2 of our AI Harbor series. We spent a lot of time trying to get these episodes right, so if you are enjoying them, go ahead and leave a review or a tele -friend. We'll also have an animated video version up on our YouTube channel soon, but for now you can find some of our recent videos like my conversation with Waymo's chief product officer in a Waymo, or a conversation I recently had at the Aspen Ideas Festival, where we discussed the classroom of 2050. As always, thank you so much for listening.
From the publisher
With the world constantly generating more data, unlocking the full potential of AI means a constant need for faster and more resilient hardware.
In this episode – the second in our three-part series – we explore the challenges for founders trying to build AI companies. We dive into the delta between supply and demand, whether to own or rent, where moats can be found, and even where open source comes into play.
Look out for the rest of our series, where we dive into terminology and technology that is the backbone of the AI, how much the cost of compute truly costs!
Topics Covered:
00:00 – Supply and demand
02:44 – Competition for AI hardware
04:32– Who gets access to the supply available
06:16– How to select which hardware to use
08:39– Cloud versus bringing infrastructure in house
12:43– What role does open source play?
15:47– Cheaper and decentralized compute
19:04– Rebuilding the stack
20:29– Upcoming episodes on cost of compute
Resources:
- Find Guido on LinkedIn: https://www.linkedin.com/in/appenz/
- Find Guido on Twitter: https://twitter.com/appenz
Stay Updated:
Find a16z on Twitter: https://twitter.com/a16z
Find a16z on LinkedIn: https://www.linkedin.com/company/a16z
Subscribe on your favorite podcast app: https://a16z.simplecast.com/
Follow our host: https://twitter.com/stephsmithio
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Stay Updated:
Find a16z on X
Find a16z on LinkedIn
Listen to the a16z Podcast on Spotify
Listen to the a16z Podcast on Apple Podcasts
Follow our host: https://twitter.com/eriktorenberg
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
