In short
a16z Podcast Episode Notes: The True Cost of Compute
Episode Overview
- Podcast Title: a16z Podcast
- Episode Title: The True Cost of Compute
- Episode Description: Discusses the increasing importance of hardware alongside software, especially as AI utilization grows, and examines the costs associated with training AI models.
Key Topics Covered
- The Cost of Compute
- Training large AI models is increasingly expensive, often requiring millions of dollars.
- Current industry trend indicates that training costs may rise to tens of millions of dollars.
- Sustainability of Costs
- Discussion on whether the high costs of computational resources are sustainable for startups and smaller companies.
- Many AI companies are reportedly spending over 80% of their capital on compute resources.
- Computation Requirements
- Factors influencing model training costs include:
- Batch size
- Learning rate
- Duration of training process
- Complexity of the computational problems faced.
- Relationship Between Compute, Capital, and Technology
- The significant capital required raises questions about competition in the field.
- The expectation is that as hardware becomes more efficient, training costs may stabilize or decrease slightly.
- Model Training vs. Inference Costs
- Training costs are dramatically higher than inference costs.
- Example: GPT-3 has around 175 billion parameters, leading to an estimated training cost of several hundred thousand dollars depending on various factors.
- Future of AI Compute
- The challenge of scaling AI models with the increasing data limits.
- Prospects for cost reductions as technology advances.
Key Takeaways
- AI Model Training Expenses:
- Training modern AI models can exceed tens of millions, making it a substantial expense for companies.
- Companies need to reserve capacity and take into account inefficiencies in utilization.
- Inference Costs:
- In contrast to training, inference is relatively cheaper, costing only a fraction of a cent per query.
- Peak capacity provisioning can inflate costs during heavy usage times.
- Technological Advancements:
- The dialogue implies ongoing innovation will be necessary to sustain the AI boom and manage costs effectively.
- The relationship between the amount of compute and the performance of AI models indicates a need for continued investment in both hardware and software.
- Market Accessibility:
- The high costs of training may hinder smaller entities from entering the AI space, yet advancements in hardware and competition could lower barriers in the future.
- Data Utilization:
- There's an intrinsic relationship between model size and the amount of training data; larger models require proportionately more data for optimal performance.
Final Thoughts
- The podcast wraps up its AI hardware series and emphasizes the ongoing transformation of the tech landscape as hardware evolves alongside software capabilities.
- Encouragement for listeners to engage further with the content, including previous episodes and upcoming animated versions on YouTube.
Resources
- LinkedIn: [Guido Appenzeller](https://www.linkedin.com/in/appenz/)
- Twitter: [Guido Appenzeller](https://twitter.com/appenz)
- Follow a16z: [a16z Twitter](https://twitter.com/a16z), [a16z LinkedIn](https://www.linkedin.com/company/a16z)
- Subscribe: [a16z Podcast](https://a16z.simplecast.com/)
Disclaimer The content is for informational purposes and should not be considered as legal, business, tax, or investment advice.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:01There's very few computational problems that complex that mankind has actually undertaken. How do you think about the relationship between compute, capital and then the technology that we have today? Yeah, that's the million dollar question or maybe a trillion dollar question. The expectation at the moment is that the cost of training these models may actually sort of top out or even go down a little bit as the chips get faster but we don't discover our new training material as quick way. With software becoming more important than ever, hardware is following suit. And with the world constantly generating more data, unlocking the full potential of AI means a constant need for faster and more resilient hardware.
0:45But how much does all of this really cost? In this final segment of our AI hardware series, we tackle that question head on. But if you're just catching up, be sure to check out part one and part two, where we explored the emerging architectures and the momentous competition for AI Hardware. And today, we're joined again by A16z Special Advisor, Guido Appenzeller, someone who is uniquely suited for this deep dive as a storied infrastructure expert, with experience like. Intel's data center group dealing a lot with hardware and the low -level components. It's given me so if I think a good insight how large data centers work, what the basic components are that make all of this AI boom possible today.
1:31Here is Gido touching on the reality of these models and how much they cost today. Training one of these large language models today is not a hundred thousand dollar thing, it's probably millions of dollars thing. Practically speaking, what we're seeing in industry is that it's actually more than tens of millions of dollars thing. As a reminder, the content here is for informational purposes only. Should not be taken as legal, business, tax, or investment advice, or be used to evaluate any investment or security, and is not directed at any investors or potential investors in any A16z fund. Please note that A16z and its affiliates may also maintain investments in the company's discussed in this podcast.
2:10For more details including a link to our investments, please see A16z .com slash Disclosures.
2:20In Gido's recent article, navigating the high cost of AI compute, Gido even noted that access to compute resources has become a determining factor for the success of AI companies. And this is not just true for the largest companies building the largest models. In fact, many companies are spending more than 80 % of their total capital raised on compute resources. So naturally, this begs the question. Is this really sustainable? The core technology that you're building in the early days towards more complete product offering, right? There's just a lot more boxes to check and features to implement and all the administrative parts of the application if you're getting to the enterprise.
3:02So probably you'll have more normal software development that's not AI, right? The classic software development happening. You'll probably also have a larger head count of people that they have to pay. So at the end of the day, I would expect as a percentage that will go down over time, right? As an absolute amount, I think you'll be going up for some time just because this AI boom is still just an insingency. The AI boom has just begun, and in part two, we discussed how it's unlikely for compute demand to subside anytime soon. There, we also discussed how the decision to own or rent infrastructure can make a non -trivial difference to a company's bottom line.
3:38But there are other considerations when it comes to cost. Batch size, learning rate, and the duration of the training process all contribute to the final price tag. How much does it cost to train a model depends on the murier the factors, right? Now the good news is we can simplify this a little bit because the vast majority of models that are being used today are transformer models, right? That was a transformer architecture. Huge breakthrough in AI. They've proven to be incredibly versatile. They're easier to train because they parallelize a little bit better than previous models. And so in a transformer, you can sort of approximate the inference time as twice the number of parameters, floating point operations, right?
4:17And the training time is about six times the number of parameters. So if you take something like GPT -3, which is open AI SPIC model, they have 175 billion parameters. So you need twice as much. So 3 and 50 billion floating point operations to the one inference. And so based on that, you can sort of figure out how much compute capacity you need, how this is going to scale, how you should price it, you know, how much it will cost you at the end of the day. This also gives you for model training and idea how long the training is going to take. You know how much your AI accelerator can do in terms of floating point operations per second.
4:53You can sort of theoretically calculate how many operations it is to train your model. In practice, the math is more complicated because there are certain ways to accelerate that, so maybe you can train with a reduced precision, but it's also very hard to achieve 100 % utilization on these cards. If you naively implement it, you probably can work below 10 % utilization, but you can probably get into the tens of percent with a little bit of work. This gives you a rough swag. How much capacity you need for training and for inference, but at the end you probably do want to test it before you make any final decisions on these things.
5:24make sure that your assumptions hold on how much you need. Now, if all those numbers confused you, that's okay. We'll walk through a very specific example. GPT -3. GPT -3 has about 175 billion parameters. And here's Gido on the computation requirements for training the model and ultimately inference. That's when you're prompting the already trained model to elicit a response. So if you just do very naively the math, let's start with training, right? We know how many tokens it was trained on, how many parameters the model has, so we can do a soft napkin math, and you end up with something like three times 10 to the 23 floating point operations.
6:02That's a completely crazy number. It's like a number with 23 digits, so it's like hard to write down. There's very few computational problems, that complex, that mankind has actually undertaken. It's a huge effort. Then you can be like, okay, so let's take say an A100, the most commonly used card, we know how many floating point operations that can do per second, we can divide that. Let's give us an order of magnitude estimation, like how much time it will take. Then we know how much one of these cards cost, like renting an A100 cost to you between $1 and $4, probably, or depending on who you enter from.
6:36And you end up with something in the order of half a million dollars, right, with this very naive analysis. Now, there's a couple of things there, or we didn't take your account optimization. We also didn't take your account that you probably cannot run this at full capacity because of memory bandwidth limitations and network limitations. And last but not least, you probably need more than one run to get this right. Do you probably need a bunch of test runs? They're probably not going to be full runs and so on. But this gives you an idea that's of training one of these large language models today.
7:02It's not a hundred thousand dollar thing. It's probably millions of dollars thing. Practically speaking, what we're seeing in industry is that it's actually more for tens of millions of dollars thing. And that's because you need to reserve capacity. Right? So if I could get But all my cards for the next two months would only cost me a million dollars, but the problem is they wanted to a reservation. So really the cost is 12 times as high, and so that basically adds a zero to much, many cost. Right. And how does that compare to inference? So inference is much, much, much cheaper. Basically, my training said, for modern text model, for example, the training set is about a trillion tokens, right?
7:39And if I run inference, each word that comes out is one token. So a factor of a trillion or so faster on the inference part. If you run the numbers like a large language model, you actually add a fraction of a cent, like a tenth of a cent or a hundred of a cent, somewhere in that ballpark for the inference. Again, if you just naively look at this, for inference, your problem is usually that you have to provision for peak capacity. So if everybody wants to use your model at 9 a .m. on a Monday, you still have to pay for Saturday night at midnight when nobody is using it, that increases your costs substantially there.
8:10For some of them on specifically image models, what you can do for inferences is that you use much much cheaper cards because the model is small enough that you can run it on essentially the server version of a consumer graphics card and that can take a lot of cost. And unfortunately, as we discussed in part one, you can't just make up for these inefficiencies by piecing together a bunch of less performant chips, at least for model training. You need some very sophisticated software for that, right? because the overhead of distributing the data between these cards would probably outweigh any saving you get from cheaper cards.
8:43Infrance on the other hand? For inference, you can often do the inference on a single card, so that's not really a problem. If you take something like stable diffusion, write a very popular model for image generation that runs on a Macbook, for example, that has enough memory and enough compute power, so you can generate an image locally. So that'll run on a relatively cheap consumer card if you don't have to use an A100 for it to do inference. So when we're talking about the training of the models, clearly the amount of compute is just drastically more than the inference and something else that we've already talked about is the more compute often, not always but often the better model.
9:18And so does this ultimately these factors all ladder up to the idea that heavily capitalize and comments when this race or this competition or how do you think about the relationship between compute, capital and then the technology that we have today? Yeah, that's the million dollar question or maybe trillion dollar question. I don't know. So, first of all, training these models is expensive, right? For example, we haven't seen yet really good open -source, large language model. And I'm sure a part of the reason is that training these models is just really, really expensive, right? I mean, there's a bunch of enthusiasts out there who would love to do this, but you need to find kind of a couple of million or $10 million of compute capacity to do it.
9:57And that makes it so much harder, right? It means you need to create a substantial effort before something like that can happen. All that said, the cost for training these models overall seems to be coming down. And in part, I think it is because it seems to me like we're becoming data limited, right? So it turns out there is a correspondence between how big your model is and what the optimal amount of training data is for the model. It's having a super large model with very few data, doesn't get you anything or you get ton of data with a small model, also doesn't get you anything. You're talking the size of your brain needs to roughly correspond to the length of your university education here, right?
10:30I didn't know why I said, it doesn't work. What this means is that because some of the large models today already leverage a good percentage of all human knowledge in a particular area. I mean, if you look at GPD, there was probably trained on something like 10 % of the internet, right? And all of Wikipedia and many books, like a good chunk of all books, right? So going up by a factor of 10, yeah, that's quite possible. Going up by a factor of 100, that's not clear if that's possible. I mean, we as mankind just haven't produced enough knowledge that you could absorb all of that into one of these large models.
11:02And so I think the expectation at the moment is that the cost of training these models may actually sort of top out or even go down a little bit as the chips get faster, but we don't discover new training material as quickly. I mean, unless somebody comes up with a new idea of the general training material. And so if that assumption is true, I think this means that the mode that's created by these large capital investments. This is actually not particularly deep, right? It's more of a speed bump than something that prevents new entrance. I mean, today, training a large language model is something that is definitely within reach for a well -funded startup.
11:39I'd say, and for others, we expect to see more innovation in that area in the future. All right, that is a wrap for our AI hardware series. We genuinely hope you came away with a little more knowledge about this increasingly important space. Because if software is indeed eating the world, well, hardware is coming along for that ride. And as a reminder, if you haven't yet listened to part one where we explore the emerging architectures and who's creating them, or part two where we dive into the future AI stack and how founders can participate, well, those are already live and ready for consumption.
12:16And as always, thank you so much for listening. We'd actually like to leave you with a fun fact from GPT -4 itself, commenting on the technology that created it. And yes, we did fact check this, and this is also AI -generated audio from 11 Labs. We'll see you next time. Chat GPT and its sibling models are trained on diverse internet text. However, the exact amount of data used can be hard to comprehend. If we were to print all of the data used to train these models, it could fill a large library. Consider that one single book may contain around 1 million characters. If we estimate that the training data is hundreds of gigabytes of text data, let's take a conservative estimate and say it's 100 gigabytes.
13:02Considering that one character is approximately one byte, this would mean the model was trained on approximately 100 billion characters. If each book has 1 million characters, then the data used to train chat GPT is equivalent to the text in approximately 100 million books. If we take the size of a large library, such as the New York Public Library, which has around 53 million items, not just books, the training data is equivalent to the text in almost twice the number of items in that library. Thanks, Jaggie B .D., a quick note to close out that many models today are even bigger, with Lama 2, for example, being trained on two trillion tokens or about eight trillion characters.
13:44Now that is a lot of libraries. Thank you so much for listening to our full AI hardware series. We spent a ton of time trying to get these episodes right, so if you did enjoy them go ahead and leave a review or Tell a friend who would so appreciate that and you can also look forward to a video animated version of these up on our YouTube channel soon. But for now you can find some of our recent videos there like my conversation with Waymo's Chief Product Officer in a Waymo or or a conversation I had at the Aspen Ideas Festival where we discussed classroom 2050. As always, thank you so much for listening.
From the publisher
With software becoming more important than ever, hardware is following suit.
As the world generates more data, unlocking the full potential of AI means a constant need for faster and more resilient hardware.
But how much does this all really cost? In this final segment of our AI hardware series, we tackle that question head on.
Be sure to check part 1 and 2, where we explore the emerging architectures and the momentous competition for AI hardware.
Topics Covered:
00:00 – The cost of compute
02:20 – Is this sustainable?
03:23 – The cost to train a model
05:39 – Computation requirements
09:05 – The relationship between compute, capital, and technology
11:15 – GPT4 commenting on the technology with help from ElevenLabs
Resources:
- Find Guido on LinkedIn: https://www.linkedin.com/in/appenz/
- Find Guido on Twitter: https://twitter.com/appenz
Stay Updated:
Find a16z on Twitter: https://twitter.com/a16z
Find a16z on LinkedIn: https://www.linkedin.com/company/a16z
Subscribe on your favorite podcast app: https://a16z.simplecast.com/
Follow our host: https://twitter.com/stephsmithio
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Stay Updated:
Find a16z on X
Find a16z on LinkedIn
Listen to the a16z Podcast on Spotify
Listen to the a16z Podcast on Apple Podcasts
Follow our host: https://twitter.com/eriktorenberg
Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.
Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

