AI Hardware, Explained

27 Jul 2023 · 16 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

a16z Podcast Episode Summary: AI Hardware, Explained

Episode Overview In this episode of the a16z Podcast titled "AI Hardware, Explained," the discussion revolves around the critical role of hardware in the context of AI advancements, particularly in light of the generative AI boom. The episode emphasizes the importance of understanding the technology that underpins AI models and sets the stage for a deeper exploration of hardware in the upcoming episodes.

Key Themes and Concepts

  1. The Interplay of Software and Hardware:
  2. The episode opens with a nod to Marc Andreessen's famous quote about software eating the world, highlighting how AI software now requires equally robust hardware.
  3. The conversation emphasizes that AI applications depend heavily on hardware capabilities to perform computations.
  1. Understanding AI Hardware:
  2. Terms such as GPUs (Graphics Processing Units) and TPUs (Tensor Processing Units) are clarified, along with their roles in AI processing.
  3. The episode discusses how GPUs, traditionally used for graphics, have become essential for AI due to their ability to perform massive parallel operations.
  1. Market Dynamics and Supply Chain Challenges:
  2. The demand for AI hardware has significantly outstripped supply, leading to constraints faced by even established AI companies.
  3. The episode hints at future discussions on supply and demand mechanics, such as inventory access and the feasibility of "printing" more chips to meet needs.
  1. Ecosystem of AI Hardware:
  2. Key players in the chip market are identified, including Nvidia, Intel, and Google, with Nvidia currently holding a dominant position in AI hardware.
  3. The importance of a mature software ecosystem, particularly Nvidia’s CUDA, is discussed, which allows for optimized performance across AI models.
  1. The Future of Chip Architecture:
  2. The concept of Moore's Law is revisited, with an exploration of whether it is still applicable or if the industry is facing new challenges in chip design.
  3. Power consumption of chips is highlighted as an increasing concern, necessitating innovative cooling solutions.

Detailed Breakdown of Topics Covered

  • AI Terminology and Technology (0:00)
  • Chips, Semiconductors, Servers, and Compute (3:44)
  • CPUs and GPUs (4:48)
  • Future Architecture and Performance (6:07)
  • The Hardware Ecosystem (7:01)
  • Software Optimizations (9:05)
  • Future Expectations (12:23)
  • Upcoming Episodes on Market and Costs (14:35)

Expert Insights

  • Gido Eppenzeller: The episode features insights from Gido Eppenzeller, an infrastructure expert with a background in both software and hardware, particularly in the context of data centers and AI.

Key Takeaways

  • Importance of Parallel Processing: The performance of AI hardware is largely due to the ability to execute operations in parallel, which is a strength of GPUs and TPUs.
  • Demand vs. Supply: The podcast expresses that the demand for AI hardware exceeds the current supply, impacting the broader AI market.
  • Software-Hardware Integration: The maturity of software ecosystems like CUDA is critical for optimizing hardware performance in AI applications.

Conclusion The episode serves as an introduction to a mini-series focused on AI hardware, setting the stage for deeper discussions on market dynamics, cost, and supply chain challenges. The nuanced relationship between hardware capabilities and AI software performance is a central theme, pointing towards an evolving landscape in technology development.

Resources and Further Reading

  • Follow Gido Eppenzeller on [LinkedIn](https://www.linkedin.com/in/appenz/) and [Twitter](https://twitter.com/appenz).
  • Stay updated with A16z on [Twitter](https://twitter.com/a16z) and [LinkedIn](https://www.linkedin.com/company/a16z).
  • Subscribe to the A16z Podcast [here](https://a16z.simplecast.com/).

This episode provides a vital understanding of the underlying hardware that supports AI technologies today and the implications for the future.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00If you look at the pure hardware statistics, so how many floating point operations per second can these chips do? There's others that are very competitive, but Nvidia has. Are we now at the limits of photography? I think it's very surprising. Who would have thought that my gaming PC or my Bitcoin miner would eventually become a good AI engine? Power is becoming an issue. He just becoming an issue and we need to rely more and more on parallel processing. In 2011, Mark Andrewsson said software is eating the world. and the decade that followed just solidified this notion, with software infiltrating nearly every aspect of our lives.

0:37The last year in particular introduced a new wave of generative AI, with some apps becoming some of the most swiftly adopted software products of all time. And just like all the other software that came before it, AI software is fundamentally underpinned by the hardware that runs the underlying computation. So if software is becoming more important than ever, then hardware is following suit. Plus, the world is constantly generating more data, and unlocking the full potential of these technologies from longer context windows to multi -modality means a constant need for faster and more resilient hardware.

1:18And it's equally important for us to understand who builds and controls the supply of this resource, especially since many of even the most established AI companies are now hardware constrained. With some repeatable sources indicating that demand for AI hardware has stripped supply by a factor of 10. That is exactly why we've created this mini series on AI hardware. We'll take you on a journey through understanding the hardware that has long powered our computers, but is now the backbone of these AI models absolutely taking the world by storm. And in this first segment, we dive into the terminology and technology from GPU to TPU, including what they are, how they work, the key players like Nvidia competing for chip dominance, and also we address the question is Moore's law, death.

2:09But make sure to look out for the rest of our series where we dive even deeper, covering supply and demand mechanics, including why we can't just print our way out of shortage, how founders can get access to inventory, whether they should think about owning or renting, or open source plays a role, and of course, how much all of this truly costs. And across all three videos, we explore with the help of E16z Special Advisor, Gido Eppenzeller, someone who is truly uniquely suited for the Steve Dive as a storied infrastructure expert. I spent my last couple of years mostly in software, but most recently before joining Android's and Horowitz.

2:48I actually was CTO for Intel's data center group dealing a lot with hardware and the low level components. It's given me so I think a good insight how large data centers work, what the basic components are that make all of this AI boom possible today and that really underpins this great technology ecosystem. Gido has also spent time at UBICO, VMware, big switch networks, and more. But let's get into it. As a reminder, the content here is for informational purposes only. Should not be taken as legal, business, tax, or investment advice, or be used to evaluate any investment or security, and is not directed at any investors or potential investors in any A16z fund.

3:28Please note that A16z and its affiliates may also maintain investments in the company's discussed in this podcast. For more details, including a link to our investments, please see a16c .com slash Disclosures.

3:44We are increasingly hearing terms like chips, semiconductors, servers, and compute, but are all of these the same thing and what role do they play in our AI future? If you're running any kind of AI algorithm, I just AI algorithm runs on a chip. And the most commonly used chips today are AI accelerators, which are in terms of how they're built, they're very close to graphics chips. So the cards that these chips are on that are in these servers often refer to as GPUs, which stands for graphics processing unit. We're just kind of funny, right? They're not doing graphics, obviously, but it's a very similar type of technology.

4:19If you look inside of them, they basically are very good at processing very large number of math operations per cycle in a very short period of time. So very classically, like an old -fashioned CPU would run one instructions every cycle and then they had multiple cores. So maybe now modern CPU can do a couple of ten instructions. But these sort of modern AI cards, they can do more than a hundred thousand instructions per cycle. So they're extremely performant. So this is a GPU and these GPUs run inside of servers. They think of them as big boxes. I have a power plug on the outside and a networking plug and and then the server sit in data centers where you have racks and racks of them that do the actual compute.

4:58Let's quickly recap. CPU is central processing unit, and GPU is graphics processing unit. And while both CPUs and GPUs today can perform parallel processing, the degree of parallelization is what sets GPUs apart for certain workloads. So for example, CPUs can actually do tens or even thousands of floating point operations per cycle, but a GPU can now do over 100 ,000. The basic idea of a GPU is that instead of just working with individual values, it works with vectors or even matrices, right, or tensors more generally. TPU, for example, is Google's name for these kind of chips, right, and they call them tenser processing units, which is actually a pretty good name for them, right, the cores and these more than GPUs often call tenser cores, like that's how the GDI calls them because they operate on tensors.

5:47And basically, the core of their value propositions is they can do matrix multiplication. So if you remember metrics like the roles and columns of numbers, they can for example multiply two metrics in a single second. So in a very, very fast operation. And that's really what gives us the speed that's necessary to run these incredibly large language and image models that make generative AI today. Today's GPUs are far more powerful than their ancestors. Whether we're comparing to the earliest graphics cards in arcade gaming days 50 years ago, or the GeForce 256, the first personal computer GPU unveiled by Nvidia in 1999.

6:23But is it surprising that we're seeing this chip design applied so readily to the emerging space of AI or should we expect a new architecture to evolve and become more performant in the future? And more way, I think it's very surprising. Who would have thought that my gaming PC, my Bitcoin miner would eventually become a good AI engine, yeah? At the same time, what all of these problems have in common is that you want to execute many operations in parallel, right? And so you can think of a GPU or something was built for graphics, but you can think of them also just as something it's very good in performing the same operation at a very large number of parallel inputs, right?

6:59A very large vector or very large matrix. All right, so perhaps it's not so surprising that in videos, prize GPUs are aligned to this AI wave, but they're also not the only company participating. Here is Gido breaking down the hardware ecosystem. The ecosystem comes in many layers, right? So let's start with the chips at the bottom. And videos came off the hill at the moment right there. A100 is the workhorse that powers the current AI revolution that coming up with a new one called the H100 which is up there at the next generation. There's a couple of other vendors in this space. Intel has some called Gaudi, Gaudi 2, right?

7:34As well as that graphics card with ARC. They're seeing some usage, AMD has a chip in this space. And then we have the launch code. clouds that are starting to build or in some case have them building for some time their own ships, move over to the TPU you mentioned before, it is my popular and Amazon has a chip called Trainiam for training and inferencia for inference and we'll probably see more of those in the future from some of these vendors. But at the moment, Nvidia still has a very, very strong position as the vast majority of training is going on on their chips. And when we think about the different chips you mentioned like the A100s or the strongest and maybe there's the most demand for those.

8:08But how do they compare to some of these chips created by other companies? Is it like double the performance or is there some other metric or factor that makes them much more performant? It's a great question. If you look at the pure hardware statistics, so how many floating point operations per second can these chips do? There's others that are very competitive, but Nvidia has. Nvidia's big advantage is that they have a very mature software ecosystem. So imagine you are an artificial intelligence developer or engineer or researcher, or you're often using a model that's open source, that somebody else developed, and how fast that model runs, in many cases depends on how well it's optimized for a particular chip.

8:47And so the big advantage in video has today is that there's software ecosystem which has so much more mature, right? I can grab a model, it has all the necessary optimizations for Nvidia to run out of the box, right? I don't have to do anything, but with some of these other chips, I may have to do a lot more of these optimizations myself, that's what gives them these strategic advantage at the moment. So as we've touched on, AI software is heavily dependent on hardware. But what Gido is pointing towards here is the performance of hardware being heavily integrated with software. So, Nvidia's CUDA system makes it easier for engineers to plug in and make optimizations, like running with lower precision numbers.

9:24Here is Gido speaking to the kind of optimizations that do exist. And what does that actually look like in terms of those software optimizations? Like, what kind of developers are working on that? because that also seems to be maybe an emerging space where different companies are having the higher developers to actually facilitate that integration. Yeah, and it happens that all layers of the stack, some who just come in from academia, some of it is done by the large companies that operate in the space, some of them is trying to get them by enthusiasts that just want to see the water run faster.

9:52But to give you an idea of how this works, like for example, typically a floating point number is represented in 32 bits, right? And some people figured out how to reduce that to 16 bits and then somebody was like, well, actually, we can do it in 8 bits. And you have to be really careful how you do it. You have to normalize to make sure it doesn't overrun or underrun, right? But if you know what I said, we can use much, much shorter floats or integers for these calculations. There's many tricks like that that, you know, like the really good AI developers use to squeeze more performance out of the chips that they have.

10:22So to reiterate Gito's point, floating point numbers are typically represented in 32 bits. That's 32 zeros and ones, with the first bit being for sign, the next eight for the exponent, and the next 23 for the fraction. This gives a fairly large range between the smallest possible value and the largest possible value while also allowing many steps in between. Now, developers can choose to encode numbers in other systems with fewer bits, but the trade -off comes with precision. So depending on the numbers that you're working with, this may or may not have much consequence, but this This does require some checking and normalizing, plus an eye for overrunning.

11:04That's when you get a number so small or so large that it can't be properly encoded in the system. And just to give a sense for size, the range of 32 bit floats lies between 10 to the power of 38 and 10 to the power of negative 38. That's a pretty big range, while 16 bit floats operate in a precision range of 10 to the power of 4 and 10 to the power of negative 5.

11:30Now, when many people think of semiconductors, they naturally think of morsela. That's the term that describes the phenomenon observed by Gordon Moore, by the way back in 1965, where the number of transistors in an integrated circuit doubles every two years. But despite our collective success, for decades to continue to push more computation onto smaller chips, are we now at the limits of the photography? For example, an Apple M1 chip from 2022 has 116 billion that's billion with a B transistors, and if we compare that to the R1 processor from 1985, that had 25 ,000. And by the way, the Apple M1 chip is not even the highest transistor count today, I believe that belongs to the wafer scale engine 2 by Siriris with 2 .6 trillion transistors.

12:23So looking ahead, are we at the point where we really don't see the same kind of advancement and at least the physical architecture of chips? And if so, where do we see advancements moving forward? Is it in the software? Is it in the specialization of these chips? How do you see this industry moving forward? Yeah, great question. There's something to tease apart there. Like Moore's Law is actually still as of today alive and kicking, right? And Oslo talks about the density of transistors on a chip. And we're still increasing that, right? Now, the scale of transistors going down, I guess it's exactly the same speed, I don't know.

13:00But as of today, if you plot the curve, it seems to be intact. There's a second thing called denards scaling, which used to, basically say, it's just as the number of transistors I can squeeze onto a chip, right, doubles every 18 months or so. It essentially meant that the power, at the same time, would decrease by the same factor. It says something about frequency, but let's see the net outcompassist power. And that for the last 10, 15 years or so, no longer is true. If you look at the frequency of a CPU, it hasn't moved much over the past 10, 12, 15 years. And the net result of this is we're getting chips that have more transistors, but each individual core doesn't actually run faster.

13:41And what this means is we have to have lots and lots more parallel cores. And this is why these tenser operations are so attractive, right? Like on a single core, I can't add numbers more quickly, but if I can do a matrix operation instead And especially do many of them in parallel at the same time, right? The second big consequence of that is that our chips are getting more and more power If you look at even a graphics card for gaming PC today, right? You have these graphics cards like hundreds of watts of power to the 500 watt card, right? Which is much much more than they used to be and that's trying to just gonna continue and we're seeing What's happening data center seeing more and more things like liquid cooling at least being experimented with our in some cases getting deployed, where basically the energy density is for these AI chips, so it is getting so high that we need novel cooling solutions to make them happen.

14:27So Moore's law, yes, but power is becoming an issue. He just becoming an issue and we need to rely more and more on parallel processing. So it sounds like Moore's law is indeed not quite dead, but perhaps a little more complex than it once was. Performance increases continue as we integrate parallel cores, but we're also seeing chips become a lot more Power Hungry. All of this will continue being dynamic as demand continues to outpace supply for high -performance chips. So as we look ahead, what does all this mean for competition and cost? You'll learn a lot more about that in the rest of our AI hardware series tackling the questions that everybody is asking, including...

15:07We currently don't have as many AI chips or servers as we'd like to How do you think about the relationship between compute, capital, and then the technology that we have today? Yeah, that's a million dollar question, maybe truly a million dollar question. We'll see you there. Thanks for listening to the A16z podcast. If you liked this episode, don't forget to subscribe, leave a review, or tell a friend. We also recently launched on YouTube at youtube .com slash A16z underscore video, where you'll find exclusive video content. We'll see you next time.

From the publisher

In 2011, Marc Andreessen said, “software is eating the world.” And in the last year, we’ve seen a new wave of generative AI, with some apps becoming some of the most swiftly adopted software products of all time.

So if software is becoming more important than ever, hardware is following suit. In this episode – the first in our three-part series – we explore the terminology and technology that is now the backbone of the AI models taking the world by storm. We’ll explore what GPUs are, how they work, the key players like Nvidia competing for chip dominance, and also… whether Moore’s Law is dead?

Look out for the rest of our series, where we dive even deeper; covering supply and demand mechanics, including why we can’t just “print” our way out of a shortage, how founders get access to inventory, whether they should own or rent, where open source plays a role, and of course… how much all of this truly costs!

 

Topics Covered:

00:00 – AI terminology and technology

03:44 - Chips, semiconductors, servers, and compute

04:48 - CPUs and GPUs

06:07 - Future architecture and performance

07:01 - The hardware ecosystem

09:05 - Software optimizations

12:23 - What do we expect for the future?

14:35 - Upcoming episodes on market  and cost

 

Resources: 

  • Find Guido on LinkedIn: https://www.linkedin.com/in/appenz/
  • Find Guido on Twitter: https://twitter.com/appenz

 

Stay Updated: 

Find a16z on Twitter: https://twitter.com/a16z

Find a16z on LinkedIn: https://www.linkedin.com/company/a16z

Subscribe on your favorite podcast app: https://a16z.simplecast.com/

Follow our host: https://twitter.com/stephsmithio

Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.

Stay Updated:

Find a16z on X

Find a16z on LinkedIn

Listen to the a16z Podcast on Spotify

Listen to the a16z Podcast on Apple Podcasts

Follow our host: https://twitter.com/eriktorenberg

 

Please note that the content here is for informational purposes only; should NOT be taken as legal, business, tax, or investment advice or be used to evaluate any investment or security; and is not directed at any investors or potential investors in any a16z fund. a16z and its affiliates may maintain investments in the companies discussed. For more details please see a16z.com/disclosures.


Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.

More from The a16z Show

All 489 episodes
AI Hardware, ExplainedThe a16z Show · 16 min
Listen in VO