In short
Computer architecture and energy constraints, comparing RISC vs CISC, then CPU vs GPU vs Google TPU for AI/ML. Patterson argues CISC “winning” is misleading due to x86 software inertia, while modern workloads are energy bound and increasingly domain-specific.
Guest
David Patterson, Turing Award winner known for major contributions to computer architecture (with John Hennessy at Stanford; Patterson references their work and textbooks).
Key claims
Moore’s Law slowed because power scaling (Denard scaling) stopped around 2005, forcing multicore and later specialization. RISC prevailed in practice: ARM (from Acorn RISC Machine) dominates mobile and is growing in cloud. CISC’s complex instructions were often not used by compilers, making microcode overhead unjustified. GPUs/TPUs succeed by tailoring hardware and data formats to ML matrix multiply, not by general-purpose design.
Notable examples
1980s RISC vs CISC debates; Acorn RISC Machine → ARM → Apple Newton and later Nokia adoption; AlexNet 2012 on GPUs; Google TPU (2016) using matrix-multiply-centric design and “bfloat16,” reported as ~30x better inference than GPUs and ~80x than CPUs. MLPerf benchmark effort; CUDA “moat” via NVIDIA libraries.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding RISC vs. CISC
0:36 to 1:38
Explore the historic debate between RISC and CISC architectures.
“Maybe you could explain RISC versus CISC.”
The Evolution of Microprocessors
1:38 to 3:59
Discover how microprocessors evolved and the significance of Moore's Law.
“So what they were doing with Moore's Law was building more and more sophisticated instructions.”
Philosophical Perspectives on Instruction Sets
3:59 to 6:31
Delve into the philosophical debates surrounding instruction sets in computer architecture.
“So three or four X speed up was the potential for risk over CISC.”
Impact of ARM Architecture
6:31 to 8:02
Learn about the rise of ARM architecture and its implications in mobile devices and beyond.
“this one was called the early forerunner of the iPhone called the Newton, they came to this company and said, boy, I really like it.”
CISC vs. RISC in Modern Computing
8:02 to 9:27
Understand the current standing of CISC and RISC architectures in modern computing.
“while the risk architecture market is growing leaps and bounds.”
Microcode and Control Signals
9:27 to 12:52
Examine the role of microcode and control signals in computer architecture design.
“But it was worth it financially for that extra overhead because of the value of the PC software base.”
The Role of Compilers in Architecture
12:52 to 14:00
Learn how compilers impact the efficiency and performance of computer architectures.
“Can you explain the role of the compiler and how it manages the relationship between software and hardware?”
RISC vs CISC: A Fundamental Debate
14:00 to 16:40
Explore the key differences and implications of RISC vs CISC architectures in computing.
“Because register allocation by the compilers was so poor, they had to put in the ability for the programmer to give hints of what variables should go on registers.”
The Evolution of Processing Units
16:40 to 19:40
Learn how GPUs emerged as specialized processors and the impact of Moore's Law.
“there'd be a special instruction built in to do all this work that what the architects thought the compilers wanted and the compiler designers say, well, we don't need that.”
Denard Scaling and Its Impact on Computing
19:40 to 22:20
Understand Denard scaling and how it influenced microprocessor development and power efficiency.
“What happened is about 2005 or so, Denard scaling stopped working.”
Show all 27 chapters
Domain-Specific Architectures and AI
22:20 to 26:20
Examine the rise of domain-specific architectures, especially in relation to machine learning.
“I worked for Google, but I think it's fair to say Google was the one who really saw...”
The Game Changer: AlexNet
26:20 to 27:00
Discover how AlexNet revolutionized machine learning and demonstrated the power of GPUs.
“But Jensen Wang, his vision was try to get these teenagers in the basement who are playing the games to learn how to program these things.”
Designing the TPU for Machine Learning
27:00 to 28:00
Learn about Google's TPU design tailored specifically for machine learning tasks.
“So we had the potentially much more performance.”
The Rise of GPUs and TPUs in Machine Learning
28:00 to 30:06
Explore the evolution of GPUs for machine learning and the impact of TPUs.
“And within a few years, everybody switched over, and not only were they doing neural networks, but they're using GPUs.”
The Shift in Computing Architecture and AI Focus
30:06 to 31:58
Learn how the focus on AI has changed processor design and industry investment.
“like 30 times better at inference than the contemporary GPU and 80 times better than CPU by having this dedicated one.”
Moore's Law: Current State and Future Implications
31:58 to 35:34
Understand the implications of Moore's Law slowing down on technology advancements.
“How should I change it for this important application?”
Benchmarking in AI: MLPerf and Its Challenges
35:34 to 37:54
Discover the role of MLPerf in AI benchmarking and the industry dynamics involved.
“And now it's maybe 10 years between factors.”
Lessons from a Career in Academia and Technology
37:54 to 42:00
Gain insights on career lessons and advice for researchers and academics.
“But the main components of a big matrix multiply unit, the so-called high bandwidth memory, a vector unit to go with the matrix unit, the basic building blocks are still there.”
Learning from Bad Talks and Careers
42:00 to 43:39
Discover how ironic approaches to teaching can lead to valuable lessons.
“So the story is early in my career, my friend, a good friend and I said, how can we teach grad students how to give a good talk?”
Prioritizing Family and Happiness
43:40 to 45:55
Understand the importance of family and happiness over wealth in career choices.
“let your career overcome the needs of your family, but you just, you gotta just remember that.”
Focusing on What Really Matters
45:56 to 47:55
Learn about the importance of prioritizing important tasks over urgent but unimportant ones.
“Other career advice things is I remember I kind of woke up one morning and it was like God spoke to me.”
The Value of Relationships in Careers
47:56 to 49:49
Explore the significance of relationships over material successes in one's career.
“something like that and that's how you judge your life okay good luck you know it's pretty hard to make it, it's pretty hard to do.”
Courage and Standing Up for Ideas
49:50 to 52:46
Understand the role of courage in confronting challenges and defending ideas.
“why are these people, very wealthy people, so mad about things?”
Optimism in Relationships and Life
52:47 to 55:26
Learn how maintaining optimism can influence personal relationships and life outcomes.
“And so if it's a specious argument, even if a person is a leader of a company and it just doesn't make sense, I feel it's kind of the responsibility.”
Advice to My Younger Self
55:27 to 56:00
Reflect on the lessons learned from imposter syndrome and starting a career.
“And this applies to both, both partners in a relationship, not just one partner.”
Reflections on Imposter Syndrome and Professional Growth
56:00 to 58:10
David Patterson shares his experiences with imposter syndrome and its impact on his early career.
“I mean, I think when I got here, because, you know, the imposter syndrome, I was a UCLA graduate student and suddenly I'm a Berkeley professor.”
Addressing Harassment in Academic Conferences
58:10 to 58:43
Patterson reflects on the issue of harassment at conferences and his desire to have intervened.
Transcript
Automatic transcript. May contain errors.0:00Well, if Moore's Law puts a lot more transistors in a chip, how come they're not getting hotter and hotter?
0:05The Peterman Pod Host:This is David Patterson, Turing Award winner, famous for his contributions to computer architecture. And we discussed the historic RISC versus CISC debate. And he said, and CISC won. Well, that's a pretty myopic view. Everything is energy bound now. Register allocation by the compilers was so poor. And his thoughts on modern GPU and TPU architecture. What's the high-level differences between a CPU, a GPU, and a TPU? What they realized for machine learning and AI... Here's the full episode.
0:41The Peterman Pod Host:Maybe you could explain RISC versus CISC. Kind of like what that debate was and why was it controversial? What happened is in kind of... You know, the microprocessor was invented in the 1970s. But in the beginning, it was basically a toy. It would be something that would go in microwaves and things like that. Those of us who believed in Moore's law believed eventually the microprocessor would be the way we did all our computing, with the doubling of transistors every year or two. Eventually, there'd be enough transistors that on one of these microprocessors would be a serious computer. So the people who were designing microprocessors like Intel and Texas Instruments weren't really computer architects.
1:26So they just imitated what the big companies did. So leading companies like IBM, or at the time Digital Equipment Corporation, IBM for mainframes, digital equipment for what were called mini computers. They kind of drove the architecture, the instruction set design. So what they were doing with Moore's Law was building more and more sophisticated instructions. They had, you know, mini computers be the size of a refrigerator or a couple of refrigerators, and mainframes would be the size of many of those. So that's what they were doing with Moore's Heart Resources. And the philosophy at the time was that by having a more sophisticated instruction set, you kind of raise the level of abstraction closer to the software.
2:10So that had some inherent benefits over having something at a lower level. Now, from some of us who were in that area, like my friend John Hennessy at Stanford and I, we thought that wasn't necessarily the right thing to do. Compilers would go from programming languages down to the instruction set. So why couldn't compilers hold that? And then the question is, what's the right instruction set for this emerging microprocessor? And the prevailing wisdom of these more sophisticated instruction sets, complex instruction sets, and if we think of it like a vocabulary, it was like having lots of polysyllabic words in your vocabulary.
2:52And the alternative was going to be what we call the reduced instruction set computer, was having lots of reduced instructions like monosyllabic words. So you might guess that for a program to execute these instructions, if they're simpler, they'd take more of them. And if they're sophisticated, they take fewer, but the sophisticated ones might take longer to run. So kind of the question came down to what was that ratio? And in the beginning, in the 1980s, when these debates were happening, there was this kind of very vociferous debates, partly about what those ratios are going to do. But the big part was kind of philosophical.
3:30Weren't you doing damage to the software industry by lowering the instruction set level? So the gap between the programming languages and the instruction set was larger. That was kind of the ferocity of the debates. Well, after the dust settled a few years later and we started getting numbers, it turned out you needed like maybe 30%, 40 % more simple instructions to execute in a program. But you can run them four or five times faster. So the net was, you know, or five times faster. So three or four X speed up was the potential for risk over CISC.
4:05The Peterman Pod Host:You mentioned that, I guess, philosophical difference, like that gap. Why would that be controversial? I mean, unfortunately, I'd say computer architecture in the 1970s and even 1980s, a lot of it was kind of people designed this by using their intuition or using their gut. Their feelings was this is the right thing to do. Even the textbooks of the time were like catalogs. They'd say, here's a computer and just list all its features, and here's another computer listing all its features. It was pretty dissatisfying what was going on because these debates, it would seem like we ought to be able to decide this scientifically with numbers that are on there.
4:49But absence of that, it was more like a philosophical debate, like how many angels on the head of a pin, right? You could have arguments qualitatively, but you couldn't settle the arguments. And because there wasn't ways to settle the arguments quantitatively, you know, people would just argue about it.
5:06The Peterman Pod Host:Would you say that, you know, RISC versus SISC, did one side win the war? Yeah, I was just reading there's some guy online who who revisited it, and he said, and CISC won. Well, that's a pretty myopic view. In the PC era, because of the importance of distributing software in binary, so once the x86 was established and people in the PCs would ship software in binaries, that was very hard to overcome. That was a huge impediment to changing the instruction set. So PCs are defined by the x86 architecture largely. But also in the 1980s, there was this company in England that wanted to do a personal computer.
5:53It was the Acorn personal computer. And they decided they needed their own instruction set of architecture to do that, their own chip. The chips they had available at the time they were doing it weren't fast enough. And they were influenced by the papers that we did at Berkeley. And so they built what they called the Acorn risk machine. And then one of the benefits of this kind of reduced instruction set is that it could be simpler. It would take less resources and take less energy to execute. And then several years later, when Apple was looking for a microprocessor that could power one of their personal devices, this one was called the early forerunner of the iPhone called the Newton, they came to this company and said, boy, I really like it.
6:40Let's get rid of the Acorn name. So they renamed it the Acorn Risk Machine, ARM, and they rechristened it the Advanced Risk Machine to get rid of the Acorn. And then Apple used it in the Newton. Now, the Newton wasn't a commercial success, but it demonstrated the benefits for risk architecture for mobile devices. So the Nokia came along just a few years later with their GDM cellular phone, which is one of the first popular ones, and they embraced ARM. And so ever since, you know, ARM has dominated all the mobile devices. So I think I just checked. There's been 350 billion ARM processors, microprocessors with ARM technology in it today.
7:24So it's like today 99 % of all processors in computers are a risk. Even in personal computers, Apple switched over to ARM from the x86 architecture. So even PCs, there's risk architecture is significant. And it's starting to get in the cloud. The cloud's largely been defined by the x86 server architectures. But Amazon and I think Amazon, Microsoft, and Google all have their own development of ARM processors. So ARM is becoming, or risk processors, becoming more popular in the cloud. So right now, I'd say the x86 architecture market is shrinking. while the risk architecture market is growing leaps and bounds.
8:12The Peterman Pod Host:Well, how could someone say CISC 1 when you mentioned like 99 %? Yeah, well, that's just if they're doing something in history and they define computers as personal computers and maybe servers and they go up through around 2000, if the story ended then, well, it looks like a CISC 1. But once we get into this post-PC era, I don't understand how somebody would reach that conclusion. You mentioned the energy expenditure. So risk makes sense in places or maybe on mobile devices, things like that. Well, even in the cloud, everybody cares about... Everything is energy bound now. So these days, of course, you're not dealing with millions of transistors, but billions of transistors.
9:05So it matters somewhat less today. There's so many things going on that you can hide that. I mean, even what the x86 architecture did is to compete is that it translated the x86 instructions in hardware into risk instructions. And so you had to pay that extra overhead of that translation step to get risk instructions. and then you could X any good ideas that the risk people had that the X86 could do. But it was worth it financially for that extra overhead because of the value of the PC software base. So it made a lot of sense for Intel to do that, and they did that in the early 2000s. It was a great commercial idea.
9:49The Peterman Pod Host:Is there any niche use case where CISC makes sense? Is this an engineering tradeoff or CISC is objectively worse? Yeah. So if we want to go kind of one level deeper into all of this, what was actually going on in designing computers, the hard part is the control. And so what happened is in the beginning, control was kind of ad hoc. You would figure out, you put the gates together to make it to work. one of the computing pioneers, Maurice Welks, figured out a more elegant way to design control. And he said, well, we could just list all the control signals as the output of a memory, and we could have something would keep track of where we were in the memory and to issue those control signals.
10:40And he called this effort, the instruction, or the control signals, you could think of instructions. So he called that a micro-instruction. And he called the programming of those instructions microprogramming. So where technology was in the 1960s, that made a fair amount of sense. And so IBM built these so-called micro-programmed computers. So it was basically an interpreter with very simple instructions that would interpret this much more sophisticated instruction set above it. But you would play this interpretation overhead. And classically, in computer science, interpreting versus compiling is something like a factor of 5 or 10.
11:23But given the latencies of the memory technologies and the possibility of doing this out of read-only memory, it made sense up until the 60s and 70s. But then the question came up as it came around 1980 is like, is this still a good idea? Should we have this microcode interpreter inside there? And so the alternative is to think of, well, rather than we've got this microcode interpreter in there, why don't we just compile directly into those instructions? And that's pretty close to the risk ideas. The microinstructions themselves used to be, you know, like 100 bits wide and really complicated. So if you make them not quite so long, you make them kind of more natural, still we could skip the interpretation step.
12:10So, going this level deeper, kind of the question would be today, would people invent an instruction set that was so sophisticated and needed a microcode interpreter? And probably they wouldn't do that. I mean, you could do it. Nothing prevents you from doing it, but you wouldn't want to design an instruction set that was forced to use a microcode interpreter. It might make sense in some very tiny applications, possibly where you have tens of thousands of transistors. Maybe this would micro-coded interpreter work, but I think nobody today, I don't think anybody's invented an instruction set in the last 20 years that has anything that would need a micro-coded interpreter.
12:51The Peterman Pod Host:You mentioned the compiler multiple times here, and it seems like that's a critical piece that kind of makes risk work so well. Can you explain the role of the compiler and how it manages the relationship between software and hardware? Well, you're writing in a programming language like C or C++ or maybe Python. But the quality of the code that gets generated is up to the compiler. A specific example is that it's useful when you construct computers to have registers. And those registers are actually kind of visible in the assembly language or the machine language programming. There can be eight, 16, or 32 of these registers for the people programming that they'll be able to use.
13:36Well, it used to be very difficult for compilers to allocate registers efficiently. Computers just weren't fast enough. We didn't have the algorithms that we could look at a section of code or a subroutine or something and say, how can we most efficiently do registers? In fact, the C programming language, which was invented to do systems programming in C, before that, to write an operating system, people wrote them in assembly language, believe it or not. But Unix people showed, Ken Thompson, Dennis Richey showed that if we had a low-level language, pretty low-level language, we'd get the benefits of writing in something that's much easier for humans to understand and debug.
14:21bug. Because register allocation by the compilers was so poor, they had to put in the ability for the programmer to give hints of what variables should go on registers. It was just easier for them to have the programmer step in and say, if the machine had eight registers, I want these six variables to be in these registers. Don't leave them in memory because it runs so much slower. Registers are so much faster. So, a big part of the risk-cisk argument was that compiler algorithms were getting better. They could handle these low-level destructions. They could allocate registers efficiently. And that was another reason in these debates about why it made sense to have a simpler architecture.
15:04What we did in the risk architecture is, well, if registers are really important, register allocation are really important, one way to make it easier for the compiler was just to have a lot more of them. So typically, the CISC architectures at times would have eight or maybe 16 registers. So we put in 32, and that was one of the arguments, right? Is, well, if it's hard to efficiently use a small number, let's give them plenty. So even if it wasn't that good, there'll be enough registers. So most of the time, and registers weren't that much more expensive to include in machines because of Moore's law.
15:37The Peterman Pod Host:I saw somewhere in one of the talks that he'd given that somehow the compiler, it's more easy for it to optimize the code for RISC, but in CISC, there were these complex, bigger ones, and it almost never used them. Is that a shortcoming of the compiler? So what happened, if we go back in the compiler days, there's this argument is that by having more sophisticated instructions, this would raise the level of extraction. There'll be a smaller gap, they'll make it easier to compile it. But that was a philosophical argument. It wasn't something that necessarily compiler people could work. Compiler people weren't making that argument.
16:18It was the architects who were making that argument. And as it turned out, when we looked at the programs when we were doing the early research on CISC versus RISC, is the compilers didn't really use those instructions. Often the compiler writers, the architects would come up with a sophisticated instruction and the compiler writers would say, well, we don't need that. In fact, we found examples where like for a procedure entry, there'd be a special instruction built in to do all this work that what the architects thought the compilers wanted and the compiler designers say, well, we don't need that.
16:49It's faster. It's actually faster to use it with separate instructions than you use your sophisticated instruction. And so we found a bunch of examples like that. So you pay the extra overhead of the microcode interpreter to have these sophisticated instructions. And then the compiler doesn't even use it. Right. So this is kind of a nonsensical situation that we're in. And, you know, this is kind of like, you know, why are there startups or why are there scientific, you know, breaking points like that? But as we looked at all the technologies with Moore's Law, ideas like caches, where compilers were using sophisticated instructions, the ability to allocate registers more efficiently, it made sense to change the directions away from these microcoded instruction set architectures to these simple architectures.
17:38The Peterman Pod Host:When we talk about instruction sets, I mean, we're kind of assuming that we're talking about these general purpose computer, these CPUs. And I know GPUs are kind of talked about a lot, and maybe there's other forms of computing. Are there instruction sets for those types of machines as well? So what happened is around 2000 is when GPUs came along, and these were what we call domain-specific architectures. So a GPU is a graphics processing unit. It had one job. It didn't need to do everything that a general-purpose processor needs to do. It didn't have to support virtual memory. it didn't have to support a compiler, which is a pretty radical idea for architectures.
18:20It was just for graphics. And so around 2000 is when NVIDIA and the other GPUs out there, because they were trying to give you the graphics for games and give you the graphics for movies. So it was a niche product. So what happened kind of in the computer industry is it was driven by Moore's Law, which I've mentioned many times, and also this lesser known law called Denard scaling, because kind of an interesting question, well, if Moore's law puts a lot more transistors in a chip, how come they're not getting hotter and hotter as we double the number of transistors every year? And the reason was this observation made by Bob Denard that as you added more transistors, people would also lower the threshold voltage, which is the distinction between a zero and one, and that had like a squared effect.
19:12So you would end up doubling them transistors, but you'd lower the threshold voltage. So microprocessors stayed at like 20 or 30 watts. In fact, we did a, John and I eventually did a textbook and I was just checking and it came out in 1990. We didn't even talk about power as an issue for the first three editions. You know, the third one came out in 2000. Power wasn't even a topic because Mark Denard scaling was going on. And so microprocessors stayed in the tens of watts even as they got faster and faster. What happened is about 2005 or so, Denard scaling stopped working. And that was a shock.
19:50So Intel actually had a microprocessor that failed one of their generations because they just couldn't get the power down enough. It was too hot to be able to do that. And then, so then what that forced us to go to multicore is that before that, it was easiest for the programmer for everybody if there was one very sophisticated processor that did everything, but we couldn't do that anymore. So it went from one sophisticated processor to two and then four and then eight simpler processors. Then it was up to the programmer to deliver on the potential of Moore's Law by paralyzing their code. So we stayed that way for about another 10 years, and then Moore's Law started slowing down.
20:36So the general purpose microprocessor was barely improving. It got a little better, but it wasn't getting dramatically better. In the 1980s, 1990s, 2000s, you have a laptop, and your friend's laptop would be like four times as fast as yours. And you were jealous, so you would throw away perfectly good hardware because your friend's thing was so much faster because of the rapid change in performance. Well, that all ended in the 2010s where you wouldn't throw away a laptop because the new ones were hardly that much faster than the old ones. You'd throw them away when they break and slow down. Nevertheless, programmers were used to this dramatic improvement in performance every few years because they could add more features to their software and stuff like that.
21:26So what were architects going to do? They'd already done the multi-core trick in 2005 or so. So in about 2015, the idea was we would do domain-specific architectures like the GPUs. So if you tell me I only have to run a narrow class of programs and I don't have to necessarily run all the operating systems, everything else. Well, yes, I could shuffle those resources and do something much more efficient for something well. So you could do some things well, and as other things, either you do poorly or not even at all. So then the question is, okay, what domain? So just coincidentally, where we were technologically, right around 2012, 2015 is when machine learning AI burst on the scene.
22:13That is evident what domain should we do, and the domain was machine learning AI. Just coincidentally where we were technologically, this new domain came along. I worked for Google, but I think it's fair to say Google was the one who really saw... It was certainly the first big company who understood the potential of machine learning AI and bet that they or feared that they were going to be swamped with demand and needed to do custom hardware. So the Tensor Processing Unit GPU that Google debuted in 2016 really kind of shocked the world and got people to realize that we should be designing hardware for machine learning, and we could continue to improve performance dramatically for that one domain.
23:07The Peterman Pod Host:What's the high-level differences between a CPU, a GPU, and a TPU? The CPU has got to be this general purpose thing, and basically, even today, they have these general purpose cores that each core is very similar to what the instruction sets used to be like 20 years before that. It's not a surprising design and it's just got lots of cores. And a modern ones today could have 50 or 100 cores in there. So that's the feature. For graphics, what they decided to do, which to do the graphics job is they need to get a lot of performance of the memory system. So they went to multi-threading. So they have hardware threads that you would send a memory request out and the hardware would switch over to do something else while the memory is coming out.
23:57So it's this highly threaded architecture and that's what they developed and could run well for graphics. And graphics it turns out doesn't need very, doesn't need a powerful floating point, it needs, doesn't need wide floating points. So 32-bit floating point is plenty for graphics, even 16-bit. So they were doing, pushing graphics with 16 and 32-bit floating point in this multi-threaded kind of architecture. And because it was kind of its own unique thing, it has its own set of terminology all to itself. It wasn't kind of out of the main branch of computer architecture. So when you look at our textbooks, we have kind of like a Rosetta Stone, which says, here's the terms that NVIDIA uses to describe GP2s.
24:45This is what it means in kind of normal, well, normal in the mainstream processor design. So they were this niche product that happened to do pretty fast single precision floating point and even half precision floating point. And they were pretty cheap. They were like hundreds of dollars. So when people, so kind of, but some people were thinking, boy, for some applications, if I could turn my program into like pretend that it was doing, creating an image, if I could turn my problem into image generation, I could use these pretty cheap GPUs, which had very good floating point performance per dollar compared to anything else.
25:31And so people started playing around with that. And the founder of CEO, Jensen Wang, really liked that idea. So in 2006, he funded an effort to create a programming language that would kind of handle this multi-threaded hardware architecture that's for graphics to make it easier to program. That led to the invention of CUDA, which I can't remember its acronym. It's really a proprietary programming language for the multi-threaded GPU architecture to make it easier to program. It was a C-like language, but you can't just compile C programs and run it, but it was C-like. People actually liked it.
26:14It was certainly much better than trying to turn your program into an image generation, but people liked it. But Jensen Wang, his vision was try to get these teenagers in the basement who are playing the games to learn how to program these things. And he had a couple of markets that he was interested in, like some of the Department of Energy labs, like fluid flow and some of the things. And he would build special purpose libraries that could handle that domain and then use his GPUs to do that kind of special purpose computing but breaking out of graphics. But that's kind of the heritage there.
26:51Now for the TPU program, and so what's happened, people started right from the very beginning, people started using GPUs because they had much better single precision floating point performance than the CPUs. They had a lot more processors on them. So we had the potentially much more performance. So kind of one of the breakthrough moments was in 2012 when within the machine learning community, this neural networking piece, which had a few advocates, but a lot of people didn't believe in it. They went into a competition to see who could do the best image recognition. In 2012, the so-called AlexNet beat all the competitions.
27:33This was this historic moment in machine learning and neural networking. And once they did it, and the guy who did that had taken a coup de course at the University of Toronto to learn how to use it. And so he says, well, as long as I'm doing this, I'll do it on a GPU. And so he was able to explore a lot more space on this cheap, fast GPU. So he entered it in the competition 2012. And he was the only one using neural networks, and he crushed the competition. And within a few years, everybody switched over, and not only were they doing neural networks, but they're using GPUs. So GPUs have this heritage of a graphics engine, but more programmable, and it started getting used for machine learning.
Read the full transcript
28:21When Google came along, they decided it was a clean slate. They didn't care about graphics. So at the heart of neural networking is a matrix multiply. That's the thing. So they designed a processor for at the time, microprocessor at the time, had a giant matrix multiplying unit. That was the main thing. And then they threw a bunch of stuff out that they didn't need. So a lot of general purpose computing is what's on the chip is maybe three levels of caches to try and have. So you don't spend all your time going to the relatively slow memory. Well, for machine learning, they knew when their memory accesses were so they could schedule that.
28:59So a hardware cache didn't make any sense. They would just have a memory that the software understood and would transfer in time. So those are some of the innovations. Also, they innovated on the floating point format. You didn't need hot scientific computing, cares a lot about precision. Most of it's done in 64-bit floating point where the exponent is less than 10 bits and most of it is the precision. you know, the fraction can be 50 some bits. Well, what they realized for machine learning and AI, they don't need all that precision. They needed the range. So, Google did the first floating point format where the exponent was bigger than the fraction.
29:40That was a radical idea, that so-called brain float 16, the part of Google that was doing this was the brain research group. So, they brought out this architecture. They had this narrow floating point format. They had big matrix multiply unit. It only had one processor in it, unlike the other ones. It was just dedicated for machine learning and it just kind of blew the doors off of everybody in the field. It was like 30 times better at inference than the contemporary GPU and 80 times better than CPU by having this dedicated one. So when Google made this announcement at their annual retreat a year or so after they had deployed it inside, it just shook everybody up.
30:26Intel started buying companies. NVIDIA started modifying the design to be much more enhanced machine learning. And then a bunch of other competitors started, hyperscalers started doing their own efforts. So I think that was the watershed. I think the TPU announcement was the watershed moment there.
30:47The Peterman Pod Host:So it's almost like levels of specialization, like the CPU is the most general. Yeah, if we could still, you know, if we still had Denard scaling, you know, Moore's Law and Denard scaling, where would we be today with general purpose processors? We should have 100 terahertz, you know, microprocessors. If we could build 100 terahertz microprocessors, that's what we'd do. GPUs would still be a niche, you know, if we could do that, you know, it raises all boats, that would be fantastic if we could do that. But that's long in the history. We haven't been able to do that for 20 years. And so there were probably two or three gigahertz microprocessors in 2005, and that's kind of where they barely improve today.
31:33So we can't do that. So there's still areas where it's important to have CPUs. They're the right, you know, you need operating systems, compilers and things. They're the right solution. But things have been specialized. And then now in terms of industry terms, huge emphasis is being put into the AI accelerators or just AI in general is where the money is being invested. And so the question for any processor designers was, how does my processor design fit into this AI universe? And what should I do? How should I change it for this important application?
32:12The Peterman Pod Host:So you mentioned that kind of Moore's Law is slowing down. And I watched this other talk from Jim Keller saying that Moore's Law isn't dead, but is it controversial to say that it's slowing down? No, I mean, not to real engineers. Moore's Law is very simple. It says in a chip, the number of transistors will double, originally said every year, and then he minuteed it to every two years. Just look, do the chips, are the number of transistors doubled? And no. Now, I think what people assume that means if we are no longer on Moore's Law, that technology is not improving. That's not the same thing.
32:54It's not improving at the rate that Moore projected. And what was amazing about Moore's Law, which lasted 50 years, is that it guided the investment of semiconductor manufacturing. We need to deliver on doubling transitions every year or two. How are we going to do that? How are we going to build the equipment to do that? So it was a guideline for the whole industry, which was kind of remarkable. So what we are doing now is there's pieces of the technology to get better, and there's pieces that don't improve at all. So one of the big pieces, important pieces on a chip is the static RAM, SRAM. So that's hardly improving at all.
33:35logic gates still continue to improve that they are the gates are getting better so if you're doing like uh adders or multipliers those are getting there so pieces it's not uniform improvement anymore and we're also going to more exotic packaging to be able to deliver it uh so people not uh it you know forever it was a single ship was the best way to package everything fits on one chip and that's what we're going to build. And now there's these ideas of chiplets or packaging multiple chips together, the latest GPUs. And I think the latest TPUs has actually two, what are called full reticle design dies package put into a package.
34:18And, uh, so a full reticle design is, uh, the maximum you can build on a, well, on a semiconductor, it, they have a step and repeat motor and there's, you know, in the old days, you'd have many chips inside that. Now there's just one. So two maximum reticle designs form a node. So they're using packaging. So if you look from outside, you can say, well, look how many transistors. It's, quote, Moore's law is continuing. But if you see what's inside the chip, that has tapered off. And I think part of it is for a fair amount of the industry. If you were in a semiconductor manufacturer and people ask you what you do is I make Moore's law.
34:58I sustain Moore's Law what I do and if you've been doing that for decades and somebody says Moore's Law is over it's like my career is over so I think there's an emotional side of it but you know just look at the data the data doesn't back up what Keller says but I also didn't say that the technology is not improving the technology is continuing to improve and specifically domain specific architectures but if you look at the general purpose architectures or even other memory technologies. You know, it used to be the DRAMs would improve by a factor of four in density every three years like clockwork.
35:34And now it's maybe 10 years between factors. So you can see plenty of evidence that Moore's law no longer applies.
35:41The Peterman Pod Host:Is there some new version or some analog to Moore's law that is kind of guiding, you know, microprocessor design and architecture these days? I guess a question is both NVIDIA and Google are continuing to deliver much faster processors for machine learning, tremendously better. What are they doing? Part of it is what I said about packaging, you know, to be able to get more transistors and you keep them closer together because distance matters. Part of it is innovating on the floating point formats. So, you know, unlike the supercomputers of 64-bit, it's not even 64-bit, not even 32-bit, not even 16-bit, but 8-bit and 4-bit floating point is going on.
36:33So narrowing of the data types. The, you know, having kind of what's called the matrix multiply instructions. So these powerful units that are put in there that can do special purpose applications that are very important. making those bigger and faster. So those are the things that are going on, but we don't have this simplifying guideline kind of underlying all this. You have to be aware of each piece of the technology, gauge how fast it's improving, whether they can deliver on what's going on, and then you assemble that together and make your bets. There's a paper that we've just got approved that I think we're going to put on archives soon, which is talking about the Google TPU line.
37:24And pretty remarkably, the Google TPU line, particularly for training, has stayed pretty constant in the basic architecture design. Things have gotten bigger and faster. But if you look at the design of it going back to the first training TPU, that block diagram still works all these, you know, a decade later. So the people who designed that in 2015 or two did a really great job. But the main components of a big matrix multiply unit, the so-called high bandwidth memory, a vector unit to go with the matrix unit, the basic building blocks are still there.
38:06The Peterman Pod Host:I think in your career with the CPUs, benchmarks played a huge role in measuring which computer architectures were performant and not. In this new space of floating point operations for AI, is there a benchmark that people use for GPUs? And can you also apply it for TPUs? I and some friends were helped involved in what called the MLPerf effort. So it was inspired by the spec CPU effort with the spec benchmarks that there were these competing companies and they'd all make claims about theirs was better than the others. And they realized that wasn't good for the industry. So they agreed on a set of benchmarks.
38:48So the ML perf effort that's run now, but what's called the ML comms, is an attempt to do that. If we're going to do this comparisons, let's not argue about what the benchmarks are. So that's a serious effort in that. Kind of interestingly, what's happened, I'd say, is because there's two pieces to the design. There's the machine learning libraries that are to implement a lot of features that you need for these applications. And the libraries will even be rewritten for specific applications rather than just having general libraries. This has turned out to be a big advantage for NVIDIA because they have a large corporation with lots of people that are available, lots of engineers who could build these libraries for them.
39:39So it's turned out not quite a, it's not a neutral evaluation, right? It's the architecture plus the libraries that go with it and the compiler too, but specifically the libraries that you tailor each time. So NVIDIA brings out, when they announce a new architecture, they create a new set of, they modify the libraries to run that really well or to run applications really well. So that's a powerful combination. Kind of in business terms, people refer to NVIDIA as having this CUDA moat. And part of it is CUDA, the programming language, but a big part of it is the libraries that NVIDIA makes. So this has made it difficult for startups to be able to compete with NVIDIA, partly because, you know, they didn't put enough emphasis in the software and partly because, you know, NVIDIA just has many more engineers than they do to be able to tailor the libraries.
40:42So it's the libraries plus the architecture that's this powerful advantage why most people do things on GPUs. Now, Google has been able to develop their own libraries. They don't have as many engineers. They use compilers more than I think NVIDIA does. So they have a set of libraries that they can do. But up until recently, Google has, it's all been internal. It's just for Google to use, or you could use it being the cloud. But these startups have to have a difficulty. As a result, that's why the MLPerf hasn't been as popular. Not everybody runs them because NVIDIA runs them really well, really better than everybody else.
41:29And the startups have a hard time showing off what they can do, given they don't have the engineering effort going into libraries that you need to do well in MLPerf.
41:39The Peterman Pod Host:You gave this popular talk about how to have a bad career, and it's kind of the negation of advice. to kind of have a good career. And I was just wondering if you could kind of summarize maybe the top three things that you kind of think people should take away. Yeah. So, yeah, well, there's been a few things. So the story is early in my career, my friend, a good friend and I said, how can we teach grad students how to give a good talk? And we thought it'd be funny to explain how to give a bad talk. And then if you didn't want to give a bad talk, this is, so this is how to do it badly. And if you don't, here's the things you do not to do it badly.
42:17And so then I later did a how to have a bad career. That's pretty much focused towards academia, you know, so for researchers to be able to do that. I later did a how to have a bad, how to build a bad research center, how to build a bad research lab. And then recently I've been giving how to give AI a bad carbon footprint. That's my latest one. But I think the thing that might be more relevant is at the end of my talks, I would kind of reflect on my career and talk about lessons learned. So I've written a paper called, I think it's Life Lessons from the First Half Century of My Career. I wrote that recently.
42:56And so those are divided up into kind of career advice and personal advice there. On the personal side, I'd say if you have a family, you know, make sure your family's first, keep your family first. The technology we've invented makes it really easy for your, you know, to take your work home with you and not pay attention to your family. Uh, I had a, when I was first here at Berkeley, uh, I gave a senior faculty member a ride home, dropped him off late at night coming up to Silicon Valley. And he said, well, Dave, if I had to do it all over, all over again, I wish I'd spent more time with the family.
43:32And I never wanted to say that. And nobody in their deathbed says, you know, I wish I'd spent more time in the office. So, so you gotta, you know, whenever it's easy to kind of, you know, let the, let your career overcome the needs of your family, but you just, you gotta just remember that. I would say in my life, kind of growing up in the fifties, you know, the idea was to be happy, you had to be wealthy, but those are actually two different goals, wealthy and happiness. So I always made decisions towards happiness versus wealth. And I felt very good about that. And even now, why would I pick something that made me wealthy and unhappy?
44:15Why would you do that? And in the field that we're in, it turned out there was a lot of wealth to go with it. So optimizing happiness didn't mean you had to suffer beyond that. I think it's important to have fun personally. I think when you're a kid, you don't have to tell kids to play. But as a adult, you get so busy, you don't think you have time to have fun. But, you know, you only get to do this once. So I play soccer. I run my bicycle to the interview. I lift weights. I body surf. I, you know, do things with my wife and family and my sons. So it's important to have fun. I think on the career side is one of the pieces of advice is by Akanthe.
45:05He wrote this book, The Habits of Very Effective People, I think. And he has a little quadrant and he divides it of urgent and not urgent and important and unimportant. And kind of there's so many things in our technology like email and texting for you to focus on the urgent things. but you really shouldn't be spending a lot of time on the unimportant and urgent things. And it takes self-discipline to set aside time for the important non-urgent things. But if you don't block that out, you can just not have time to do anything. I get to see other people's calendars at Google, and there's managers that every half hour from eight to six, five days a week are scheduled.
45:48And I don't see how you have time to think and reflect on things like that. Other career advice things is I remember I kind of woke up one morning and it was like God spoke to me. I was like thunderstruck. And it says it's not how many things you start, it's how many things you finish. It seems relatively obvious, but, you know, that's not the way I was acting. I had they had many things going on. But after that, I was like, there's one main thing I'm doing at a time. So when Hennessey and I wrote a textbook, that was the main thing. when I was department head here, that was the main thing I did.
46:26I would do some other little things there. And John Hennessey, he wrote a book to a kind of a career advice book based on his presidency. And one of the things he said in there, you know, you're only going to be remembered for the five or six things you've done in your life, not for the hundreds of little things. So to give yourself a chance to have some things you're really proud of, it's better to concentrate on a few of them, hoping that some of them will turn out to be a big deal rather than scatter yourself to many things. But if you're interested, you can, yeah, if you look for life lessons, David Patterson, first half century, you can see the whole list of, I think there's 16 lessons altogether.
47:11The Peterman Pod Host:You mentioned that you didn't want to reflect back on your life and feel like you didn't have enough family time. And yeah, I've, I don't think I've ever heard anyone say the opposite. Why do you think that is? Why is it that everyone looks back on their life and they never regret, you know, a lot of family time? I imagine there's got to be at least one person that says, I spent too much time with my family. My career suffered. Well, you know, what's, you know what's it's it's pretty philosophical i mean what's what's what's life all about what's success right you have to you have to figure out what that means for you i mean if you if you have i don't know if you have financial goals of being a you know a millionaire or billionaire or something like that and that's how you judge your life okay good luck you know it's pretty hard to make it, it's pretty hard to do.
48:09But I think it's the, you know, when I was finishing my PhD, I read this book by Studs Terkel called Working, where he interviewed all these people in his careers and they'd look back and what they liked and what they didn't like, what they felt about their careers. And what I got out of it, the people who worked with people like ministers or teachers or doctors felt really good about what they did with their careers. And the people who did more ephemeral stuff, you know, like, you know, technology things or airplanes or something like that are long gone, didn't feel as good about it. It was the people that they worked with that they really cared about.
48:43And I went out with a retired engineering dean here and to a meeting who I knew. And he said, you know, Dave, as I flink my hair, it wasn't the projects, it was the people that I worked at that mattered. And I thought, I knew that from a long time ago. So that's, I mean, this is kind of senior persons offering advice. I think you're going to care more about the people you've worked with and people you've helped will be a bigger deal. That's part of, you know, one of the other things about personal happiness is they studied happiness. Psychologists used to just study crazy people, but they started like, why are people happy?
49:17And they know what the reasons are, you know. Have a job that you like, you know, have friends and family. Helping other people. Helping other people makes you happy. They know this. Having something kind of either a religious side of it, it doesn't have to be a formal religion, but like contact with nature, the grandeur of nature. But the kind of the list of things you need to do to be happy is well understood. And, you know, there are these hats, and our world is filled with unhappy billionaires, right? If money was the thing that made you happy, why are these people, very wealthy people, so mad about things?
49:56So yeah, this is me passing on advice.
49:59The Peterman Pod Host:In one of your talks, you had, it said, what worked well for me? And it was kind of some reflections. And one of the things in there I thought was unique and interesting, you mentioned that courage was a big part of your career. And I don't hear that too often. I was curious why you say courage is so important in a career. Yeah, I think that's, I mean, that might be partially my personal makeup. But, you know, I was kind of the youngest kid in my class and kind of small. It took me a while to grow despite age, so I was always small. But my parents encouraged me to go out for wrestling. And wrestling gives you physical self-confidence because you, you know, spend years doing that.
50:46I did it in high school and college. And so I think partly, you know, technically, you know, having courage to do things, it kind of goes along with the advice is fortune favors the bold. That's this that goes that's advice is 2000 years old. I mean, it's hard to figure this out. Helen Keller wrote, you know, even even trying to play it safe, you still get caught. And so it turns out you might as well you might fail no matter what. And if you take a big chance, you can succeed. If you don't take the chance, you probably won't succeed if you play it safe. So fortune favors the gold, and it takes courage to do that.
51:28I think also just for me intellectually, I feel like if there's something not right, I need to stand up and confront it. And I think that kind of ironically comes from the wrestling side of my personality, where if I see somebody getting picked on or something like that, I'm going to stand and try and stop it. And I feel that same responsibility intellectually if people are making bad arguments or doing something that we need to stand up and do it. And I feel good about that. The cautionary start about that is because I guess one of my senior faculty members saw this nature of me. He said one of the sayings is friends come and go, but enemies accumulate.
52:11This is an old saying. So if you think about it, you kind of, people you went to high school with a while, friends, you kind of forget them. But somebody who you really does dislike, you never forget that you dislike that person. So standing up when it's important, but be careful when you make enemies because, you know, they're going to stick around for a long time.
52:29The Peterman Pod Host:When you say something's gone bad technically, do you mean someone was incorrect? Yeah, when they're weak, either politically or technically, when it's a weak argument. I think I really like there to be a marketplace of ideas and we hone the ideas by arguing them. And so if it's a specious argument, even if a person is a leader of a company and it just doesn't make sense, I feel it's kind of the responsibility. It's better for the company if somebody stands up and points that out than to just let them get away with it. And then, you know, it's a little bit confrontational. But, you know, as long as people all agree that, you know, this is for the greater good, we need to get the right ideas out there.
53:22And so let's argue about the ideas to see, you know, polish them to make them stronger. I think that's important in science and engineering and kind of in life too. There's a lot of stuff going on right now in the country that is worrisome. And I've certainly stood up and wrote op-eds about things that I think are wrong and need to be corrected. And if people are afraid to do that, it's hard to be optimistic about the future. of people are afraid to stand up when there's wrongs and try and stop them.
54:02The Peterman Pod Host:You also mentioned optimism in the talk, and you had this story I wonder if you want to share. So I would say in engineering, it's hard to know, right? But I think you need to be kind of optimistic or have a positive outlook because so many things could go wrong. And then some of my story, personal story that illustrates it's going back to high school when I'm 16, I'm dating this very attractive girl. And I screw up my courage and ask her if we would be exclusive. At the time, the phrase we used was going steady. And she looked at me and said, and she was 16. She had dated other guys and thought we were pretty young.
54:42And she said, well, Dave, you're such a nice guy. I don't know how to say no. For me, as a logical person, not a no sounded like a yes. And so I hugged her and said, great. And so she, in her mind, she thought, well, I'll let him down gently later. But we've been married 59 years now, and she hasn't let me down yet. So that was the case where optimism paid off.
55:06The Peterman Pod Host:I think everyone that hears a healthy relationship for that long, they might wonder how you did it. I used to tell people, you know, if you go to weddings, the marriage vows are really great, right? But nobody can remember their wedding vows. I used to say remember your wedding vows, but nobody remember that. So we boiled it down to nine magic words. And it's just three sentences. And they start I, you, I. And you got to say all three. And it's I was wrong. You were right. I love you. Okay. Those are the, those nine words. And this applies to both, both partners in a relationship, not just one partner.
55:43But yeah, if you can say them all and no substitutions, I was wrong. You're right. You're a jerk. You know, you can't do that. If you can remember those nine words, that can help you have a long relationship like my wife and I have.
55:56The Peterman Pod Host:And then last question for you, like knowing everything you know now from your career, if you could go back to yourself when you had just entered the industry and give yourself advice, what would you say? I mean, I think when I got here, because, you know, the imposter syndrome, I was a UCLA graduate student and suddenly I'm a Berkeley professor. So that just doesn't seem like that was very intimidating. But after a while, I just thought, well, I'm probably not going to get tenure, so I should just have a good time. So I think I already had a right attitude about it. I think that first year, I think it was very stressful while I was trying to handle the imposter syndrome and be a Berkeley professor.
56:40But I think after that, I handled it pretty well. I did all the things with the kids and stuff. So there's a version of that question is like, is there anything I would do over again? There's one thing I would have. I was chair of the architecture community, the SIGarch, as it's called. And they have an annual conference this year. And I was, this was in the 1990s, I think, I was the chair. And what I wasn't aware is at these conferences, there were men who were harassing young women at this conference. I just didn't think, you know, people like, you know, young people, people like me, nobody would do that.
57:22Only an idiot would do that. That can't possibly be happening. But it was happening. And I wish somebody had said something to me about it. And because I would have straightened out any man doing that. I would have threatened his life if he were to do that today. The only comforting thing is Sarita Abwe, who is a famous computer architect. And she said she also was not aware that that was going on. It became clear later, you know, five or 10 years later, it became more clear that this was going on and there were mechanisms. But that's the one thing I wish, you know, if I could go back in time, I would have figured that out and I would have straightened men out who were doing that.
58:05And they wouldn't, that would have stopped. I believe that would have stopped them.
58:08The Peterman Pod Host:Yeah. Thank you so much for your time today. I really appreciate it. All right. And thanks for the interview.
58:43The Peterman Pod Host:split keyboard so there's two sides this is in the case but yeah we launched on kickstarter and we hit our goal within eight hours of launching i really appreciate it if you were one of the people who grabbed one of the early units we're now working on the long journey of building the tooling now and so if you still want to pick one up i've left the late pledges open on kickstarter so you can grab one there i'll put a link in the description thank you again for watching the podcast and i'll see you in the next episode
From the publisher
David Patterson is a Turing Award winner famous for his contributions to computer architecture. I interviewed him about his past work, thoughts on GPU/TPUs and career advice from half a decade of experience.
• My ergonomic keyboard project I mentioned, you can follow along here: https://read.compose.llc/
• The Kickstarter page for it: https://www.kickstarter.com/projects/ryanlpeterman/compose-simple-ergonomics-beautifully-done
Podcast links:
• YouTube: https://youtu.be/Pn4ZwlEh5nw
• Apple: https://podcasts.apple.com/us/podcast/the-peterman-pod/id1777363835
• Transcript: https://www.developing.dev/p/turing-award-winner-tpu-vs-gpu-vs?r=n49ky
Timestamps:
(00:00) Intro
(00:42) RISC vs CISC
(12:51) Compilers
(17:38) GPUs
(23:07) GPU vs TPU vs CPU
(32:12) Is Moores law dead?
(38:04) GPU benchmarks
(41:40) How to have a bad career
(49:59) Courage and optimism
(55:56) Advice for his younger self
(58:15) Outro
Where to find David:
• Wikipedia: https://en.wikipedia.org/wiki/David_Patterson_(computer_scientist)
Where to find Ryan:
• Newsletter: https://www.developing.dev/
• X/Twitter: https://x.com/ryanlpeterman
• LinkedIn: https://www.linkedin.com/in/ryanlpeterman/
• Threads: https://www.threads.com/@ryanlpeterman
• Instagram: https://www.instagram.com/ryanlpeterman
• TikTok: https://www.tiktok.com/@ryanlpeterman
Referenced in this episode:
• AlexNet paper: https://proceedings.neurips.cc/paper_files/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf
• David Patterson's “How to Have a Bad Career” talk: https://www.youtube.com/watch?v=Rn1w4MRHIhc
• “Life Lessons from the First Half-Century of My Career”: https://cacm.acm.org/opinion/life-lessons-from-the-first-half-century-of-my-career/
• The 7 Habits of Highly Effective People (book): https://en.wikipedia.org/wiki/The_7_Habits_of_Highly_Effective_People
• Working (book): https://en.wikipedia.org/wiki/Working_(Terkel_book)




