Andrew Feldman: Building the World’s Largest and Fastest Computer Chip for AI

2 May 2024 · 44 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Notes: Generative Now - Episode with Andrew Feldman

Episode Overview Title: Andrew Feldman: Building the World’s Largest and Fastest Computer Chip for AI Host: Michael Mignano, Partner at Lightspeed Guest: Andrew Feldman, CEO and Co-founder of Cerebras Systems Description: A discussion on Cerebras' latest innovation, the Wafer Scale Engine 3, and insights on AI supercomputing.

---

Key Topics Discussed

Introduction to Cerebras Systems

  • Cerebras Systems focuses on building AI supercomputers using Wafer Scale Technology.
  • Andrew Feldman, the CEO, has a background in hardware systems with multiple successful startups including SeaMicro and Force10 Networks.

Wafer Scale Engine 3 (WSE3)

  • Launch Announcement: The WSE3 is the largest and fastest chip ever made, approximately the size of a dinner plate.
  • Specifications:
  • Contains 4 trillion transistors and 900,000 cores.
  • Aims to revolutionize AI compute by offering unprecedented performance.

Vision and Purpose

  • The founding vision was to address the increasing demand for AI compute, driven by large language models and other computationally intensive tasks.
  • Focused on improving communication in computing rather than just computation itself.

Unique Challenges in Chip Design

  • Discussed the obstacles in designing such a large chip including:
  • Cooling: Developed innovative water cooling systems to manage heat.
  • Power Delivery: Addressed significant energy demands unique to larger chips.

Market Landscape and Competition

  • Noted that while companies like NVIDIA and AMD are key players, many are playing catch-up in the AI chip market.
  • Highlighted how the strategic missteps of larger companies (like Intel's focus on x86 processors) allowed Cerebras to carve out a unique niche.

The Future of AI Compute

  • Emphasis on the growing importance of compute capabilities in training and inference of AI models.
  • Predictions about the continued hunger for compute and ongoing investment into AI technologies.

Applications and Use Cases

  • Cerebras is working with diverse industries, including:
  • Healthcare: Personalized medicine and drug discovery.
  • National Laboratories: Significant projects like predicting virus mutations.
  • Energy Sector: Simulating reservoir behavior for major oil companies.

Reflections on AI’s Evolution

  • Discussion on how AI will integrate into various sectors over time until it becomes a standard part of life, similar to existing technologies like CRM systems.
  • Importance of continuing to innovate and solve challenges in AI to improve societal outcomes.

Company Philosophy and Future Directions

  • Cerebras aims to be 10x better than competitors, focusing on both hardware advancements and software efficiency.
  • Continuous innovation is emphasized, with plans for future iterations of their technology already in motion.

---

Key Takeaways

  • The WSE3 represents a significant leap in AI hardware capabilities, driven by innovative thinking around chip size and design.
  • Cerebras' approach combines a strong understanding of the technical requirements of AI workloads with a vision for market needs.
  • As AI technology matures, the industry will shift from focusing on the technology itself to discussing its impact on society.
  • The collaborative use of Cerebras technology across various sectors is creating pathways for groundbreaking advancements in AI applications.

---

Conclusion The conversation with Andrew Feldman illustrates the exciting developments in AI chip technology and the broader implications of these advancements on industries and everyday life. The journey of Cerebras serves as a case study in innovation and foresight in the rapidly evolving landscape of AI and computing.

---

For more insights and updates, follow [Lightspeed Venture Partners](http://www.lsvp.com/) on their social media platforms and subscribe to the Generative Now podcast wherever you listen.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:04Hey everyone and welcome to Generative Now. I am Michael Magnano. I'm a partner at Lightspeed And this week I'm sitting down with Andrew Feldman. Andrew is the CEO of Cerebrus, the hardware company that builds AI supercomputers. He's a leader in the AI industry, creating the next generation of compute hardware. And this was a fascinating conversation. Just a few weeks ago, Cerebrus announced their latest and greatest innovation, the Wafer Scale Engine 3. This is the largest and fastest chip in the history of chips. It's about the size of a dinner plate and promises AI compute that is faster and easier to use than anything else on the market.

0:43And so without further ado, please enjoy my conversation with Andrew Feldman. Hey, Andrew. How you going, Michael? Good, good. Thanks so much for doing this. Happy to do it. Happy to do it. So you are the co-founder and the CEO of Subrebrus. And this is a very, very exciting company and very relevant, I think, to what's going on in the world right now with all the excitement around AI. You recently had a huge launch that, at least from my perspective, seems to change things a bit in the world of chips and AI. And I want to get into that story, but maybe just to level set for the audience. We'd love to hear a little bit about your story and sort of the background of Cerebris.

1:26Sure. Well, this is my fifth startup. An inveterate builder of hardware systems, of chips and printed circuit boards and cooling systems and embedded software and software operating systems. And so this is what I love to do. And 2007, we got together and we started a company and we ended up selling it in 2012 to AMD. and we worked there for a little while and some of its leaders got back together in 2015 and we wrote two things on the whiteboard we wanted to work together and we wanted to move an industry and we began thinking about what what we might want to build and as computer architects we we tend to ask the same question again and again it's it is there an interesting workload and if we built a machine, could we make it faster?

2:23There's one question in computer architecture, is how to make the particular software problem faster. And we saw on the horizon AI. And we didn't know it would be as big as it is today, but we knew it would be big and important. And that we thought, even in late 2015, that we could build a chip, could build a processor that would be vastly faster at doing AI work. And we had an idea that the hard part of AI isn't the actual calculation, it's the communication. And that it would take, as these problems got bigger, it would take more and more compute. That compute wouldn't sit on a standard chip.

3:12And more and more pressure would be on the interchip communication. And so we had an idea that if we built really big chip, if we went from the traditional limit that everybody had in their mind of how big a chip could be to 50 times that to the whole wafer, we might be able to radically improve performance, reduce power drill, and make it easier to build big systems. And in 2019, we yielded the largest chip ever made and uh not by little bit it was a a chip with more than a trillion transistors and it's 57 times larger than the largest gpu that was at 16 nanometer and two and a half years later we announced a seven nanometer version and a few weeks back we announced a five nanometer version uh it is the size of a dinner plate whereas most chips are the size of a postage stamp It has 4 trillion transistors and 900 ,000 cores.

4:17Those are little engines that drive your AI work. And so we built a whole system around these. And the system delivers power and cooling and the data. And then we started tying these systems together. As the problem got bigger, we built first clusters and then supercomputers worth of these machines. And we now build the largest AI supercomputers in the world. Then we started, we had so much compute capacity, we began training models like crazy. And have trained, you know, dozens of production models and thousands of test models along the way. So, I mean, this is fascinating because it seems like there is a big problem in the world today.

4:58There's so much demand for these chips. And you, as you mentioned, like in the founding of this company, you even said that you guys had the foresight to know that this was going to be a really, really big deal. But it seems like some of the biggest companies in the world that have been at sort of the forefront of chip making and, you know, some of the highest market cap companies in the world almost seems like they're playing catch up. So, like, how did you guys have the foresight to know this would happen? And maybe from the outside looking in, it almost feels like they didn't. In retrospect, we got it right.

5:28Trust me, made many bets of my career that turned out way less well than this one. And when you build a company, you make many bets that are wrong. As a CEO, you're making decisions every week. And if you get most of them right, you're doing really well. There were a collection of idiosyncratic elements in AI. First, in about 2015, as this was exciting, Intel decided that they could attack this problem with an x86 processor. And that was a big mistake. AMD was, under Lisa's leadership, under Lisa Su's leadership, building its processor business. They had come out of a period where they had made some really bad decisions.

6:14And Lisa had taken over in about 2013 or 14. And they were really focused on building a better processor. Not a better part for AI, but they're a core business of theirs. And they got that right. But what that meant is every ounce of their energy was really being focused on a different workload. And that left among the big three that left NVIDIA alone. You could see by AMD stock price over that time, they got the CPU part right. They built tremendous parts and were a customer of theirs. And now they've turned their attention to building GPUs for this market. And so I think that was the large company environment.

7:01I think there were a group of startups who saw some of the same things that we did. Intel bought one. They've bought actually several now. But we've done a little better than most of the others. Been able to differentiate ourselves, produce significant advantage, vastly better performance, easier to program. Things that we bet on turned out to be really important in the market, especially with the rise of large language models. Why does this matter right now? Maybe level setting for the audience, like what's happening in the market right now? Obviously, you know, if somebody is listening to this podcast, they're probably aware that there's a lot of demand for GPUs.

7:44Just don't swipe right. Yeah, exactly. Exactly. Maybe set the stage for why and how the wafer scale engine three could have such a dramatic effect on this market right now. So I think it's important to note two things. The first is that the type of AI we're doing right now, particularly language models and generative AI, is enormously computationally intensive. So that's the first part, that this is a brute force methodology. You and the amount of compute you need is approximately the product of the number of parameters in your model and the amount of data tokens you show for training. that is the defining characteristic and so as we've built bigger models and as we learn that showing the more data will improve the accuracy of the model we've ratcheted up the amount of compute we need and you know we now do ai with uh an amount of compute that previously was only in the nuclear physics domain, right?

8:51Previously, these supercomputers were at a small number of oil companies and at a small number of government labs. And now that's the basis for training models. On the inference side, the amount of compute you need is a function of several things, but primarily the number of parameters in the model. It is a linear function of the number of users. So as we began to train models, these models got bigger. They became more useful. We began to do inference on them, and more people used them. And so the amount of compute on the inference side went through the roof. And so the importance of compute is if you can train faster and use less energy, less electricity in your training, better for the society, you can test more hypotheses, you test more ideas in less time.

9:44and that's how these models get better. They aren't sort of born like Zeus's children, sort of fully formed from the heads of smart people. They are instead born of a set of iterations, tests of small models, hyperparameter tunings, and these go on. And in order to announce a big new model, you know, like Lama 3 or one of these, they've done hundreds, maybe thousands of small tests. and this takes vast amounts of compute and takes a long time. And if you can shrink the amount of time it takes, you get to get to good answers faster. Yeah, and so now the Wafer 3 exists and how do we go from where we are today to a point where developers can start leveraging this instead of having to use a broker to find whatever clusters they can of H100s?

10:39A GPU cluster in Botswana? Yeah. Yeah, exactly. We've stood up over the last nine months, eight exaflops of AI compute. We will stand up eight more in the next four or five months. We've stood it up for customers. We've stood it up for our own cloud, for our customers who are CSPs, who offer cloud services. It's available now. We have customers in North America, in non-China, Asia, in Europe, in the Middle East. The idea of going big was to solve two problems. The first is that when you break up a problem, when a problem doesn't fit on a GPU, you have to break it up into component elements. and you then spread that over many little GPUs, right?

11:38And you have to tie them together and you have to move information back and forth. And this is a process called distributed compute. And not only does it put tremendous pressure on the communication fabric because it is now central to the compute itself, but also it puts pressure on you to know how far away the slowest of, the furthest of the GPUs is because the slowest response determines when the answer is, right? You break up a problem, you send it out, the results are aggregated and you wait until the last result comes. Then you can find the answer and move on. And that's Amdahl's law. And that meant that you couldn't use traditional cloud technology.

12:31It meant that you needed to know, right, exactly how far away and have to be deterministic. And this is why distributed compute with GPUs is so challenging. I think Andre Iquipathy has this statement about just how hard it is in practice to do. Now, we went big so that you never had to do that. You never had to break up work. You could keep work on a giant chip. and when you move work around the chip, you move it at about the lowest power. We know how to move data and we move it about as fast as we know how to move data. And so by keeping it on the chip, you avoid the hard parts of distributed compute and you get to work at low cost and low power.

13:19Now, while we were going big to solve the communication problem, NVIDIA was aware of the communication problem. They bought Mellanox, right? So important was the control of the movement of information that they paid, what, six or seven billion dollars for a switching company. And so this is a fundamental problem. We solved it with innovation, making a giant chip, and they solve it with Ethernet or proprietary Infiniband. Yeah, I wonder if that's a business decision, right? It's like it enables them to bundle the solution to the problem with other products and services, right? That they can offer.

14:00If you're a switch maker right now, the worst thing on earth for you is if NVIDIA is selling your customers hosts because they want to rip your networking gear out and swap in theirs. Right. Right. And so this larger chip, you know, if I'm understanding it correctly, it basically reduces the need to cluster as much as you might normally have to, which probably means that you could build far more efficient configurations. That's exactly right. Okay. That's exactly right. The idea here is that the hard part of distributed compute is as you go from one or four or eight or 16 little nodes to 100 or 1 ,000, you have to keep track of where they all are.

14:43You have to divide your work up. And this isn't AI. You've made your model. This is now distributed compute. This is putting your model onto the compute substrate, onto the computational cluster. That often takes weeks or months. And with us, it's a keystroke. And so that whole amount of time that's called tensor model parallel or pipeline model parallel distribution goes to zero. Once you stand up, once you've done this, we're faster in the straight up training or lower latency in the straight up inference. So the WSE3, it seems to almost like blow Moore's law out of the water. It did. And I wonder, is that because of what you're talking about here and the increased, like the increased size, is that what enabled that to happen?

15:34Yeah, it is. And is this Feldman's law? This is not Feldman's law. Feldman doesn't need a law. We will leave lawmaking to others. To Gordon Moore. That's right. To far more important people. I think what happened was what we saw was you would need more silicon, right? And you can get more silicon one of two ways. You can make bigger chips, or you can take big whole wafers, dice them up, put them on expensive motherboards, put them in different systems, buy expensive switches, and tie them back together again. And all chips are made off a wafer. And so in the fabrication process, everything begins on a wafer.

16:23And then traditionally, these were chopped up. We call it dicing into 800 square millimeter chunks at the biggest. And these were then put on a motherboard and you might use a thousand. We're like, why are you cutting them up? What if it's, you know, why are you cracking hunt, dumpy, spreading it out all so that you can get them to be back together again? We're like, why don't we leave it as it is? Can we invent technology that allows us? And this was the first time in the 70-year history of the computing industry that you could do this to build a chip this big. And then you get all the benefits of communication instead of across switches and Ethernet and with Nix.

17:08You get it at the speed of a silicon wire. Yeah. Yeah. Assuming Moore's law does continue, even though I guess these clusters can now be built more efficiently because of the size of the wafer, we will still need as much compute as humanly possible. And so I wonder if this ends up, physically speaking, covering even more ground. Does that make sense? Because we're never going to run out of our hunger for compute, it seems like, right? I don't foresee, certainly in the next three to five years, us running out of our hunger for AI compute. This is trillions of dollars is going to go into AI compute.

17:53And I think one of the things that's different, that really, I was in the mid-90s, I was building the first generation of hardware-based switches. And that was extraordinary as we moved, sort of, we made IP networking. free. And there were about 20 companies and we were driving and we had no idea that later WhatsApp and all these other apps would sit up on top of this and change the world. We were just thought we could make switching really, really cheap. Here, nation states are playing, right? You have sovereign clouds being built. You have them built. One of our partners in the Middle East and the UAE is called G42, and they're building together with us in strategic partnership, massive clouds.

18:40You have in Singapore and Australia and the UK, you have nation states standing up giant super compute clusters for AI. And you have that. It used to be that if you wanted your compute, you went to the hyperscalers. And now you have dozens of choices who have huge amounts of compute for rent and for sale. um and so the whole landscape is changing and i i think the ability to do ai and do it at massive scale has become sort of a strategic imperative will our hunger subside um maybe when uh the scaling laws for llms ultimately asymptote or i mean i guess we don't even know if that's going to happen?

19:25I think that's the wrong question, Michael. I mean, we don't know. What we've seen is that, and we've seen it sort of behind the scenes with open AI as their models got bigger, but we're seeing it now with Lama 1 move to Lama 2, move to Lama 3. I think the models are getting better as they get bigger. Now they get more expensive to do inference on. And this is the really curious part, right? There's a cost to that twice. First, when you train it, and second, when you do inference on it. And so I think that the models that will go to production will be trained differently. They will be smaller, but trained with more data.

20:00And they will be domain-specific. And whether you do that as a mixture of experts or a standalone domain-specific models, it is difficult to, and extraordinarily expensive, to do inference on 100-plus-billion-parameter models. And I think if you can get most of the gain by going 10 times smaller, but with better data, I think most production users would take that in a heartbeat. Yeah. Yeah. And I guess the other thing that could slow it down is you mentioned data, if and when we run out of data, right? Or we can't, you know, we can't solve data through synthetic data or something like that. The use of synthetic data is hugely important.

20:44It sounds strange or scary, but if you think about the way we train airplane pilots, right? You train them in simulators, and that's a synthetic environment. And flying a plane most of the time is there's not a lot to learn, right? And so when you want to train a pilot, you want to train landings and takeoffs, you want to train when problems occur. and the simulator can do that, can create that environment extremely well. And so it's a really important part of a pilot training. And that's, you know, the environment there is generating data and the pilot is the thing being trained. And so to the extent we can do that in self-driving cars, it's the same.

21:29Driving on the freeway is not a very difficult challenge for a car. an unprotected left turn in the snow is a much harder problem for a car. And it's like that in 25 different domains where the data collection might be extraordinarily expensive. Human subject trials, for example, we do a lot of work with the Mayo Clinic. They have some of the great medical data in the world. unbelievably expensive to gather. Unbelievably expensive to gather. 30 years of patient records, right? To the extent that one can augment with synthetic data, we will be able to do even better things. Yeah, and I guess crossing over that synthetic data event horizon would probably be really good for your business because it means that we can keep training, right?

22:28The cost of data plummets. And right now, some people scrape the internet for their data. Maybe one way to think about the landscape of AI is there are three fundamental pillars. There's an algorithm, a training model. There's data and there's compute. And when the model hits the data on compute, it learns. And that's what's happening. And so you can find in that just about everybody's strategy. Some began writing algorithms. They scraped the internet. So they got their data for free and they partnered with hyperscalers for their compute. We came at it from the compute side. We built extraordinary compute.

23:14We had so much compute that we gained tremendous experience on the algorithm side. And many of our customers have unique data sets, right? And others, many of the pharma companies, many of the life sciences have extraordinary data. And they're looking to find insight in their data. And they're partnering companies like us to build models with them and for compute. And so this sort of dynamic. And so to the extent that good data for a particular problem becomes less expensive to gather and can still teach a model, it's enormously valuable to the system. Yeah. Going back to the WSE3, you know, maybe talk a little bit about some of the other challenges that come with building these chips.

24:01You know, for example, cooling. You know, how does this thing, you know, if it's so much bigger, how does that affect heat and the problem of having to cool these things as they overheat? I think your listeners might be interested. You know, we did fundamental innovation and that is expensive and time-consuming and you need a very patient board. You need to raise a great deal of money. And when you do that, even when you succeed, often you find yourselves in a position where nobody's ever got what you have. So there's no received wisdom. And there are no components. You can't go to a catalog and order a heat sink for a part 57 times larger than had ever been built before.

24:44Nobody's got that in their catalog. Right. Right? And so not only do you have to invent in your primary domain, you frequently have to invent in components, in test fixtures, in the software testing structures. All of those have to be invented as well. Now, for 70 years in the computer industry, nobody had been able to build a chip this big. And there were a litany of reasons why you couldn't do it. And in a grand irony, we solved all of those in about 10 months with$20 million. And none of them turned out to be the hardest problem. And it's very interesting. It's an analogy that, you know, if you were having coffee at the base of Everest with the first guys who failed to get up there, they came back down and they said, there's this really hard problem about halfway up.

25:35And I said, all right, you can make a note yourself. You get up, you get to the top, you come down, and you're having coffee with the same guys. You say, that wasn't the hard part, right? The hard part was near the top coming down. the hard part was different than because you knew because you knew that would be a problem or no because nobody knew nobody had even faced the problem everybody had been stopped two-thirds of the way there i see not to face the hard problems right right you never yeah no one has ever tried to cool uh 17 kilowatts to a chip right right i see never tried to deliver that across the motherboard nobody ever tried to cool a chip or package a chip of that size because nobody ever had a chip of that size.

26:17And so it turned out that these problems that hadn't been encountered before turned out to be easier than the problems that had broken every previous effort to solve this problem. And that was sort of interesting and non-intuitive. And so the problem of powering and cooling, we have dozens of patents around that. Our inventions are significant. Two interesting things happen. First, per square millimeter, we use about the same power as other people at five nanometer. We have about the same number of transistors per square millimeter as others. We just have more square millimeters. It's just bigger.

26:59It's just bigger. They have 800. We have 46 ,000. Yeah. Now, this does two things. First, when you cool a chip, generally you put a little fan on it called an impingement fan and you can only afford a very cheap cooling system because the chip's standing alone because we had consolidated all this silicon we could afford to use water and so we designed a water cooling system for it and so we use a cold plate on the back and we run cold water across the back of the cold plate And so the result was we could amortize a more expensive cooling system because we had so much compute to amortize it over.

27:41And the result is we run the wafer at a lower junction temperature. And that means in electronics, you have a longer lifetime. Got it. Got it. But all of these were really hard problems. And each time you thought you were in the clear, right, each time you thought you were through the tunnel, you know, you encountered another difficult problem. And that was sort of the fun part of this project is we were well beyond where anybody else had ever been. You know, on the old maps, the old cartographers used to write here the dragons where they didn't know, right? That's where we were. We were out there with the dragons.

28:22Well, what's cool is now when you build the next one, you already know that some of these problems are actually solvable. I mean, I'm sure there'll be new problems, right, on the next mountain. We ended up being among the world's leaders in a collection of areas that we didn't know we would need to be good at when we started. Right, right. And we did it through sort of relentless engineering discipline by carefully making mistakes and understanding which failure analysis, why we failed, not making the same mistake again and step by step. And this was true for not just the wafer scale design, but the packaging and the power delivery and the cooling and the compiler software.

29:02We made a ton of mistakes there. Over time, it just got better and better and better and better. Maybe talk a little bit about like the unique properties of the wafer scaler relative to your peers and competitors probably enables you to work with some interesting different types of customers, right? I understand you work with national laboratories and healthcare giants and obviously in the enterprise. Like maybe talk a little bit about the customers and some of the cool use cases you've seen as a result. There's some really extraordinary use cases. And we designed the chip and the system and the software that goes on it and the compiler chain that allows your ML practitioners to write PyTorch and run on our machine.

29:49And then what we sell is either a system for your premise or time on systems through our cloud or through our CSP partners' clouds. And so when you are this fast, we've been able to solve problems using generative AI models to predict mutations in viruses like COVID. Some of our work in partnership with Argonne National Lab was given the top award in the industry for that. We have a publication coming out shortly with another one of the national labs. We're doing work with, we published work with Gladys Smith-Klein, where we did AI work on epigenomic data. Working with our strategic partner, G42, we trained the largest Arabic model.

30:39GPT-4 is really weak in Arabic, and there are 400 million Arabic speakers worldwide. And so we trained an Arabic model. And then it's now, it's called JACE, which is one of the tallest mountains in the UAE. And it's now one of the foundations for Microsoft's Middle Eastern service, served by them. That's really cool. Yeah, we're working on extraordinary models with the Mayo Clinic. Models that allow you to look at an individual's genetics and make estimations about which rheumatoid arthritis drug will be most effective. How cool is that? This true personalized medicine. We worked with Totala Energies, one of the largest oil companies in the world, in reservoir simulation and using AI for the estimation of reservoir behavior.

Read the full transcript

31:33I mean, it goes on, the amount of interesting things. We designed a model, a small LLM, it could be served off a cell phone. as the top performing small one. We are attacking interesting problems. The data types we've used, I mean, we've used satellite data, we've used MRI data, we've used genomic data, we've used text, we've used text in Arabic and in Portuguese and in Catalan. We use Japanese. We've taken in different data types and used different model types. I mean, it's really, it's been extraordinary. And what is it about the Waiver Scale Engine 3, besides the efficiency of it, that enables you to do these things in a way that other chips can't?

32:21Well, I would say there are a couple of things. And I think one is, it is my view as an entrepreneur that I don't want to be a little bit better when I sit out than the big dog. You want to be 10x better. That's right. Being a little better is a disaster because you've worked really hard and they can get you with price movement. They have no problem moving price or bundling, as you described earlier, the switches with the GPUs. And you have to go out and be vastly better. And for me, that generally means taking as much engineering risk as I can in pursuit of being better. And I know VCs like to say, what is your risk in the building?

33:02What are you guys, what can you control? And as an aside, I think those are the more interesting engineering problems as well. So we don't want to be a little bit better. We're willing to work and attack problems and even fail in pursuit of extraordinary results. And that was sort of what we set out to do here. The result was, you know, we're vastly faster. We really, you can configure us for the equivalent of a thousand GPUs with a couple keystrokes, not months of playing around and doing distributed work. And so you test more ideas in less time, your models get better faster, and so you end up in less time with a higher accuracy model.

33:52And that's what everybody wants. They want working models faster. and they want their models to be more accurate. And that's what we've been able to show around the world. What comes next? I mean, I know this only just launched. Are you already thinking about the next one? The minute you tape out a chip, so the chip sort of was taped out, I don't know, a year ago, you turn to the next one. The one was at five, you turn to a three nanometer project. The minute you finish one, you're on a treadmill. And that's sort of the life of someone who makes chips is you are always needing to go bigger, faster, cheaper.

34:32And the minute you finish one project, you jump on the next. And now that doesn't mean there's not a huge amount of opportunity to improve on the algorithmic level, right? What you've generally seen, both with GPUs and with us, is over time, the machine goes faster, mostly because you're getting to be more efficient with your software, right? The software, the utilization of the machine increases, and as a result, you perceive faster performance. And so the software has made huge leaps and bounds with us, and at the same time, you're trying to make new chips, you're trying to make your software more efficient and faster.

35:13You're trying to find algorithmic advantage. One such advantage is sparsity, which is something we're very, very good at. And that's when many of the models we use today are over-parameterized, right? And what that means is they do an all-to-all connection. And it's probably the case that not all of the independent variables, not all of the parameters actually impact learning. And so over time, there might be a zero in that weight matrix. And when you multiply by zero in a computer, it's about the dumbest thing you can do, right no new information is gained you know i don't have to do that work to know the answer zero and so if you can compress those zeros out you can reduce the power used and reduce the amount of compute used to get to the same answer and so there are all these really interesting techniques that look like compression and one of them is sparsity and And that's something we're very, very good at and further drives efficiency gain.

36:15So the improvements will be on multiple dimensions. They'll be on the hardware dimension. They'll be on the algorithm dimension. And they'll be on the model dimension. Our models will get better as well. There's a lot of talk now about how the model architecture could change. Obviously, everyone is talking about transformer models and diffusion models, but there's a lot of investment going into building the next type of architecture. How will that ultimately impact what you're doing on the hardware side? You have a choice. One of the fundamental questions in computer architecture, when you begin with your clean sheet of paper, is what to be good at and what not to be good at.

36:55And when we began with our architecture, convolutional networks were the rage. and one can choose to design special features for the convolutional network. Now, that's great if you do convolutional networks, but you pay a penalty for that. Just like you can have a four-wheel drive in your car, you actually pay a penalty when you're not using it. And I guess that's what we're kind of seeing with the H100s, right? They've made a bet that transformers will continue to be important and they put a positive circuitry for transformers. And the minute you try and do something that's not a transformer, they're performance dips.

37:35Right, right. Right. And so you have to be very, very sure at what the future will hold to use that, to do it that way. And we made a different set of decisions. We said underneath all of the AI today is a particular type of sparse linear algebra. And if we're good at that, whether you do transformers, which weren't invented when we laid this architecture out, or convolutions or any type of AI, we'd be really good at it. And so you make a different decision about where to specialize and where to be general. And we made a set. These were good decisions we made. We're happy to talk about plenty of bad ones I made as well.

38:17I mean, that's true. Every CEO knows it. And you make 100 decisions a week. And if you get 87 or 92 right, you're having a great week. Of course. And this was a good one we made. And the result was we were a very good machine as the market changed. And so as we move from straight up transformers to maybe a mixture of experts or other sort of architectures, we've remained really good at these. Whereas others who had done, you know, specialist things for a particular architecture, they've struggled to adapt. Speaking of making predictions and getting things right versus making mistakes, you made a great prediction, as we talked about at the beginning of the episode.

38:59And now everyone in the world, this is all anyone's talking about, right? Transformers, AI, generative AI. What are we going to be talking about in a couple of years from now? You know, what are your predictions for what's next? You know, I think a couple of things happen. in the adoption of technology is over time, it becomes embedded in life in a way that nobody talks about it. You go from a time where nobody's heard of it and nobody's using it to everybody's talking about it and nobody's using it to the point where everybody's using it and nobody's talking about it. That is just the life cycle of cool technology.

39:35When Pete, we sat on a plane and someone said, what do you do? And you said, you built hardware for AI. And they're like, is that like a computer? right you know then chat gpt came out and it's like do you build stuff for like chat gpt it's like yes well you know my nephew's using it to write papers for high school good for him that's great he should learn all he can but i think over time nobody talks about the engine behind search or google translate or maps or the the the technology behind the recommendation engine and Amazon or Netflix. They're all AI. And so what happens is the technology creeps into your daily life.

40:15Nobody talks about the technology behind your CRM. It enters life and begins being productive and stops being the center of a huge amount of conversation. It stops being scary. you know you don't need congressional uh uh panels convened to discuss crm technology it's no less cool it's just it's now embedded in the fabric of the ecosystem where right now we're new we're the new kid on the block and it's sexy and so what happens is we get better things that matter let's take health care is about 17 percent of gdp um what can ai do to to turn that into 14%, right? What can we do as an industry, not Cerebris, but all of us together, what can we do to deliver better healthcare for less money?

41:11And I think AI is at a crucial point there. So will we be talking about some of the applications, you know, in a few years, maybe we will be talking about some of the applications. The results of the applications. My doctor just built me a drug on demand that treats my rheumatoid arthritis. That's right. Rather than the doctor running through a series of experiments that take two or three years with you taking different drugs of rheumatoid arthritis and checking every two months to see if they work, maybe he gets it right with the first or second one. And in four months, rather than in three years, your rheumatoid arthritis is under control.

41:49All right. Maybe we can see things in an echocardiogram that a cardiologist previously missed. not because they weren't good cardiologists because they were looking for something else and tiny little perturbations actually that a machine could pick up that humans couldn't actually predict things down the road and there's there's some really interesting published work in mail about that um that it should make our lives better and whether we're talking about it or not I'd hope we're talking about the things that it made better, right? It made it easier to age. I mean, something really trivial. I mean, my mother-in-law is 92, and she asks Alexa to play a Frank Sinatra playlist.

42:35And, I mean, she doesn't have to get up, which is sort of she gets up slowly, and it hurts her. And Alexa plays a list, and she will say, I didn't remember how much I liked that song. Yeah. And how nice is that? and I think it will impact our lives in hundreds of ways. And I think I'm looking forward to the time when we're not talking about the AI, we're talking about the good it did. Yeah. Right, and we're benefiting from the good it did. Right now, mostly we're talking about AI and we've got some work to do as an industry to deliver on the promise. Well, that seems like a great place to leave it.

43:13Very optimistic future. I'm looking forward to that. So Andrew, thank you so much. This has been fascinating. I learned a ton. I'm sure the audience has as well. I really, really appreciate you making the time. Thank you for having me. And thank you listeners for listening in. And congrats again. Thank you so much. Thank you so much for listening to Generative Now. If you liked what you heard, please rate and review the podcast on Spotify and Apple Podcasts. That really does help. And if you haven't already, please subscribe so you get notified when we drop new episodes. If you want to learn more, follow Lightspeed at LightspeedVP on YouTube, X, LinkedIn, Instagram, and everywhere else.

43:53Generative Now is produced by Lightspeed in partnership with Pod People. I am Michael Magnano, and we will be back next week with another fascinating conversation. See you then.

From the publisher

Cerebras recently announced the launch of The Wafer Scale Engine Three. This is the largest and fastest chip in the history of chips. 


This week on Generative Now, Lightspeed Partner and host Michael Mignano talks to Andrew Feldman. Andrew is the CEO and cofounder of Cerebras, the hardware company that builds AI supercomputers known for its Wafer Scale Technology. The Wafer Scale Engine Three is about the size of a dinner plate and promises compute design for AI. Andrew and Michael talk about what it takes to have the conviction to make big bets and push the limits of hardware innovation.


Prior to Cerebras, Andrew co-founded and was CEO of SeaMicro, a pioneer of energy-efficient, high-bandwidth microservers. SeaMicro was acquired by AMD in 2012 for $357M. Before SeaMicro, Andrew was the Vice President of Product Management, Marketing and BD at Force10 Networks which was later sold to Dell Computing for $800M. Prior to Force10 Networks, Andrew was the Vice President of Marketing and Corporate Development at RiverStone Networks from the company’s inception through IPO in 2001.


Episode Chapters
(00:00) Welcome to Generative: A Deep Dive into AI Supercomputing
(00:26) The Wafer Scale Engine 3
(00:56) Andrew Feldman's Journey: From Startup to AI Supercomputing Pioneer
(04:53) The Foresight in AI Chip Development: Outpacing the Giants
(07:32) The Wafer Scale Engine 3: A Technological Leap in AI Compute
(17:12) The Future of AI Compute: Expanding Horizons and Synthetic Data
(23:18) Challenges in Chip Design
(23:51) Innovations in Cooling and Powering Chips
(25:02) Overcoming Precedent Challenges
(26:40) The Revolutionary Impact of Wafer Scale Engine 3
(29:08) Diverse Customer Applications
(34:05) Future Directions in Chip Technology and AI Applications
(39:11) Closing Thoughts


Stay in touch:

More from Generative Now | AI Builders on Creating the Future

All 90 episodes
Andrew Feldman: Building the World’s Largest and Fastest Computer Chip for AIGenerative Now | AI Builders on Creating the Future · 44 min
Listen in VO