In short
Cerebras CEO Andrew Feldman explains why Cerebras built the largest AI inference chip ever (wafer-scale), how it beats GPUs on “decode” speed, and why AI demand is constrained by memory, foundry processes, and advanced packaging. He also discusses the OpenAI data-center deal, multi-silicon competition, and CUDA’s shrinking “moat.”
Guest background
Andrew Feldman is co-founder and CEO of Cerebras. He previously worked on a company sold to AMD and has led chip development across ~20 years, including low-power computing work.
Key claims
- AI “speed” is tokens per second per user; inference latency is dominated by sequential decode.
- GPUs are slow at decode because they repeatedly move model weights from HBM DRAM to compute.
- Cerebras uses a wafer-scale chip packed with SRAM to make weight-to-compute movement much faster (claimed ~2.5k×).
- CUDA is losing durability because state-of-the-art models increasingly train/operate outside CUDA flows.
Notable examples
- OpenAI deal: up to ~750–760 MW via Cerebras cloud/hardware; also AWS deal splitting pre-fill (AWS Tranium) vs decode (Cerebras).
- “Three bottlenecks”: HBM/HBS DRAM sold out, TSMC COOS process constrained, and TSMC 3nm/advanced factory capacity constrained.
- “Grok” acquisition by NVIDIA is cited as evidence GPUs can’t deliver fast inference.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOIntroduction to the Largest Chip
0:00 to 0:36
Learn about the unprecedented scale of the largest chip in computing history.
“This is the largest chip built in the history of the computer industry.”
The Importance of Speed in AI
1:31 to 2:24
Discuss why speed is crucial for AI and how it affects productivity.
“So has speed become the dominant conversation for AI today?”
Understanding AI Metrics
2:26 to 4:26
Explore the key metrics for measuring AI performance and user experience.
“Because those are more productive tokens.”
The Landscape of Chip Technologies
4:38 to 6:36
Analyze the evolution of chip technologies and the role of various players.
“So I'd love to talk about the landscape of the chip industry right now to help people visualize and guard where you guys are.”
The Role of ASICs in AI
6:41 to 9:05
Learn about application-specific integrated circuits and their impact on AI performance.
“Can you maybe define what that term means?”
China's Position in the AI Chip Race
9:18 to 14:00
Discuss China's advancements in AI and chip technology as a global competitor.
“which is a very large customer of yours, and Broadcom, where does it fit in that picture?”
Multisilicon Ecosystem and Data Center Focus
14:00 to 15:01
Discussing the constraints of mobile processing power and the necessity of cloud data centers for heavy computing tasks.
“Is that something that you guys think should be part of the multisilicon ecosystem?”
Market Demand and AI Bubble Concerns
15:02 to 18:06
Examining the market dynamics and potential concerns regarding demand for chips amidst public market fluctuations.
“So you mentioned how powerful the demand was and demand at all production.”
Bottlenecks in Chip Production
18:07 to 21:15
Identifying the current bottlenecks in chip production related to memory and manufacturing processes.
“Whereas what's different about AI right now is we're all trying to catch up.”
Advancements in Chip Technology
21:16 to 22:36
Explaining the evolution of chip technology and the significance of larger chips for AI applications.
“etches transistors that are a certain distance apart and the smaller the distance apart the more transistors per square millimeter and the more you can do per square millimeter.”
Show all 27 chapters
The Role of CPUs in AI Systems
22:37 to 25:26
Discussing how the increasing demand for AI-driven actions impacts CPU consumption and the overall infrastructure.
“And therefore, you get answers in less time.”
Cerebras' Journey and Vision
25:27 to 28:01
Reflecting on the inception of Cerebras, the vision behind its chips, and lessons learned from prior ventures.
“And CPUs experience the same memory shortage?”
The Journey of Chip Development
28:01 to 30:01
Learn about the evolution of chip technology and the unique challenges faced.
“I think around, you know, we have, as a team, built dozens of chips over the past 20-clack years, and the returns to experience in chip making is in Argus.”
Overcoming Early Challenges
30:01 to 35:34
Explore the hurdles faced in the early days of Cerebras and the strategies used to overcome them.
“So at first, we, at the beginning, we were honest with our VCs and told them we're going to attack a really hard problem.”
The IPO Experience and Its Significance
35:34 to 38:04
Understand the emotional journey and pride associated with going public as a company.
“And over that 18-month period, we became the best in the world at it from approximately zero.”
Innovations in Chip Cooling and Reliability
38:04 to 40:51
Discover the innovative cooling techniques and reliability measures implemented in large chips.
“Everybody who'd been with us more than nine years and their families, and we all rang the bell together.”
Understanding Inference and Performance
40:51 to 42:00
Learn why Cerebras' architecture is fundamentally faster than traditional GPUs for AI tasks.
“But again, we had to invent the technique to cool off a chip this big.”
Understanding Model Weights and Processing Steps
42:00 to 45:00
Learn about the process of inference in AI models, including pre-fill and decode steps.
“They can be code, they can be pixels, but call them words.”
Evolution of AI Models and Reinforcement Learning
45:01 to 48:00
Explore how AI models evolve and the role of reinforcement learning in training.
“Do you think that models need to really evolve as well in that fast AI world?”
Challenges in GPU and Distributed Computing
48:01 to 52:48
Discuss the complexities of using GPUs for trading and the advantages of Cerebras' approach.
“And so as the best models, all of them, whether they're domestic or Chinese, whether you're open AI or anthropic or Gemini, whether you're any of the Chinese models, they all move to a reasoning approach.”
Cerebras' Role in AI and the OpenAI Deal
52:49 to 56:00
Delve into Cerebras' business model, capabilities, and the significant partnership with OpenAI.
“Let's talk about the super business a little bit, because you guys are obviously a chip maker and provider, as we discussed.”
Data Center Dynamics with OpenAI
56:00 to 58:39
Learn about the strategic data center capacity deals between Cerebras and major AI companies.
“And that's why Anthropic did a huge and sort of very expensive deal with Elon for data center capacity.”
Cerebras' Unique Cloud Solutions
58:40 to 1:01:06
Explore how Cerebras integrates chips and cloud solutions for AI applications.
“Like talking about this industry is how it seems that flexibility is so important.”
The Shifting Landscape of AI Training
1:01:07 to 1:03:28
Discuss the evolution of AI training models and the decline of GPU dependency.
“of you know well i don't think kudos are most correct korea we should talk about that yep I think two years ago, every state-of-the-art model was trained in a CUDA flow.”
Supply Chain Strategies in Chip Manufacturing
1:03:29 to 1:09:54
Understand the challenges and strategies of Cerebras in chip manufacturing and supply chain.
“And so that's where we're building sort of our strength and that's how we're to deliberate value to customers.”
The Future of AI and Business Tools
1:10:01 to 1:12:07
Explore how AI is transforming business processes and tools.
“like in the next year or two well we know some things we we know that the model we use it you're using today, Fable or Gini Petit Flash 6, will be the worst model you ever use.”
Closing Remarks with Andrew Feldman
1:12:07 to 1:12:17
Andrew shares his thoughts on the discussion and thanks the host.
“What a journey and what an exciting future.”
Transcript
Automatic transcript. May contain errors.0:00This is the largest chip built in the history of the computer industry. It's 58 times larger than a GPU. And for AI, bigger chips process information more quickly, and therefore you get answers in less time. For AI worm, big chips are undoubtedly the best way to go. There's no mode in inference. It takes you eight keystrokes to move from a GPU to us in the cloud. We solved a problem that nobody in the history of computers solved, and we delivered it in 2020, and nobody cared. Nobody cared. Nobody bought it in, nobody cared. Everybody said we were crazy, it would never work, so then we built the next one.
0:40Hi, I'm Matt Turb. Welcome to the Mad Podcast. My guest today is Andrew Feldman, co-founder and CEO of Cerebrus, the company that built the largest chip in the history of computing and just pulled off the biggest semiconductor IPO of all time. Andrew has been everywhere talking about the headlines, the$20 billion plus OpenAI deal, the IPO, but this conversation is a bit different. We started from what is a wafer and built up step by step, why GPUs struggle with fast inference, the three shortages nobody talks about, the decade in the desert when nobody wanted this chip, and why Andrew believes that CUDA is no longer a moat for NVIDIA.
1:18If you want to actually understand the current chip landscape and how AI inference works at the silicon level, this episode is for you. Please enjoy this fantastic conversation with Andrew Feldman. I thought a fun place to start would be to talk about speed. So has speed become the dominant conversation for AI today? Yeah, what happened, I think, was for a long time, AI was sort of a novelty, right? It was like a parlor trick. It was cool, but not useful. And what happened somewhere around the middle of 2025 was the AI got smart enough such that people began to use it. And we remember, we make AI with training, but we use it with inference.
2:05And suddenly people wanted to use it. And the minute you want to use it, the minute it's productive, speed matters. Fast tokens are more productive. And so the conversation moved from everything else to how do we make our inference faster? How do we deliver tokens more quickly? Because those are more productive tokens. We get more done in less time and therefore it's more valuable. Great. And what does speed mean? Is that a question of token speed? Is that completion of the task? What's the right metric? Sure. The right metric is tokens per second per user. That's how fast you get the first token all the way through the last token in your response.
2:52And it's true for queries from chat, but it's also true for agentic floats, right? if they're sort of multi-cycle turns, waiting is amplified. And so what you want is blisteringly fast responses so that the AI feels like it's in real time. You can engage with this. So it's the broadband moment for AI. I think that's right. I think that's a very good analogy. I think if you think of something like Netflix, when the internet was slow, Right. Netflix delivered DVDs and envelopes. You would get a DVD and an envelope. And when the internet became fast, they didn't get more efficient at delivering DVDs and envelopes.
3:43They became a movie studio. right the speed enabled them to become something completely different and that's what speed does uh in general and in particular for for ai it opens up a whole new domain that allows you to use the ai differently you will stay longer you will come more often and you work on harder problems yep so it's literally a question of the ux right that's just nobody wants to wait a few seconds That's right. I mean, how big is the market for slow search? How big is the market for dial-up? It's zero. How long will you wait for a website to resolve? Will you wait eight seconds? Nobody waits.
4:23And so it's the exact same with AI. Yeah. So no more people are waiting with their laptops open while the agent... That's right. Well, it's running and running and running. I think that is not what people want. Okay, great. Wonderful. So I'd love to talk about the landscape of the chip industry right now to help people visualize and guard where you guys are. so um there used to be basically this concept of one chip to do it all uh and uh obviously this is evolving dramatically um this you guys uh this this grok uh this there are gpus there are then people have a part of tranium people around as tpu so help us compare and contrast who does what for what?
5:15Well, I think there used to be one chip to do it all. It was called the CPU. That was right. And there emerged a coprocessor to do discrete graphics. And as the AI workload became interesting, we focused chips more and more on that particular workload. And today, several companies make traditional GPUs. So NVIDIA and AMD, they make very standard GPUs. There are a group of companies who the hyperscalers make some of their own parts. So the TPU. So to me is Google. TPU is Google and Tranium is AWS. And then there were a group like Strabris. and we were among the pioneers to build a part from scratch, optimized for AI and nothing else.
6:14And that we weren't inside of a hyperscaler, we weren't optimized for a hyperscaler's problem or for one lapse problem. We were building a chip for a collection of AI problems. And all our thinking was around how to accelerate AI. And that's sort of the landscape to this day. Yeah. There's even more specialized ones, right? Like developed specific transformers. Those are called ASICs as well. Can you maybe define what that term means? ASIC is an application-specific integrated circuit. And it's a word that now has a wide range of meaning. It means that you have made a series of choices away from general towards a narrower class of problem solving.
7:11And that you've made some decisions in your architecture that make it much better at some things and much, much worse at others. Right? And that's choices that are made across the spectrum. So, TPU has made some choices like that. it can't do graphics. It's very good at matrix multiply. It's not very good at a collection of other things that we do in mathematics. Same for the GPU. We've all made different choices. Right now in production, there are Iridi and AMD, there are a collection, there's the TPU for Google, Traniums, just coming up as a Maya part, is a part from Microsoft, Cerebrus, and there was one other Grock that got acquired by NVIDIA.
8:08And to the general chip versus pistol-aged chip, it was actually interesting because you guys and your CTO celebrated when Grock was acquired. Was that, what was that? Was it a recognition by NVIDIA that your vision was right all along? Yeah. I think one of those ideas sort of most durable modes was the perception that the GPU could do everything and it was all you needed for AI. And the acquisition of Grok for$20 million and the structure and the speed with which they chose to do it made clear to everyone that that wasn't true. That the GP architecture couldn't do, could not do fast inference and that this market was large and growing quickly.
9:00And we were the fastest at it and the largest. And, you know, our sales were more than 10 times the rocks, and they paid$20 billion for the number two collect. So that was a good day. That was a great day. And the recent announcement of Holotino between OpenAI, which is a very large customer of yours, and Broadcom, where does it fit in that picture? at one of those highly specialized azx types right grok because it's for inference that's right jalapeno is a is a part uh that has been a long time coming um it was uh announced with broadcom before uh openai did the deal with us remember we did a huge deal this is quite the largest deals in Silicon Valley history, north of$20 billion.
9:58I think they have a yawning need for silicon. And one of the things that OpenAI has been, I think, the best in the industry at is looking at an exponential curve, adoption, rate of growth, not being afraid of what it says. right? Others have had to go out and strike really bad deals to get capacity. Where open AI saw this coming, they struck big deals for memory, for compute, with us, with others. They've really been sort of visionary in understanding what it means to extrapolate from an exponential curve. And that curve is AI usage, right? It is growing so unbelievably fast. the whole industry is chasing demand things right we usually well often it's the other way often people are building it hoping it will come and in our case all of us are are chasing what people already want to do let alone what they might do in the future hey ed as a point of opinion that i mean i mean um because of this such a interesting topic um to to finish on gps versus inference versus ASICs.
11:20Is that the permanent situation from your perspective that we're going to be in this forever multi-silicon kind of environment? Yeah, multi-silicon environment is a healthy ecosystem, right? I don't think anybody would say that the x86 environment was a healthy place. For 20 years, there were two players. And if you notice, when a new workload came around, and cell phone workload. It was very similar, but required low power and battery and both lost, right? Intel got zero share. MD got zero share. And a company nobody previously heard of became the largest seller of compute in the world. And that's our A.
12:04And that healthy ecosystems have lots of different ways to solve problems. Where does China fall in all of this? So this Huawei send chips for DeepSea can this emergence of a full stack Chinese AI factory for like a better term. Is that is is that divided happening or is that overstated? No, no, no, that that divide is real. No, I think they are an industrial adversary that I take. They have made really interesting investments that made investments in power. the grid which is a real weakness in the u.s um if you're in france you have nuclear which sort of turns out to be pretty cheap power and pretty clean um but uh in china they made huge investments in their grid uh they are behind in chips um but their approach was at the next level as open source models where they're producing some extraordinary models not as good as as as gpt or anthropic or google's gemini but very good hey and i i think um uh they don't have the the same chip strikes that the us does but they've got leadership in other domains and potentially in power which is what we need for beta centers okay what does the china fall in sero race strategy we don't sell questions uh for geopolitical reasons for business reasons for regulatory reasons for regulatory reasons and and for some uh geopolitical reasons okay um is there a role for local AI and local chips?
13:57NVIDIA had some announcement around just building chips for Windows computers. Is it for inference? Is that something that you guys think should be part of the multisilicon ecosystem? Yeah, I think if you look at the way the ecosystem for apps and cloud emerged, everything you can do, you should do on your phone or your laptop. But the ability to get real processing power to a phone or to a laptop is constrained because they're generally working off a battery. And so you want to do as much as you can, as close to the data as you can. But the truth is, in most situations, for real compute, you have to go to the cloud to the data center.
14:44And that's exactly the way it's going to be with say we're going to do a lot of work on the cell phones, on the laptop. But for the big work, you're going to go to the data center. And that's where our focus is. Our focus is data center compute free. Great. So you mentioned how powerful the demand was and demand at all production. So just to ask the inevitable question around market equilibrium bubble or not. And so the most recent thing we all saw is that there was a little bit of a flash crash earlier in June where 1.5 trillion in market value was lost. And NVIDIA lost 300 million that day. Now that was corrected within 48 hours.
15:33so the market seems uh very jittery uh around uh that that general concept but just uh a little bit towards senior last one just based on what you said those sounds like you're you're reasonably on the other side of that it's okay concern uh i i i think that that if you spend a lot of time staring at the at the public markets um and watching their ups and downs every day you're watching the wrong thing. I think that the wisdom I received from the great thinkers of public market investing is that in the short term, the public markets are voting mechanisms, who's most popular. But in the long term, they're weighing mechanisms, who's created the most value, the most weight.
16:25So what we can do as an only public company is we cannot look at the day-to-day fluctuations and we can focus on just building extraordinary technology, winning new customers, making our existing customers happy, and sharing they buy more, being better every week. And so I think the day-to-day fluctuations are out of our hands, and that's about the voting. Yeah. And maybe just to play devil's advocate on the demand, it seems like a lot of the demand comes from the big labs, which themselves are financed by venture capital, private equity, hedge funds, where you call it, sovereign investors. is there any concern that demand for chips comes from the labs which are maybe artificially financed and if you add a few circle of deals here and there then you sense any infragility?
17:26That's not exactly our experience. I mean obviously we have enormous demand from Polopanam but we have dozens of other customers who are trying to place in very big orders
17:45And historically, bubbles were when supply got out ahead of demand, right? When in the 90s, we built out telco infrastructure, right? We built out fiber years before it was going to be used. And it took six or eight years and it all got used. But it was a sort of if you build it, they will come mentality. Whereas what's different about AI right now is we're all trying to catch up. We're trying to build data centers faster. We're trying to increase our demand, our supply chains for demand that's already here today. And only a very small portion of the world are using AI anywhere close to its potential.
18:32And we're already sort of overwhelmed with compute. We're overwhelmed with the demand for memory, which is a real weakness in the GPUs. It's not a problem we face. In the ecosystem, there are constraints left and right. And that doesn't feel like a bubble. Did you want to actually double check on that? Because that's so interesting, right? It seems like the bottleneck keeps moving. So what's the current problem of memory shortage that people may have heard of? So there are three major bottlenecks right now. The first is all GPUs and all ASICs except us use a type of memory called DRAM and a particular flavor of that memory called HBS.
19:30And that memory is made by three companies in the world. The Hydex, Samsung, and Microsoft. Three companies. And they're sold out. That's the number one problem. We don't use it. And they're sold out as in, like, the only way to, as all things hardware, the only way to increase supply wouldn't be to build new factories? Build more factories. And so you have this huge problem for GPUs, and even now for CPUs and you can't get this memory. Because of our architectural choices, we don't use it. The second constraint was a process at TSMC. Process was called COOS. And this used silicon as a motherboard.
20:22And this is what Inxadia and AMD and others use as a motherboard. So they put their chips in the memory chips on it. And it's more efficient than a traditional motherboard. That process is sold out. It's highly constrained. We don't use it. Third limitation in the space is three-danimeter factory space at TSMC. So this is the most aggressive factory. So with the most advanced technology, our chips are at five gattamater so we don't use the three gattamater technology and so and that's so fascinating so to ask the layman's question so this this what does that mean 3 millimeter factory that's a factory that's solely dedicated that's right they build factories and the factory etches transistors that are a certain distance apart and the smaller the distance apart the more transistors per square millimeter and the more you can do per square millimeter.
21:34And so the history of the computer industry making chips has been we've gotten better and better at putting transistors closer and closer. And so we used to be at 16 nanometers apart and then we went to 10 and to 7 and to 5 and to 3 and they're working on 2, 1.8. And so right now, the bulk of the GPU world is off-free. And there's tremendous congestion at the factory. And that's literally different tooling at the factory? Tooling at the factory. And we use the 500-meter factory. And this is our shit here. Yes. This is the largest chip built in the history of the computer industry. Everything is 58 times larger than the GPU.
22:28It has, you know, two and a half or 3 ,000 times more memory bandwidth. And for AI, bigger chips process information more quickly. And therefore, you get answers in less time. And obviously, you don't want a bigger chip for a laptop or a cell phone or for lots of other things. but for AI worm, big chips are undoubtedly the best way to go. Okay. That's unique. All right. So I want to do a bit of a deep dive on the chip itself in a second. But to close on this, so three bottlenecks. You also mentioned CPUs a couple of times in this conversation, and there seems to be a theme around the emergence of a CPU shortage.
23:19as well is that so one is that true to what causes it so agentic ai is a world in which ai doesn't just provide answers it initiates action so that action might be go to a website it might be learn gather some data from a website bring it back take another action those actions are done by CPEs. And so as AI gets better and better at doing things, at making instructions, calls for things to get done, we're using more and more CPEs. And that is driving up the consumption of CPEs and therefore the demand for CPEs. And so this sort of huge push for more CPEs is being driven by AI on machines like ours and GPUs doing agentic work and asking the CPUs to take an action, to go to a website, to order a burrito, to find a piece of information, to pull it from storage to all that work is being done by the CPUs.
24:39Including in a system like yours? Including in a system like ours. Oh, it's interesting. So what does it mean? So you have your chip and CPUs on the side, and those will be a system? It means that AI, when we do AI, the agentic, the AI processor is like the brains, and the CPUs are like the body. They're taking action. They're doing things in the digital world under the direction of the AI which is running on the accelerator, on the streamer system. And so as we do more and more AI work and more and more agentic work, we're making more and more calls to CPUs. And therefore, the demand for CPUs is through the roof.
25:30And CPUs experience the same memory shortage? to say his memory sure it's the exact same number that's right okay yeah let's go back to the products in a second but um uh let's talk about uh you and your journey a little bit uh for people to have context so you know obviously you've been pulled and all of an incredible um idea always perfect timing um and i know i know it was a little bit uh the way you have perfect timing is to have horrible timing for 10 years. So you have perfect data. Very much to this point, so you guys started building this bigger chip focused on inference over a decade ago.
26:132016. 2016, which makes perfect sense in retrospect, given how long it takes to build a technology like this. But what was the vision? 2016, I guess it was like four years after ImageNet, deep learning was a thing. It was all vague. It was all kind of revolutions. And that's why we don't believe the right thing to do is to embed the latest and coolest model into your circuitry. That's a mistake. We're the fastest in the world of transformers and our architecture was set before transformers existed. We're the fastest in the world of diffusion before an architecture was set before What you want to do is take the underpinnings of those so that when the market moves, you can be good at that as well.
27:04Otherwise, because of two-to-date short lines. To the ASICs question earlier, so that's what others do, they build the architecture of the model into this? Some have. Some have. Historically, that's been a structural mistake. It's been a structural mistake. So you're doing the brawny or as Intel platform on which any kind of model would want. you want to think about what is the underlying calculation that the underlying calculation at all this work is sparse linear algebra yeah and if you can accelerate that whatever the bottle makers invent you can make faster okay and that was our approach yeah so you started in 2016 um you had a prior company that you sold to your amd what were some of the lessons you run there that you took into your cerebris?
27:52I think the lessons are large and many. I think one of the things, I guess, experience is another name for having made mistakes and learning from them, right? I think around, you know, we have, as a team, built dozens of chips over the past 20-clack years, and the returns to experience in chip making is in Argus. And we built a different type of computer at C-Micro, a type of computer optimized for low power and optimized for a workload that was very different than AI for something like web browsing. But the fundamental underpinnings, the questions you ask as a computer architect are always the same.
28:41What can I do to make this work faster? And is there enough of it to make it worthwhile? These are the two questions we ask. Should we build a part for it? What could we do to build a chip optimized for AI? And will there be enough AI so that you can build business around it? Those are the questions we asked in 2016. And, you know, the flip side of that was, wouldn't it be a surprise if the GPU, which had been optimized for graphics for 20 years and been pushing pixels to a monitor, was suddenly good at a new world? Wouldn't that be serendipitous? And we came to believe that it wasn't the right architecture for it.
29:34It was just better than the CPU. late and that we could build an architecture that would be faster faster they would use less power and could drive down the cost of of it and that was the journey so service was very early but uh then we spent time in the desert we wandered in the desert yeah maybe walk us a little bit through those years for any area you know founders especially ntpec founders are listening to this so what was so first of all what was the issue was it marketing was it the technology was not working uh and then how did you go about it as a team you know i get your you'll be on board and your investors and raising more rounds so as as you presumably didn't have the proof points that you one of the need.
30:26So at first, we, at the beginning, we were honest with our VCs and told them we're going to attack a really hard problem. We worked in a build, something that was a little bit better than a GPU. And that our idea, our strategy was that you will never beat a great company like NVIDIA by doing something a little bit behind the feet. That they're going to buy everything for less. They're going to have price and pressure. They're going to be able to bundle. And the right strategy would be to do something incredibly hard in engineering that was way better. 10, 15, 20, 30, 50 times better. But to do that, ordinary and obvious paths are all closed.
31:11Everybody else has taken the bar. And so what we observed was that speed in inference was going to be a function of memory. And that there were two types of memory. There's this D-RAM of HBM, and they can store a lot, but they're sore. There's another type of memory called ESSER. It is unbelievably fast, but per unit area can't store very much. And for graphics, everybody had always used D-RAM, they'd used HBM. and that the AI workload was fundamentally different. In graphics, you move data to the GPU, and then you work on it for a long time. And then you send the result. So the time spent in total of movement plus work is dominated by work.
32:15In inference in AI, it's the exact opposite. you move a huge amount of data, all the weights from memory to compute, and you need one calculation to generate the next word. And then you have to do it again. So all the time is dominated by the movement of data. So that's why GPUs have so much trouble being fast. So we observed that if we chose a strategy, in S-RAM, we could be faster. But then we have to overcome the trade-off of SRAM that it can't store very much. That led us to the solution that if we could build a part vastly larger than any part in history, the size of a dinner plate, we could stuff it to the gills with SRAM and thereby get the benefit of SRAM that it was fast and overcome the weakness that it can't store very much.
33:12And that led us to a strategy called wafer scale. This chip is made from a single wafer. It comes out of TSMC. Do you want to define what a wafer is? A wafer, all chips are cut from a wafer. A wafer is a circular piece of silicon that's 300 millimeters across. And the process of chip making stamps out, like your mother does with a cookie cutter, stamps out chips. speak. The biggest chip that had ever been built before us was 800 square millimeters, 840 to be exact. And this is 46 ,000. So we had to invent all sorts of new technology to build a chip this big. And once you build it, you have to invent ways to power it, to cool it.
34:03There are no vendors waiting for you. Yeah. Right. Because you look like nothing else ever made. And that took years. And it was a deep tech problem. It had never been solved before. And we had an 18 month period where we were spending 8 million a month and we couldn't build them. So if you're deep tech founders, you have 8 million a month. Yeah. Why so much? What was the cost just to understand how these businesses? Because what everybody thought was hard, we solved quickly. And what nobody else knew about, because they'd never actually done it, turned out to be really hard. Imagine, I tell people, imagine the first group that was going to climb Everest.
34:50And they're at base camp. And they're having tea with a group that just failed. And the group that just failed says, halfway up, there's this part, it's unbelievably hard. We couldn't do it. Okay. your team climbs up, makes it all the way to the top, comes back. They're having tea again. And the team that made it leans over the team that hadn't made it and said, that part in the middle, that wasn't the hard part. Because nobody had gotten past it. Nobody had gotten past certain things, so they didn't even know what to be afraid of. We now know. And it was something, a step called packaging. And that's how you affix a wafer to a motherboard, how you deliver power to it, and how you cool it.
Read the full transcript
35:38And nobody had done it before. And over that 18-month period, we became the best in the world at it from approximately zero. And we did that by sailing again and again and using good engineering methodology and doing a failure analysis, every single failure. So we failed differently again and again and again and again. And we told our board, you know, we met with our board over six weeks. And yeah, this is a strategy. Nobody's ever done this before. This is what we're going to do. And we had to invent new materials. We had to invent new techniques. We ended up building things that everybody else had partners who could do.
36:20But in all this to 2019, we announced we'd solved it. And, you know, my co-founders and I, the first time it worked, it was sitting in a tiny little lab. And we couldn't believe it. We were the first few things this tree gave me to make one work. And we just stared at it. You know, watching a server run is about as exciting as watching paint dry. And there we were, the five of us, just staring at this machine, not believing it might work. We made it work. Was it a bigger moment that actually rang the bell? or just completely different events? Like emotionally? Emotionally, it was a completely different thing.
37:00It was that we had made our idea work. And I think for deep tech founders in particular, there's always this little thing in the back of your mind that says, maybe it's shit. Right. Maybe we're actually crazy. That's right. Maybe it's going to fail. And maybe it's not going to work. And maybe I don't have time. Maybe we're going to run out of money. maybe, maybe, maybe. And the flip side of that is the joy that this was my co-founder's ideas. I mean, these are their ideas manifest in the world. And that is an extraordinary thing. And it's when someone's ideas take physical shape. And then the next step is when you watch other people's work sit on top of your idea, then you know you love making infrastructure.
37:50And so when we rang the bell and we went public on May 14th this year in the largest semiconductor IPO in history, and we did something unusual. We invited all our engineers who'd been with us to start. Everybody who'd been with us more than nine years and their families, and we all rang the bell together. We all got up on stage. which that was sort of a moment of, oh, a different pride that we had done this together. And that we should manage not to die. It's untied with a startup, you know, people aren't honest. They don't tell you that, but we'd avoided some, we'd made plenty of mistakes, but we've been avoided the sad ones.
38:42and we made it through to a level of success that gave us the opportunity to pursue a new level of success. That's what an IPO is. It's not the end. It's you've gotten to a plateau that you can chase a new level of success in the public market. And that felt pretty good, I'll admit it. Amazing. Thanks for sharing. So just to go a little deeper on some of the product and technical stuff. So one of the trade-offs of building a bigger chip, one that may come to mind is failure mode. Sure. If you have a lot of little chips on a big wafer, presumably you can isolate the problems. Yeah. If you have a big one, then everything won't fail at the same time.
39:37So you have to think very carefully about failure mode and we invented a technique that had about a million identical tiles. And if one fails, we can shut it down and we can use redundant ones, we keep going. And so if you're going to go big, you have to think about in the very architecture of the computer, how you're going to manage value. I mean, the JPs have a huge failure rate. So I'm sure you guys have spoken about this. Infant mortality is enormous, and they fail all the time. There's big data. Facebook put out a paper on the number of failures they get in a big cluster. Now, we can do some other things.
40:23Because we have all this compute in one spot, we can invest more to cool it. So we pioneered water cooling in AI systems, and we run these much colder. and GPMs. And the failure mode in electronics is temperature. And so by running them cleaner, we are more reliable. And so that was an advantage. But again, we had to invent the technique to cool off a chip this big. And so I think we have a system mentality, right? If you're going to do something big, you've made trade-offs. And you have to think about an architecture that allows for redundancy and repair you have to think about the pros and cons of every aspect of the architecture and that's how you do something different and new great and again just to drive the point home um to make sure this i guess clear takeaway uh you know for people from this conversation.
41:31So explain it to me, like I'm maybe not five, but 15. Why is this faster than a JPU? Like what fundamentally means it faster and what can't you be super fast with a JPU? How tokens are generated in inference is why it's faster. In an inference to generate a word and our answers are a whole stream of words. They can be code, they can be pixels, but call them words. The weights of the model are moved from memory to compute. Calculation occurs and that generates the word. Is it pre-fill versus decode? That is both steps. Okay. And do you want to maybe explain what pre-fill and decode are? Okay. There are two steps in the computation necessary to do inference.
42:34When you type in the chat GPT, explain to me the history of this village prior to World War II. And it can't see you. Two things have happened. The first thing is your prompt has been processed. That's step one. And step two, your answer has been generated. And the way your answer is generated is called decode. And decode is sequential. Processing your prompt, which we'd call pre-fill, can be paralyzed. So you can process many of them simultaneously. But the speed with which you get an answer is a function of the decode, and it is step-by-step-by-step in sequence, and that can't be changed. And how you do that step is you move weights, which are the intelligence from the model, to compute, to generate a word.
43:47You do a calculation, and you get a word. And then that word is used to generate the next word where weights are moved from compute. So the process is one of moving weights from memory to compute. So how big are the weights? In a little model, like a 70 billion parameter model, the weights are about the size of 100 HD movies. So to generate a single word, you move 100 HD movies from memory to compute. And then you have to move them again for the next word. And you want to do this a thousand times to get a good answer. A thousand words. This is where HBM is slow. This exact step is where HBM is slow.
44:34And that exact step, it is where by having all the SRAM here, we're blisteringly fast. And so the speed of moving waits to compute is about two and a half thousand times faster here than on a world in GP. And so that's the essence of what we've been able to do here and why it's so much faster. Yeah. Fascinating. Do you think that models need to really evolve as well in that fast AI world? Or is that just a problem? No, exactly. Because remember, two things are happening. One, we want the models to be smarter. And two, one of the ways models are getting smarter is with RL. And RL uses inference inside a training.
45:29And so the faster you can do the inference inside a training, the faster you're training. One, two things. Is that a market for you guys? Yeah, that's a market for us as well. So you're not just in France. We do RL and we do traditional trading too. Not for the largest models, for the largest lab, but for the next tier. For pre-trading, for pure pre-tuning. Pre-trading, fine-tuning, a full set. So in Chengdu, pre-trading, I mean, post-trading was RL, some pre-trading for the other labs. GPUs are still better than you for what job? So GPUs in training have some challenges that have been solved by very narrow selections of the community.
46:29GPU is a very small chip. And the calculations that we need to do in trading are very large. And one of the most complicated parts of trading is the breaking up the calculations and spreading them apart on multiple GPUs. And that's called distributed compute. That has historically been the domain of the supercompute world. and is very difficult, not just because cracking a problem and having lots of others work on it is hard, but they have to constantly share information. And that sharing information is why they needed to buy Mellanox. They needed to control a fabric over which all is sharing in order to get an answer.
47:20That breaking up a big matrix multiplier, that big calculation is called running tensor model parallel. You are breaking up the tensor and spreading it apart. And the best labs in the world are good at that. But nobody else is. When we run training, we don't have to run that way. We run what's called data parallel. And data parallel is very simple. And so it allows teams who are good and very good to quickly test ideas in training and so we are easier to use and faster because we allow them to use a technique which is much simple let's talk about the energetic world maybe starting with with um with reasoning um i think i saw a blog post where you guys have a given state of that reasoning was not always the right solution for all problems how do you think about this and where does that fit in the answer to it well reasoning is a technique a little bit like when you're in 8th grade and you wrote different drafts of the paper
48:44single shot in front of the seat you get an you write a query you get an answer it the easiest way to think about reasoning is you're going to do several drafts it's going to break the problem up it's going to solve in parts it's going to bring them together it's going to review the results it's going to improve the results and then it's going to give you an answer and so that is going to take more compute And if your computer is slow, that's going to be more than an irritant. That could be crippling speed. And so as the best models, all of them, whether they're domestic or Chinese, whether you're open AI or anthropic or Gemini, whether you're any of the Chinese models, they all move to a reasoning approach.
49:38But that meant more compute was being used during inference. That was a huge advantage for us because we were fast. And it made the GP slowness stack up. And it made our speed have an even bigger advantage. And so this was a huge building for us. We think this is going to stay as a fundamental essence of the way these models are run right now. And you also wrote about verification and whether what was the bottleneck in agents was whether the reasoning was good enough, whether they were smart enough or whether there was a verification problem. Sure. I think the verification problem is a little bit like the guardrail problem.
50:34And what you'd like to do is after you've written two or three drafts of your answer, you would like to be sure that it wasn't wrong. And that's your verification step. And you can do that maybe with a different model. You can do that by asking your existing model a similar question in a different way. All of these are ways you can pressure test your result. Guard rails can work the same way. You want to review either with a model or with another technique through a scoring mechanism that this question isn't out of line. that this answer isn't about how to make biological weapons or calling upon information that you are direct to the FBI, right?
51:35And all of that takes compute time. And so whether you're trying to improve through reasoning or whether you're running guardrails by being faster, you can get results in less time. So what do you think the world is evolving for agents? Is it a bunch of smaller models running faster, doing more verification versus a large model? But I think those work together. I think your big model produces an answer. Then you want to double, you just want to check your data. I mean, in the journalism industry, people would write papers that have data checkers, right? Somebody would go and make sure back in the world, journals check data.
52:19In fact, the data checker, right? That was a job. Each claim was checked independently. that's a different model right the main model wrote the piece and then a little model checks some of the answers and i think that's a very a very good way to go by where does a multi-modality for the larger models falling in your world we just announced uh sort of that we were fastest in the world on on one of google's multi-modal models i i think the truth is is that there's very little text that doesn't have charts and graphs right you you must to understand text uh be able to understand uh illustrations graphs uh charts and so that that's sort of the the first and easiest part and then you ought to be able to create both right and then you want to understand images and uh i think the the new models are very very good at that obviously what follows that is video because a video is just a collection of images um but that takes an enormous amount of compute right now and that's one of the resources in the sort of set aside by the leading labs and so unbelievably computation intensive Great.
53:51Let's talk about the super business a little bit, because you guys are obviously a chip maker and provider, as we discussed. You're also a cloud provider, data center provider. What's one of the different parts? we may compute. And that computer is optimized for AI. It is the fastest at AI in the world. If you have a data center, we will sell you hardware for deployment in your data center. If you don't have a data center and you'd like to rent it by the month or the year, we have data centers. So you can rent our equipment through our data centers and through our cloud. And so that allows us to get AI natives as well as large enterprises and governments.
54:44And what's the pressure? It was 50-50 last year. And I think this last quarter, it was maybe 75-25 in favor of hardware sales. I think this year it might be 50-50. And as our OpenAI deal continues to unfold, it will probably be 30-70 with 30 on-premise deployments of hardware, the 70 class. Great. Let's talk about that OpenAI deal. since it's such a major historical milestone record making. So it's providing up to 750 megawatts, which is interesting, by the way, as a metric, because we're in a chip provider, but there's this power. So is that shorthand for... It's a shorthand. It turns out right now, and we didn't talk about this because there's sort of in the adjacent supply chain, we went through the shortage of memory, We had to do the shortage of a process called COAS at three nanometer capacity.
55:50The other limitation in our industry right now is data center availability.
55:57And that is a limiting factor for everybody. And that's why Anthropic did a huge and sort of very expensive deal with Elon for data center capacity. Our deal with them, with OpenAI, was because data center capacity is a limiting constraint, measured the way data centers are measured in megawatts. The deal is 760 megawatts, 250 megawatts in 26 on a multi-year lease, an additional 250 megawatts in 27 on a multi-year lease, an additional in 28 on multi-year lease. And you're doing the data center for them, or are you providing the chips that are going to the data center for them? We're delivering a full cloud solution, so they connect to us via an API.
56:52Basically, you mentioned 2026, do I say it's immediately about general warning? I am looking for data centers. My next meeting is in fact with a data center provider. we're doing a lot in Europe right now a lot in the Nordics is that because it's closer to power sources? yes, it is because there is low cost power how about that? clean, low cost power and low cost doesn't matter in their time business whether data centers is located in terms of yeah, there is an additional latency called transport latency. And that's the speed of light through fiber to get from Helsinki to New York. And you have to account for that if your customer is in New York and your data center is in Helsinki.
57:50And it's usually about two-thirds the speed of light, in case how long it takes. You would like data centers on the same kind, both of them. That's the open AI deal. There was an exciting deal with AWS as well where you... It's a code chip solution? That's right. It's the disaggregated solution you mentioned before where their pranium part is doing the pre-fill and is doing the parallelizable step. So it's pranium. It's pranium. It's doing the pre-fill step and our chip is doing the decode. And so you get a lot more bolusierably fast tokens. It's a good deal for us. It uses their data centers.
58:38So these are deployments in the AWS data centers. Yeah, it's fascinating, right? Like talking about this industry is how it seems that flexibility is so important. There's just something to bring a solution like you need chips, you get chips, you need data centers. Everybody's buying from different suppliers to reduce dependency. I think that's one of the reasons why we went from being a traditional shipping system provider to also offering data centers is that what our customers want are fast tokens. And anything we can do to make the delivery of fast tokens easier. For some of them, that's in their data center.
59:20For some of them, it's with an API. Just point your traffic to us. which will point the fire hose to a fast token back. Yeah. And just to get a sense for where you start and where you stop, is you do not at least currently provide the cloud version. So if I want to run Kimi, like I know you have like incredible stats for Kiwi and Gemma and Chitospe, and I forgot to mention them, but you don't run those as a service or do you? We do. You do. Okay. So you have a service who is still competing with the base tents and fireworks? Yeah. I think we have an on-demand service where you can come to our site and book a month.
1:00:10I think you can even buy buckets of tokens for Kimi or GLM or some of these models. Many of our customers come there and get excited about it and then you move to a dedicated offering where they take hundreds of machines for a year or two or three or four once they've proven out that um the benefit for them in their work often they do a b tests not surprising people like faster you got so it's just more like a testing it's a full environment okay um you can you can go and use it it's at three births.ai you can play around fascinating but that could become like a yet another big clock at this time okay so you have uh ships you have data centers and you have a client business running in france on um thinking about modes so you know famously uh nvidia has uh kudos and well discussed mode what's your equivalent of you know well i don't think kudos are most correct korea we should talk about that yep I think two years ago, every state-of-the-art model was trained in a CUDA flow.
1:01:30And right now, Gemini is trained without CUDA. Anthropic Law is trained without CUDA. OpenAI is trained without CUDA. So in a one or two year period, they lost 70 % share of training models or is the state of the art. Because Gemini is trained on TPs. It's very good. Out of anthropic is trained oxygen. And so I think the storage of the mode is still present, where the data show the mode is clearly shrinking. There's no mode in inference. It takes you eight keystrokes to move from a GPU to us in the cloud. and it's for a seat oh it's each cheese cheese cheese for sure that's it to move your traffic from gpu api to us and so uh obviously gluda was sort of enormously important the creation of our industry and in allowing the the graphics processing unit to be more general than graphics its process.
1:02:51But since 2324 quick live, I think it's ability to serve as a durable moat, short, substantial. And you have a whole ecosystem strategy as you think about your moat, to the extent that any moat can be present in this industry, right? You build a whole ecosystem? Is that, is that? Yeah. Yeah. We built an ecosystem. I think, um, our moat comes from the fact that, uh, by virtue of our architecture, we are doing things no one else can do.
1:03:26And it's not that they can spend more money or they can't pay more for this. If you want fast, you can't have it. Well, I mean, Jake, it doesn't work. And so that's where we're building sort of our strength and that's how we're to deliberate value to customers. I do think about supply chain. We mentioned supply chain constraints for others, but what are your supply chain constraints? Are you in your old TSMC? We are TSMC. We have very close collaboration. They were investors of us. They've been exceptional partners. I'll tell you an unusual story. In 2017, we showed up and we were about 30 guys total.
1:04:21And we showed up in August. Horrible trend in Taiwan. I don't go to Taiwan in August. It's brutal. Not that it's so nice here today. It's only 90 years. You're embarrassed today. Yeah, but also the humidity. You got put on a suit. And we met with the leadership of TSMC. We said, we would like a little pipsqueak company. We believe we can solve a problem that nobody solved in history. And here's how we would modify the way you make chips to make this possible. They thought about it. And they said, we agree. Let's do it. In the meeting. In the meeting. In the meeting. It wasn't going for a month in the meeting.
1:05:11Was that a prepared mind or they were just exceptionally fast on their feet? First, the salesperson had gathered the decision-makers.
1:05:26Second, our proposal was sort of really good at allowing them to use what they were good at. And it didn't require them to change a huge amount, but it did require them to make real changes. And I think they saw this as sufficiently bold that they would learn as they did it. And they also knew that AI was better on big checks. And so the combination of fair of mind, willingness to take some risk, bold thinking from a very large company. Fascinating. It is fascinating. I mean, that's how big companies win. Isn't that right? And how rare is that? It was extraordinary. And what happened next? Like, how long does it take between a decision in a meeting like that?
1:06:22It sounds exceptionally fast to... Two years? Two years. chip making is a long hard process and and most of the time your first chip isn't a winner and there are lots of of startups now sell them with really smart guys like their first chip will not be order the tpu google had some of the best guys in the industry first chip wasn't a winner nor the second nor the third fourth was really good now they're on their eight, then it's a really good chip. The Anna Plena team at AWS, first chip wasn't great, second, third chip was really good. It takes time. And so we built a chip, we delivered it, and this gets to an earlier question you asked.
1:07:10We solved a problem that nobody in the history of computed solved, and we delivered it in 2020, and nobody cared. Nobody cared. Nobody bought it in, nobody cared. which you were like oh my god nobody everybody said we were crazy it would never work it out works and like nobody wants on it is so then we built the next one I mean the first one we probably sold quite a year of set and nobody wanted in because the marker was not ready because I'm product was not good enough um nobody wanted it because ai was a hobby it was right and who cares if your hobby's really fast so um you care about fast when it's in production you care about fast when you use it every day and so we built another one and you know that one we sold three or five hundred yeah hey we built the third one and we sold tens of thousands right yeah amazing yeah isn't that interest to aim it um and um so going back to supply chain uh do you need to think about uh on shoring diversification so it's very hard to diversify away from tsmc chips are so hard and you actually when you design a chip part of the design is for the rules of that factory right so you can't take your design from tsmc and go to somebody else because a huge amount of the work has been to be sure your design is within their rules.
1:08:49And so England, I think only with one or two exceptions in history, each chip generation goes to one fat. So we're going to be with the SMC for our next generation as well. I think we have a supply chain that is built in many parts, but we bring the chips back from TSMC to the U.S., we package in the U.S., and we assemble in the U.S. We do our manufacturing in the U.S., and then we ship from the U.S. I think when you're growing as fast as we are, there are a whole range of garden variety supply chain challenges. A vendor screws up a batch. It gets stuck in customs. The number of ways that things can go wrong, the supply chain is unbelievable but we manage these every day and we're increasing our manufacturing through quite exponentially and so it's really uh that part of the business in chrome really incredible so maybe to zoom out as a last question um what's your best guess about where all of this is going uh you know obviously who knows in the eye in the next like fears but like in the next year or two well we know some things we we know that the model we use it you're using today, Fable or Gini Petit Flash 6, will be the worst model you ever use.
1:10:19And whatever you think is cool about it right now is going to be boring and backwards in 6+. And that is so exciting. And I think I watched the way our young engineers use it, and it's very different depending the way I'm using it. And it's sort of a fun time where you can learn from your young team members that they're using AI very differently. I think the business of dashboarding and the business of AI doesn't, the damages to the SaaS is, I think, unrepairable. You could ask your AI, build me a tool like Salesforce. 30 seconds later, You have a working tool. It's just unbelievable. And all the things that were difficult because they cut across your inside organizational silos.
1:11:18One of the things that's really hard, if you want to know for your top performing people, when was the last time they got a stock option refresh and how much holding power is left? So how much uninvested stock they have left at today's price? It was like five systems. You're in work days. You're in your stock. You're in your car. None of them can. That's what a CEO wants. How much holding power is for my top guys? And I used to have little tools I wrote for this. And I've got a little app that I had. And Spark right. Well, what a story. It's just incredible to hear all of this from you. What a journey and what an exciting future.
1:12:13So thank you very much. I learned a lot. And this was terrific. Thank you, Andrew. Thank you for having me on your show. I really appreciate it. Hi, it's Matt Turk again. Thanks for listening to this episode of the Mad Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests. Thanks and see you at the next episode.
From the publisher
AI is no longer just a race to train smarter models. As AI moves into production, the bottleneck is increasingly inference: how fast models can generate tokens, use tools, reason, verify, and act. In this episode of the MAD Podcast, Matt Turck sits down with Andrew Feldman, co-founder and CEO of Cerebras, to explain why fast inference may define the next era of AI.
Cerebras is known for building a chip the size of a silicon wafer. But this conversation is not just about one company or one chip. It is a deep dive into the AI infrastructure stack: GPUs, ASICs, memory, HBM, SRAM, data centers, power, TSMC, AWS, OpenAI, agents, reasoning models, and why speed changes what AI products can become. Andrew explains why “tokens per second per user” matters, why generating a single word can require moving the equivalent of 100 HD movies through memory, why agents amplify latency, why GPUs struggle with certain inference workloads, and why fast AI may eventually reshape SaaS itself.
This is a reference conversation on fast inference, AI chips, and the next compute bottleneck.
(00:00) Cold open & Intro
(01:31) Why speed became the AI bottleneck
(02:32) Tokens per second per user, explained
(03:16) AI’s broadband moment and the Netflix analogy
(04:35) The AI chip landscape: GPUs, TPUs, Trainium, ASICs
(06:36) What is an ASIC?
(08:08) Nvidia, Groq, and the fast inference war
(09:16) OpenAI, Broadcom, and specialized silicon
(12:10) China, power, and sovereign AI infrastructure
(15:05) Is the AI infrastructure boom a bubble?
(18:56) The hidden bottlenecks: HBM, CoWoS, and 3nm
(22:57) Why agents are creating CPU demand
(25:36) Andrew Feldman’s path from SeaMicro to Cerebras
(26:13) Why Cerebras bet on AI in 2016
(31:14) SRAM vs. HBM: why inference is a memory problem
(33:19) What wafer-scale computing actually means
(34:28) The deep-tech “Everest” problem
(36:07) The moment the first Cerebras system worked
(36:49) Ringing the bell and surviving deep tech
(39:08) How a giant chip handles failure
(41:22) Why GPUs struggle with decode
(42:17) Prefill vs. decode explained
(44:01) The “100 HD movies” problem in AI inference
(45:04) How fast inference changes RL and training
(48:08) Reasoning models and why they cost more compute
(50:08) Verification, guardrails, and small models checking big models
(52:37) Multimodal AI and the path to video
(53:51) Cerebras’ business model: hardware, cloud, and API
(55:14) OpenAI’s 750MW inference deal
(55:36) Why data centers are measured in megawatts
(58:01) AWS Trainium + Cerebras decode
(59:29) Fast tokens as a cloud product
(01:00:52) Is CUDA still a moat?
(01:03:53) How TSMC helped Cerebras build the giant chip
(01:07:41) Why nobody cared in 2020
(01:08:15) Why chip supply chains are hard to diversify
(01:09:54) Why today’s AI models will be the worst you ever use
(01:10:38) What fast AI could do to SaaS
