In short
Cerebras CEO Andrew Feldman explains how Cerebras built the world’s largest AI chip (wafer-scale silicon) and why it can deliver much faster AI inference at lower power/cost than GPUs, plus the company’s go-to-market, supply constraints, and geopolitics/export controls.
Guest backgrounds
Andrew Feldman is founder and CEO of Cerebras. The episode is hosted by Jill Weisenthal and Tracy Alloway (Odd Lots).
Key claims
- Cerebras wafer-scale chips are ~58x larger than prior chips and enable use of much faster (but less dense) memory, reducing inference latency.
- Cerebras says it is 15x faster than the fastest GPU on some workloads, and 50–100x (up to 1,000x) faster on certain problems.
- The company’s architecture targets both training and inference, but current demand is heavily inference-driven.
- Cerebras avoids key bottlenecks: it doesn’t use scarce HBM memory, doesn’t use NVIDIA-style COOS process, and uses TSMC 5nm rather than 3nm.
- CUDA is “not important” for inference; two of three leading frontier models reportedly run without CUDA.
- Data centers are the near-term constraint; AWS deployment is planned via Bedrock.
Notable examples
- Deals: OpenAI contract “north of $20B” (signed December) and AWS deployment (signed March).
- Customer: G42 (UAE) runs Cerebras systems for training and inference across its cloud ecosystem (e.g., universities, Adnoc; deployments in Santa Clara, Minneapolis, Dallas, and soon Toronto).
- Open-source example: Cerebras-powered “KimiK2” (1T parameter) on Cerebras cloud, described as 10–15x faster than others.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOAI and Its Market Impact
2:20 to 3:40
Discussion on the increasing influence of AI on markets and personal thoughts.
“But you do spend a lot of time inputting instructions, pressing the button and seeing what comes out.”
The Chip Landscape in AI
3:40 to 4:54
Exploration of the different types of chips used in AI and Cerebras' innovations.
“Seriously, though, there's a good reason to think about AI more and more, which is that a huge chunk of not just the market, but the real economy is now revolving around AI, right?”
Introduction to Cerebras' Giant Chip
4:54 to 6:00
Andrew Feldman discusses the significance of Cerebras' large wafers.
“Well, one fact about them is their chips are just enormous about the size of the dinner plate.”
Technical Advantages of Giant Chips
6:00 to 8:31
Understanding why larger chips can enhance processing efficiency.
“What is the technical reason why this actually makes sense as a superior form of architecture for at least some aspect of AI?”
Overcoming Challenges in Chip Development
8:31 to 10:26
Andrew explains the hurdles faced in creating large chips over the decades.
“Yeah, it was an ambitious undertaking, that's for sure.”
Cerebras' Position in the AI Market
10:26 to 12:18
Insights into Cerebras' market valuation and the demand for inference chips.
“Obviously, you've hit this remarkable milestone.”
The Importance of Speed in AI
12:18 to 18:16
Andrew discusses the importance of speed for AI applications and business.
“Maybe we'll get more to the theoretical training market a little later.”
Cerebras' Market Share and Supply Chain
19:24 to 22:22
Explore Cerebras' strategies to tackle supply chain constraints.
“Let's say we stipulate that this is all true and everyone wants the fastest and everyone's like, you know what?”
Open Source vs Closed Source AI Models
22:22 to 26:00
Understand the financial implications of open source vs closed source AI.
“I mean, since we're talking physical constraints, I guess I should ask you, we did an episode about helium recently, a helium shortage given the situation in the Strait of Hormuz.”
NVIDIA's Moat and Competitive Landscape
26:00 to 28:00
Examine the relevance of NVIDIA's software ecosystem in today's market.
“But if you had to make a bet, like in 20 years, is the dominant aim AI model going to be a cheap open source thing or a more expensive, incrementally better closed source model?”
Show all 22 chapters
The Shift from CUDA: Market Trends in AI Training
28:00 to 29:25
Explore the declining dominance of CUDA in AI model training and inference.
“And it has no role whatsoever in inference.”
Financialization of Compute: Insights and Innovations
29:25 to 32:07
Discussion on the evolving market for compute capacity and financial instruments.
“You know, since we're talking about the economics of inference and all this stuff, I would love to get your take.”
G42: A Key Player in AI and Compute
32:07 to 34:20
Learn about G42's role and contributions to AI using Cerebras chips.
“That as this market for data centers and compute matures, there'll be people making bets on either side and financial instruments will be created to do it.”
Corporate Demand for Inference Services
37:01 to 40:02
Examine the shift in corporate needs for AI inference providers.
“I mean, look, Anthropic every couple of days announces some new thing.”
Challenges in U.S. Semiconductor Manufacturing
40:02 to 42:00
Discuss the hurdles in establishing semiconductor fabs in the U.S.
“You know, you were talking about fabs in Taiwan earlier, and I'm now regretting not going on a fab tour when I was in Taipei, but it just didn't cross my mind at that time.”
Export Controls and Their Impact on Semiconductor Business
42:00 to 45:04
Learn about the complexities and strategic importance of export controls in the semiconductor industry.
“You want them working at today's cutting edge, but tomorrow's and next year's and in 10 years, cutting edge as well.”
Navigating the IPO Process and National Security Concerns
45:04 to 46:24
Explore the challenges and considerations involved in taking a semiconductor company public amid national security scrutiny.
“Just quickly on the sort of like current business stuff, you mentioned the deal with AWS.”
Reflections on Wealth and Success in Tech
46:24 to 48:24
Understand the personal perspective on wealth accumulation and success as a tech CEO.
“especially because of the relationship with G42 about CFIUS and some of the national security concerns, and maybe that was an issue with the IPO.”
Balancing Innovation and Financial Performance
48:24 to 51:16
Discover how a tech CEO balances financial performance pressures with long-term innovation.
“I think what you have to do is you have to love the work, you have to love the people, and you have to think every day about how to make your team rich.”
The Philosophy of Building Technology
51:16 to 52:36
Delve into the philosophy behind developing technology and the patience required in the hardware industry.
“And so I think I would love it less if you could do it in a week.”
Exploring the Future of Silicon and Artificial Life
52:36 to 53:00
Engage in a thoughtful discussion about the role of silicon in technology and its connection to life.
“Fascinating conversation right in the sweet spot of what we're interested.”
Discussion on AI Token Spending
56:00 to 56:56
Learn about the implications of companies' spending on AI tokens and future market dynamics.
“And I think one of the things that is going to happen And there have been all these stories about sort of like token shock, like how much companies are spending on tokens.”
Transcript
Automatic transcript. May contain errors.0:00Tracy Alloway:OddLots is brought to you by VanEck. For years, investors basically forgot about real assets, energy, gold, and infrastructure. But look what's driving markets now. Central banks loading up on gold, massive capex cycles, currencies doing weird things. These assets are at the center of it. RACS, the VanEck Real Asset ETF, is an actively managed one-stop shop for real assets, spanning gold, commodities, natural resource equities, and more. Go to vanegck.com slash R-A-A-X pod to learn more. Fun disclosures later in this episode. Being a small business owner isn't just a career, it's a calling. Chase for Business knows how much heart and effort go into building something of your own.
0:40Manage all your business finances, from banking to payments to credit cards, all in one place with Chase's digital tools. Plus, access online resources designed to help your business thrive. Learn more at chase.com slash business. Chase for business. Make more of what's yours. The Chase mobile app is available for select mobile devices. Message and data rates may apply. JPMorgan Chase Bank N.A. Member FDIC. Copyright 2026. JPMorgan Chase and Company. When you're running a business, the best days are the ones where priorities stay on track. For midsize and large companies, risk can affect multiple parts of the organization at once, from property and liability to cyber and regulatory challenges.
1:21At that level, managing risk becomes an ongoing discipline. At the Hartford, the focus is on helping businesses manage risk before it turns into something more disruptive. And when losses do happen, that work is paired with insurance coverage shaped by years of underwriting, risk engineering, and claims experience. Learn more at thehartford.com slash risk mitigation. Policies provided by Hartford Fire Insurance Company and its property and casualty affiliates, Hartford, Connecticut. Bloomberg Audio Studios. Podcasts. Radio. News.
2:06Tracy Alloway:Hello and welcome to another episode of the Odd Lots podcast. I'm Jill Weisenthal. And I'm Tracy Alloway. Tracy, I have to say, unfortunately, I don't have AI psychosis. I'm certain of that. Debatable. I'm pretty sure I don't have AI psychosis. I do have to say, unfortunately, like the amount of time now where it's like it feels like AI related questions and there are many of them are sort of like swallowing up the other thoughts that I have in my head of whether it's questions about which models best and why and what are the economics of inference and how much training is pre-training versus post-training for each model.
2:50Tracy Alloway:all like it's just sort of like this blob that's growing that's taking up more and more of my thoughts what is your definition of ai psychosis because one would argue that maybe thinking about ai literally all the time would be a form of psychosis well let's just say like i'm not the type who thinks that like i don't like think that the ai is a friend for one thing i'm not in love with the ai models i don't think that in collaboration with chat gpt that i'm stumbling on unified theory of physics and things like that. So like... But you do spend a lot of time inputting instructions, pressing the button and seeing what comes out.
3:29Tracy Alloway:And seeing what comes out. I'm just saying, I think I'm aware that I'm talking to a machine and that we're not establishing any great breakthroughs of which we are collaborators and partners and friends. Recognizing you have a problem is the first step towards healing, Joe. Seriously, though, there's a good reason to think about AI more and more, which is that a huge chunk of not just the market, but the real economy is now revolving around AI, right? Totally. So anyway, again, within the AI conversation, there are a lot of subcategories. One of the subcategories happens to be another of Lot's favorite topic, which is chips.
4:04Tracy Alloway:Of course, chips are used in multiple different ways. Chips are used in different parts of the AI supply chain, different types of chips have different roles. And so we have to learn more. We have to learn more. And I have to say, I'm particularly interested in the company we're about to speak to partly because the two things i know about them are number one they just had a huge ipo yep right raising something like 5.5 billion dollars at kind of insane multiple i can't even do a price to earnings multiple because they're not profitable yet but i think just on a sales basis it was like 67 times forward earnings which is pretty juicy pretty hot and the second thing I know about the company is they make giant wafers.
4:47Tracy Alloway:Yes. Which is just a fun image to have in your head. That's right. So if you were thinking it's like, okay, there is a hot entrant in this space. What is their differentiator? Well, one fact about them is their chips are just enormous about the size of the dinner plate. One might think you're reading an Onion article, but in fact, it's real. And apparently it actually has some real technical advantages. And it's different to what everyone else is doing. So everyone else is, I guess, doing this sort of like modular networking thing where you get together a bunch of chips and you connect them together.
5:20And that's how you get more compute, more memory, more power, basically. But this company has done something different in the form of the giant wafer.
5:28Tracy Alloway:The giant wafer. And if you figure that to get maximum performance, you sort of want to lessen the distance between things, then put it all on one wafer. Anyway, we're going to learn a lot more. I'm very excited to say about giant wafers and more. I'm very excited to say we do have the founder and CEO of Cerebros on the podcast, Andrew Feldman, truly the perfect guest. So, Andrew, thank you so much for coming on the podcast on the week of your IPO. Well, thank you so much for having me. What a pleasure. Absolutely. Why don't you just start us off the big giant chip? They're apparently real. They're as big as a dinner plate.
6:03Tracy Alloway:What is the technical reason why this actually makes sense as a superior form of architecture for at least some aspect of AI? I think larger chips process more information in less time. Okay. And that produces faster results. And everybody had gone to bigger chips. NVIDIA had moved from 400 square millimeters to 800 square millimeters over the course of five or six years for this exact reason. and in the compute industry wafer scale which is building a chip this for those who are by the way for those who are just listening andrew is now holding up the chip and yes it looks it actually looks bigger than its interplate to be honest but is that is a big that's a big chip that's a big chip beautiful 50 is think of it's 58 times larger than any other chip that had ever been wow And what it did was it allowed us to use a different type of memory.
7:01Okay. A type of memory that, at the beginning, there are two types of memory. There's memory that can store a lot, but it's really slow. Okay. And there's memory that can't store very much per square millimeter, but it's blisteringly fast. Okay. And historically, all graphics processing units used this memory that could store a lot, but was really slow. And that's the reason they do inference so slowly. So if you're using Claude right now or you're using anything but ChatGPT, what you'll frequently feel is you'll enter your prompt and you'll wait for an answer, right? And that's because the memory is slow and they have to move a ton of information from memory to compute.
7:45Now, by going to wafer scale, we could use this fast memory. Now, we couldn't make that memory store more information per square millimeter, but we could add square millimeters. And so by building this big chip, we were able to stuff it to the gills with this fast memory. And that's why we're 15 times faster than the fastest GPU. That's why on some problems, we're 50, 100, even 1 ,000 times faster than graphics processing units. Wait, can you explain how you actually managed to do this? Because I know there have been previous attempts to do wafer scale, and I seem to remember there was even like an early attempt in the 1980s or something to do it.
8:28How were you able to pull this off? Yeah, it was an ambitious undertaking, that's for sure. Every previous effort in the 75-year history of our industry had failed, including Gene Amdahl, who's sort of on the Mount Rushmore of compute in our industry. He failed sort of spectacularly in the mid-80s at a company called Trilogy. Not only that, but after we succeeded, people who had visited us, who'd been in our labs, tried to copy us, and they also failed. And so what we were able to do is solve a set of really fundamental problems. And those problems cut across a wide swath of technology. They cut across lithography, so we had to collaborate closely with TSMC, and they turned out to be a great partner.
9:15We had to make inventions in material and packaging. That's how you put a processor, how you put a piece of silicon on a motherboard, deliver power and IO to it. We had to make inventions in power delivery. When you build a giant chip, you're going to deliver way more power to it than if you do a chip the size of a postage stamp. We had to invent ways to cool it. We had to write new types of software that ran on it. All of these had never been done before. And it was a decade-long process. It took us five years and about$500 million to deliver the first one. And it's been an extraordinary run since.
9:56In December, we signed a deal with OpenAI north of$20 billion, one of the largest contracts ever signed in Silicon Valley. And then in March, we signed a deal with AWS, where they would deploy our systems in their data centers, in their AWS data centers. And so it's just been an extraordinary run, but it took a long time. It took extraordinary engineering. And there were certainly long periods of time when it wasn't clear we were going to make this work.
10:26Tracy Alloway:Obviously, you've hit this remarkable milestone. You have, in fact, IPO'd and so forth. and right now markets valuing your company at$64 billion early days of the IPO. Just for the listener to understand, the chips, are they solely an inference as opposed to training? When we think about AI, I think about, okay, there's training, training the model, and then answer giving, that's the inference. Are the chips just for inference? So a couple of things. I think you framed it exactly right. Training is how we make AI. And inference is how we use AI. And so what happened was that in sort of 2025, in the first part of 2025, the models we made were smart enough to be useful.
11:10And there was an explosion of use. And we use AI by doing inference. So there was this sort of tidal wave of demand on inference. And that has continued in 2026, and we think it will continue for years and years to come. And so that's what had happened. In 2015, when we began thinking about the company, we knew that AI was on the horizon and it would eat a huge amount of compute. And we made sort of two fundamental bets. We bet that it would need dedicated silicon. And graphics had needed dedicated silicon. And that's how you got the graphics processing unit. Mobile compute had needed dedicated compute.
11:55That's where you got ARM processors. We made that bet. And we made a bet that modifying the GPU architecture wouldn't be right. You needed to start with a clean sheet of paper. And so what we started with was a new vision. And that vision could do training. And it could do inference. And it was orders of magnitude faster at both. But right now, what we're seeing is such an explosion in demand for inference that a lot of the business this minute is inference, even though we're just as fast at the same amount faster than GPUs on training.
12:31Tracy Alloway:That's interesting. Maybe we'll get more to the theoretical training market a little later. Just real quick on inference. Ben Thompson, who writes a newsletter about tech, he wrote a piece in which he distinguishes between answer inference and agentic inference. So answer inferences like, you know, format my resume or whatever, or write me an essay on X or Y or answer some questions. And then agentic inference is like, OK, here's this thing that's going to go around. Do you distinguish and do services for you, not producing visual answers? Do you distinguish between those two? Is that a real divide in your view?
13:08Tracy Alloway:And can your chips do both? Our chips can do both. I think it is a divide. OK. I think speed matters equally in both. I think if you are engaged with the AI, if you're writing code, which is agentic, if you're writing code or you're doing work, nobody wants to wait. I mean, we could just turn the question around and say, well, how big is the market for slow search? Zero. How big is the market for dial-up internet? Zero. Why is that? Because nobody wants to wait. Right? So if you're engaged with the AI, speed is of the essence. But if the AI is doing agentic work and your competitor gets three times, five times, ten times as much work done in 20 minutes than you do, you're going to get smoked.
13:57And so this notion somehow that Ben proposed that speed isn't very important in agentic flows is dead wrong. That speed is important in all aspects of productive work. and that your ability to get more done in less time is a fundamental advantage that accrues over time. If while your competitor is doing one unit of work, you can do three, and in the next time they do one unit of work, you do six. This adds up over time, and you beat them in any line of work. And so speed, which is sort of our specialty, is important across the board. What do giant wafers and speed in general actually mean for, I guess, the economics of tokens?
14:47Because one way I think about it, I have this sort of vision in my head, like, okay, if I'm out shopping for toothpaste, I know I need toothpaste every once in a while and I go into like a CVS store, I get one thing of toothpaste and then maybe a week later I get some more toothpaste. Or I could go to Costco and buy a giant thing of toothpaste and take it home probably at a cheaper cost. And that's sort of how I think of the giant wafers. Maybe it's a bad analogy. But what does speed actually mean for the cost of tokens? Well, I think there are a couple observations. I think people have chosen so far to price speed a little higher.
15:27For example, Anthropic offered a premium service in which they offered tokens twice as fast and charged six times as much. And they sold it out. And they couldn't meet the demand. Now, just to give you an idea, we're 15 times faster than they're twice as fast. And so people value speed because it allows them to do more work. And they value their time. And when you can do more work in less time, you are making people more productive. That's why people have chosen to price them at a premium. They don't cost more to make. In fact, the GPU architecture is an extremely good architecture and extremely efficient at building very slow tokens.
16:13And if you don't mind slow, the cost per token on a GPU is extremely low. But the GPU has a characteristic that as you try and go faster, the cost and the power used per token increased. Sort of like as you go faster in your car, your miles per gallon decreased, right? So what happens is as you try and get fast enough to be useful, fast enough to be interesting, fast enough to keep users' intelligence focused on this product, they become extremely expensive and extremely power hungry. And so the question is, is not just what people are paying for a token, what people are choosing to price them at, but what they actually cost to make.
17:02And GPUs make very slow tokens very cheaply. And they're unbelievably expensive at fast tokens. We make fast tokens vastly less expensive than GPUs. And we use a tiny fraction of the power.
17:30Tracy Alloway:Data centers need electricity. AI needs copper. Reshoring needs steel. And gold's run may tell you something about how the world is repricing money and debt. All of those point back to real assets. The RACS ETF is an actively managed one-stop real asset shop from gold to commodities to natural resource equities, adjusting as conditions change. Visit VanEck.com slash RAAX pod to learn more. An investor should consider the investment objective risks, charges, and expenses of the fund before investing. To obtain a prospectus and summary prospectus, which contains this and other information, visit VanEck.com.
18:08Tracy Alloway:Please read the prospectus and summary prospectus carefully before investing. Rax is distributed by VanEck Securities Corporation Distributor. The thing about AI for business, it may not automatically fit the way your business works. At IBM, we've seen this firsthand. But by embedding AI across HR, IT, and procurement processes, we've reduced costs by millions, slash repetitive tasks, and freed thousands of hours for strategic work. Now we're helping companies get smarter by putting AI where it actually pays off, deep in the work that moves the business. Let's create smarter business, IBM. Everyone has been there.
18:48Your team's feedback is scattered across emails, chats, and sticky notes. It's a mess. But PDF Spaces and Adobe Acrobat gives you one collaborative workspace to streamline every file and comment. So, if you need six departments to finally agree on a proposal, do that with Acrobat. Need to turn a mountain of feedback into one plan of action? Do that with Acrobat. Want to stop searching for files and finally get everyone on the same page? Do that, do that, do that with Acrobat. Learn more at adobe.com slash do that with Acrobat.
19:24Tracy Alloway:Let's say we stipulate that this is all true and everyone wants the fastest and everyone's like, you know what? This is the solution that the Cerebras technology, one big chip. This is really where it's at. How much of like your market share for the inference market when you look out next year, the year after, etc.? How much is your market share going to be dictated by your ability to get capacity at TSMC fabs? How much is that a gating mechanism for growth? You know, TSMC is a huge part of the supply chain. Yeah. But we have some real advantages. There are three areas right now that are limiting vendors and building AI computes.
20:08Number one is HBM memory. is this memory we described earlier that can store a lot, but it's really slow. That's made by three companies approximately, Samsung, Hynix, and Micron. And it's under unbelievable supply pressure. It's extremely difficult to get. There are very long lead times. It's unbelievably expensive right now. We don't use it. The second part that's limiting is a process inside of TSMC called COOS. And this is the process that NVIDIA and other GPUs use. We don't use it. The third thing is that at TSMC, the factory that is under most pressure is their three nanometer factory. We don't use it.
20:53We use five nanometer. So we have managed to avoid some of the most binding supply constraints. Now, TSMC still has to give us a meaningful allocation, and they've been an extraordinary partner from the get-go. And they are the greatest manufacturing company on earth by far. A fab is sort of a modern pyramid. It's an unbelievable thing. And I highly recommend you or any of your listeners, if you get a chance to go to Taipei, go and see them. They are just extraordinary. Can you do fab tours? You can, actually. Yeah, you can't do that. You can go and they have a museum of innovation. And it is an extraordinary thing.
21:33They are the sort of the national champion of Taiwan. But I think today, TSMC has given us as many wafers as we've needed. Business today is constrained by data centers. And that's the grand irony, right? You invent technology that has been unbuildable, never been invented for 75 years in the history of compute. You write software. that is extraordinary. You built a product that is vastly faster than the incumbent. And what are we all constrained by? Buildings. Data centers right now are everybody's constraint in the entire industry. Powered buildings, so real estate. It is an amazing thing right now.
22:14And that is true sort of across the board. And that will not change for the next 15 or 18 months for sure. I mean, since we're talking physical constraints, I guess I should ask you, we did an episode about helium recently, a helium shortage given the situation in the Strait of Hormuz. And one of the things that helium is used for is lithography on semiconductor chips. Has that affected you at all? Or is that something that you're monitoring? We monitor, but there's not a lot we can do. And there's plenty of stuff to worry about that we can't affect. We obviously are in communication every day with TSMC.
22:53We're in communication with our entire supply chain every single day. And we stay abreast of the various issues. But it has had no impact on us. And we put that in the bucket of things that our manufacturing partners worry about also, and that we can't help.
23:12Tracy Alloway:You know, so in addition to manufacturing these chips, you actually, I didn't realize this, you have your own cloud. We do. And, Or you have your own cloud services. We do. I have a bunch of questions about that. But you have your own cloud services through which a user can actually get access to various open source models and so forth. It looks a little bit sort of visually, it looks a lot like the open router interface, roughly the same environment, except it's all like the open source. Something I'm curious about, and maybe you can speak to this, you know, in traditional software, open source, one nice thing about open source is you don't have to pay for it.
23:54Tracy Alloway:So it's free. It's a little bit different when we're talking about there's no really such thing as like free AI software, because even if it's like free, you still have to like pay for the depreciation of the chips and you have to pay for the electricity to run them. So there's no real such thing as like free open source AI software. But what I am curious about in your experience as a cloud vendor, are the open source models cheaper on a per unit of intelligence basis? If we had some way of saying levelized cost of intelligence, which I don't know if the industry has yet, are open source models cheaper per IQ point, however we want to measure intelligence?
24:35Yes, by a lot.
24:36Tracy Alloway:Really? Yeah. I think in the closed source world, you're paying a lot for that extra little bit of intelligence, right? The open source models, there are no open source models that are as good as the closed source models. Okay. Think of it as 3%, 4%, 5 % different. Okay. Something in that range. It could be a little more, it could be a little less. But the cost to you using them, right? You can jump up right now and run KimiK2. It's a 1 trillion parameter model. It's an open source model on Cerebris where 10 or 15 times faster than others. And what you're paying for is the cost of our power and some cost of the compute that took to calculate it.
25:19What you're not paying for was the cost to train it. Right. And that's a battle that is underway in the market. You have OpenAI with their coding software. You have Anthropic with their coding software. And you've got companies like Cursor and Cognition that are using open source. We power open AI and we power Cognition. You have a battle underway between closed source and open source. And I think that the winners of that battle is yet to be determined. What is clear is that the closed source is strictly better by a little bit, by how much varies, and it's more expensive. Yeah, I think we've talked about this before, but like I've heard of a lot of big companies in the US who have been like very quietly shifting from some of the closed source models to the open source models like the Chinese ones like Kimmy.
Read the full transcript
26:12Is that what it's called? Kimmy and Quinn. I'm sorry to press you on this point. But if you had to make a bet, like in 20 years, is the dominant aim AI model going to be a cheap open source thing or a more expensive, incrementally better closed source model? I don't think there's going to be one. There's not one SaaS software. There's some big dogs. There's Salesforce. There's some other giant players. And there are lots of other specialists. I can't think of many markets where we've settled onto one player. If you look at the semiconductor market, you've got x86 where you've got two major players in AMD and Intel.
26:51And then you've got a whole adjacent market owned by ARM and the companies that build ARM parts. And then you've got custom silicon around that. I think that's the way you're going to have this. We're going to have, you know, OpenAI is going to continue to do extraordinary things. They will be competitors to them, and they will be open source. I don't think any of those go away. Since we're on the topic of software, one of the things you often hear when talking about, you know, new chip entrance going up against NVIDIA is this idea that, well, you know, like NVIDIA chips, they're great and all, but the real moat of NVIDIA's business is CUDA, right?
27:30Yeah. software stack that goes with it. What's your take on that? Is that a realistic concern for someone who's trying to go up against a company as big and I guess as embedded in the software system as NVIDIA currently is? NVIDIA is probably the greatest company in the first part of this century, right? Jensen's one of the great CEOs of our era, along with Hawk Tan at Broadcom and maybe Lisa at AMD, just extraordinary. And CUDA was really important in the creating of the AI landscape. But it's not important now. And it has no role whatsoever in inference. If you want to move from running a model on GPUs today to running it on us, we can move it in 10 keystrokes.
28:18Just move, point to our API. So that's the first part. The second part is that a year ago, every major frontier lab model had been built on a CUDA foundation. And today, two of three haven't. So they lost 70 % market share. There are three leading frontier models, Gemini, Quad, and GPT. Gemini, built by Google, on TPUs, trained on TPUs, served on TPUs, no CUDA. Anthropics models, trained on Tranium, no CUDA, served on TPUs, on Tranium, and on GPUs. And OpenAI's GPT, trained on GPUs in the CUDA environment. So two of the three leading models today use no CUDA. That's a hemorrhaging a share. And so I think what was true three or five years ago in which CUDA had a dominant position with central has shrunk significantly and not important at all in inference and shrinking in its role in training.
29:31Tracy Alloway:You know, since we're talking about the economics of inference and all this stuff, I would love to get your take. One of the things that literally in the last couple of weeks, there's been this flurry of announcements of these attempts to financialize the market for compute. And so it's like, oh, you're going to buy some capacity, the H100 benchmark, et cetera, and people want maybe theoretically hedging it. But I'm not entirely convinced. It still seems to me like I it's not like maybe. But on the other hand, like an inference provider can lock in a very long term relationship bilaterally with a data center and so forth and no need for like these spot hedging markets.
30:18Tracy Alloway:Do you think the market is going to evolve in such a way that there will be significant demand for financial instruments that allow inference providers to hedge their price exposure? I don't know. I'm not a financial engineer is the first thing. OK. But we can look a little bit at history. OK. The guys at Corweave were enormously innovative in how to fund some of their massive deployments. They were some of the first to use a debt instrument that had a backstop with the GPU. And this enabled them to really leap out and sort of have first mover advantage in the neocloud space. And that was an innovation in financial engineering and extremely creative.
31:04Others followed, and now there's a big and active debt market in funding the building and the fit out of data centers. When you have a market that is that big and that active, you have people who want to make bets on either side. And I think over time, those bets normalize and regularize, and you can wrap them up, and you can make it easy to make the bet. When sort of CO2 was one of the first to loan money against GPUs for CoreWeave, this was really innovative. And not only does Corweave get credit for the creating of the instrument, but so does the other side of the deal for doing it and making a successful, innovative bet.
31:48And as sort of more and more people jumped in and these could be regularized, they could be more easily priced. And then once it's regularized and you have a market, then derivatives of that market are easy to make historically. And that's sort of the way I see this unfolding. That as this market for data centers and compute matures, there'll be people making bets on either side and financial instruments will be created to do it. Whether it's a good idea or not, I have no opinion at this time. Since we brought up finance, I was looking through the IPO filing and looking at some of the actual numbers in there.
32:29And I know you have the OpenAI deal now, but a huge chunk of your revenue comes from this company called G42 in Abu Dhabi. And I think they're both like your biggest customer and also a major investor. What does G42 actually do with all these chips? Sure. Last year, they were a really important chunk of our business, a lot of it. They're a minority investor. They are the national champion, the national AI champion of the UAE. And they build a cloud that is used across the UAE's ecosystem. So it's used by leading universities there. It's used by leading companies there, companies like Adnock. They're a leading oil company.
33:19It's used by G42's nine operating companies. The deployments to date have been in the US. We have massive data centers that run equipment for G42 here in Santa Clara, but also in Minneapolis, in Dallas, Texas, soon in Toronto. And so they're doing training and they're doing inference. The training they're doing, they have pioneered some of the leading English-Arabic models. They've done genomic work. They are doing serving of models. And they're operating as a cloud, particularly for the UAE ecosystem, but also for global companies.
34:20Sending a file is easy. Making sure your clients understand the file is the hard part. But with PDF spaces in Adobe Acrobat, you can give your clients the full picture with custom intros, audio summaries, and a helpful AI assistant to your docs. So if you want to stop the endless follow-ups, do that with Acrobat. Need to make your docs crystal clear? Do that with Acrobat. Want to make sure your clients get everything they need to hear? Do that with Acrobat. Learn more at adobe.com slash do that with Acrobat. Support for the show comes from Public. Public is an investing platform that offers access to stocks, options, bonds, and crypto.
34:59And they've also integrated AI with tools that can assist investors in building customized portfolios. One of these tools is called Generated Assets. It allows you to turn your ideas into investable indexes. So let's say you're interested in something specific like biotech companies with high R &D spend, small cap stocks with improving operating margins, or the S &P 500 minus high debt companies. Chances are there isn't an ETF that fits your exact criteria. But on public, you just type in a prompt and their AI screens thousands of stocks and build a one-of-a-kind index. You can even backtest it against the S &P 500.
35:35Then you can invest in a few clicks. Go to public.com slash market and earn an uncapped 1 % bonus when you transfer your portfolio. That's public.com slash market. Ad paid for by Public Holdings. Brokered services by Public Investing, member FINRA SIPC. Advisory services by Public Advisors, SEC Registered Advisor. Crypto services by ZeroHash. Sample prompts are for illustrative purposes only, not investment advice. All investing involves risk of loss. See complete disclosures at public.com slash disclosures. Running a small business takes everything you've got, but with Chase for Business, you're not alone.
36:09They bring together local support and a broad range of resources to more than 7 million customers. With a deep understanding of your day-to-day needs, they provide products and guidance built to help you thrive. Right now, earn$500 when you open a new Chase Business Complete Checking account for new business checking customers with qualifying activities. Offer expires June 18, 2026. Chase Business Complete Checking has the flexible tools you need to accept payments, make deposits, and manage your finances with confidence. Learn more at chase.com slash podcast biz offer. Chase, make more of what's yours.
36:43These may apply to Chase Business Complete Checking accounts. The$500 offer is available for new business checking accounts with qualifying activities through June 18, 2026. Eligibility and qualification requirements must be met. Additional restrictions may apply. Please speak with a business banker for more information. JPMorgan Chase Bank, N.A., member FDIC.
37:04Tracy Alloway:Do you think that over time, corporate users and perhaps individual users, but corporate users will want inference served from a company that's separate from the model maker such that they can be certain that they are not revealing and thus training the company that might replace them? I mean, look, Anthropic every couple of days announces some new thing. Oh, we have a new markdown file that could do this for tax or that could do this for whatever. And then a bunch of companies fall. Like, are companies that use AI increasingly going to want to use data centers and inference providers that aren't the model themselves?
37:48Well, first, I think there is a type of professional, a type of job that is most directly under threat from AI. Okay. And they're almost always white collar. And they required you to have expertise over a body of knowledge, right? That's what an accountant is, right? You have expertise over a body of knowledge of rulings, of previous examples of tax case law, etc. That's exactly what AI is good at right now. Exactly. So lawyers, accountants, these professionals who have stood between the ordinary person who doesn't know anything about IRS tax rules and the tax rules. That is under threat. And that is something that it will be very easy for companies like OpenAI and Anthropic to chew through.
38:49There are other areas, like say, drug design, genetics, genomics, where companies like GlaxoSmithKline have remarkable and unique data sets. This is true for one of our large customers, Mayo Clinic. It's true for GlaxoSmithKline and other of our pharma customers. They have unique data. Yeah. And they will be able to find insight in that data, and they will be able to get value from that data. And they will certainly not want to share that data with the foundation model makers unless they are guaranteed that it will not sort of make the general model smarter. Right. And these are companies that have spent 20 or 30 years spending tens of billions of dollars a year gathering data, right?
39:44Patient care records or test results for drug design. They're going to mine the insight in this work, and they're going to find extraordinary things. And those are much more protected because the insight's in the data, and they have the data. You know, you were talking about fabs in Taiwan earlier, and I'm now regretting not going on a fab tour when I was in Taipei, but it just didn't cross my mind at that time. Next time. Yeah, hopefully. And there have been various efforts under the CHIPS Act and some other industrial policies to try to build more chip making capacity in the U.S. In your view, what's the big, I guess, impediment to actually doing it?
40:30Yeah. A, is it happening? And then B, why does it seem so difficult to actually make happen? Right. The first thing is difficult because it's a difficult problem. They're hard. They cost 30 or 40 billion dollars and take five or six years to build. So that amount of money in that amount of time cuts across administrations, right? And that's a problem with the politics in the US, is it's hard to make policy that's durable across administrations and across time. It's the first thing. The second thing is these are remarkably complicated buildings. And we have a sort of a hodgepodge, a sort of strange latticework of local, regional building codes that a fab maker has to negotiate.
41:24Third is we're trying. TSMC has dedicated tens of billions of dollars to their fabs in Arizona and have committed hundreds of billions more. Samsung has dedicated tens of billions of dollars and committed hundreds of billions more to their fabs in Texas. But they take a long time. And we have to remain committed to building not just the fab, but the surrounding ecosystem, them, not just for three or five years, but for 20 years or 25 years, because you want not just one fab, but you want a whole trajectory of fabs. You want them working at today's cutting edge, but tomorrow's and next year's and in 10 years, cutting edge as well.
42:09And those are things that have proven really challenging in the US. And I think we need it. They're strategic assets. And I think we need to find ways to collaborate with those that have the expertise and to find ways to build policy that is durable over a length of time that can build a vibrant ecosystem in the fab and the associated elements. So the other big political economy theme, I guess, when it comes to semiconductors is this idea that they are, in fact, a strategically important technology. And so the U.S. should place some limitations on their use abroad. And so we've seen things like export controls, export restrictions.
42:53You're an actual chip company. And so I'm very curious at an operating level what your experience of these kind of export controls has actually been. Like, how much time does that take up for you? And then also, given that one of your biggest customers is an international firm in Abu Dhabi, how important is the trajectory of those export controls to your future business? I think three or four years ago, I would have said not important at all. I think today they're really important. In the last administration, I got to know the leadership in the Department of Commerce and in the BIS Division of Commerce, which oversees the licensing.
43:34I think this is an extraordinarily difficult job. And we saw really hardworking, smart people doing a job that is very, very difficult. I got to know the people in this administration and I found the same. Every single one of them is earning a tiny fraction of what they could earn in the private sector and is doing this because they believe that this is an important mission. The problem is, is that there are differing views about the right way to do this. And there are differing views on, you know, the right way to achieve the goal, which is to not give your most precious technology to your industrial enemy.
44:14And I think we can agree that today, in today's environment, China is an industrial enemy. Good, well-meaning people can disagree on whether the right strategy is to limit them from gaining access. Others argue, as those at NVIDIA have argued, is that the right strategy is to give them access and to keep them we're working on our product, on US made, on US sort of designed product. I come down on the other side of that argument. I understand they're good arguments in both directions. I think limiting the distribution, the diffusion of our most precious technologies makes sense. And I think we have to do it thoughtfully and we have to recognize that means some markets will be foreclosed to us.
45:05And I'm okay with that.
45:07Tracy Alloway:Just quickly on the sort of like current business stuff, you mentioned the deal with AWS. How does that work? Can customers right now, like, can customers of AWS pay them to have inference served specifically on one of your chips? Not yet, but soon. Okay. They will be, it will be served in Bedrock, which is their AI as a, service offering. And they will, yes, be able to go down the click down menu and get super fast inference, which will be delivered via a combination of what's called a disaggregated solution, which is using some tranium for some of the inference work and using the Cerebris technology in our systems called the CS3 for other parts of the work.
45:53Tracy Alloway:And presumably someone who scrolls down and selects that, they would pay some premium for that ultra fast inference? I think they will pay a premium. We will see this as entirely as Amazon wishes to price it. This is their product. So you IPO'd this week. It's May 2026. This is not the first time that you've tried to or look towards going to the IPO market. There were headlines going back to 2024 about wanting to try for the IPO market. And then there were headlines last year, especially because of the relationship with G42 about CFIUS and some of the national security concerns, and maybe that was an issue with the IPO.
46:34Tracy Alloway:But also last September, you got one of your, looks like G round, G round, one of the participants in the G round investor was 1789 Capital, which is of course the firm that's associated with Donald Trump Jr., which is a
46:55Tracy Alloway:I'm a cynic. So I wonder if the participation, if Donald Trump Jr.'s investment in your company made it easier to get the green light from these national security concerns to do an IPO. I wish it were that easy. No, it had no role at all. We resolved all CFIUS issues in March of 2025. I believe that was before we took money from 1789. Okay. Moreover, I wouldn't ask. that's not who I am, and that's not the way we roll. So we took money because they are a thoughtful venture firm. And we don't believe that there's only one point of political view. There are lots of political views. They all have some merit.
47:39They all have some weaknesses. And so we have right-leaning political, some investors. We have left-leaning. The fact that this firm had some right-leaning investors. We were looking only at their ability to help us build an extraordinary company. And we have asked, and we have never asked, nor will we ever ask for political access or anything of the kind. What's it like to become a billionaire in a single day? This is something I assume will never happen to me, so I might as well ask you. No, I think the honest truth is it was a big nothing for me. I had some wealth before and have some wealth after.
48:22I think this is a very difficult way to make money, being a tech CEO. I think what you have to do is you have to love the work, you have to love the people, and you have to think every day about how to make your team rich. And far more important then sort of some change in my wealth was we made more than 800 millionaires. And that's something I'm proud of every minute of every day. And at my last company, we made 100 millionaires. And at this company, through our IPO, we made more than 800. And that's something that you wake up feeling good about yourself every single day. That was going to be my last question, but actually you just reminded me in that answer, you know, the idea that getting here, I said you became a billionaire in a day, but obviously this was the outcome of years and years and years of work.
49:17And if we think about technological hardware, one of the things most people associate it with is really long lead times and really big research and development budgets. Now that you're a public company, how do you sort of balance that quarter to quarter financial performance pressure with the idea that you still need to be investing in CapEx, in new ways of designing chips, new improvements to the existing ones? First, we think the opportunity for innovation based on our way for scale engine, the best work is still ahead of us. Number one, we see an opportunity for extraordinary innovation in the years ahead to make leaps every bit as big and often bigger than what we made by building the largest chip on earth.
50:09When you love building hardware, the fact that it takes time is part of the deal, right? That what we do can't be done in a week or a month or a year. And that's what you sign up for. And that's true in every profession. You sign up for the good and the challenging. And you have to sort of make peace with that if you're a person that wants to dive in and sort of begin iterating right away and fail quickly and code up something and look at it and throw it out in the market and see if it wins. Godspeed, that's great. And that's not for me. In our business, we measure twice before we cut once. And you have to put that in your soul and you have to like it.
51:00You have to like that mistakes in our business are really expensive. And you have to like the fact that you breathe life into a chunk of silicon. And you get it to do things that nobody else has ever been able to make a chunk of silicon do. And if that's for you, then this process that takes time and money, you'll love that too. And so I think I would love it less if you could do it in a week. And I think the people that I love to work with, they feel the same way. And they like being engineers, not because it's a path to money. They like being engineers because they like building things. And they like building hard things.
51:40And I like working with them for exactly that reason.
51:43Tracy Alloway:Yeah, you mentioned breathing life into a chunk of silicon. My dad, who's a physicist, always likes to point out how carbon and silicon are right next to each other on the periodic table. They are. And they're sort of like, here are the two things that we have closest to life, and they're literally touching each other. Maybe there's something deep in that. I think that's a really thoughtful thing your father said. Thank you. And I think that's really cool. And nobody pointed that out to me, though. We've stared at periodic tables for a long time. But I think to the extent we can make artificial life, we need silicon.
52:17Yeah, and they're right next to each other. Right. Carbon is the heart of all other life. And artificial life will be founded. At least the intelligent part will be founded on silicon. Right below silicon is germanium.
52:29Tracy Alloway:Maybe the next—I don't know. What does that mean, Joe? Ask your dad. Yeah, let's keep an eye on germanium next. Andrew, thank you so much for coming on Odd Lots. Fascinating conversation right in the sweet spot of what we're interested. Really appreciate you taking your time. Thank you guys for having me. And I really appreciate it. Look forward to seeing you again soon.
52:59Tracy Alloway:That was really fun. I'm super interested in this topic. And it does feel to me like the economics of inference in particular and the market for inference, inference capacity, speed, like it's still day one. You know what I'm saying? I just like looking at the giant. It's so cool. It really does seem like an onion thing, doesn't it? It's like company solves inference with a giant chip. By building the biggest chip in the world. But it is interesting. We did that episode, of course, with Ray Wang from semi-analysis and talking about the role like memory as being this really important part of this sort of cutting edge chipsets.
53:37Tracy Alloway:And it's interesting to think it's like, OK, well, here is a bottleneck that doesn't run into that they don't have. And the idea that at least as he described it, they're not fighting to get the smallest nanometer chips. And so maybe that gives them a little bit of breathing room on capacity there, too. Yeah. I mean, I do imagine there are some downsides to having giant chips, you know, just as there are upsides that Andrew laid out. The other thing I was wondering, I know he made the case for the reason speed is very important, but like I can also imagine a world where maybe it's not that important.
54:15you know like i think at some point like the incremental speed factor just starts to become less important when weighed against like the incremental cost of generating that extra speed
54:28Tracy Alloway:i think it really this is like one of those things where it probably really depends what you're what you're using it for right so it's like if you're like you know what i'm really curious uh why pterodactyls aren't actually dinosaurs can you explain it to me then it's like i don't care about that. Like that fraction of a second is not that important. I would wait five minutes for the chat bot to tell you you're wrong, Joe. You just, buddy, you just don't really care that much. But if you're doing some sort of like a gent decoding thing or whatever, et cetera, then like, yeah, that definitely adds up.
55:00Tracy Alloway:And I will say like, as you use it more, like, it's just like everything else, the treadmill of expectations. Here's some tasks that you can do in 30 seconds, which maybe several years ago would have taken you 30 minutes. And you get a patient in that 30 seconds and you want it in 10 seconds. And that's just like that competition to shave down seconds. I think it's always going to be there. So no one ever gets satisfied with this is my point. It always eventually becomes like it feels like waiting. But to me, this feels like this is the crux of the A.I. valuation argument, which is like how much of a premium are we going to place on a model that may be a closed source model that is maybe slightly better than an open source model?
55:43How much premium are we going to place on compute that is slightly faster than this other type of compute or like other use of compute? Like that to me, it's an unanswered question. And Andrew was pretty upfront about closed versus open source. But I think on the speed question too, like we're going to find out.
56:02Tracy Alloway:We're going to find out. And I think one of the things that is going to happen And there have been all these stories about sort of like token shock, like how much companies are spending on tokens. My guess is one of the things that will happen at some point is there's going to be a lot more discussion about why are we using this ultra premium model when we could have done this? Like there is a lot of just like throw it at the AI, rack up those bills, et cetera. And at some point there's going to be this like, OK, what really needs to be served fast? What really needs to be served on the most premium closed source models?
56:38Tracy Alloway:And companies are probably going to get a lot more skilled at allocating from, you know, different forms of inference depending on the need. Yeah, I think that's exactly it. And at that point, like we could well see some of the dynamics in the market start to change in terms of valuation. Shall we leave it there? Let's leave it there. This has been another episode of the All Thoughts podcast. I'm Tracy Alloway. You can follow me at Tracy Alloway. And I'm Joe Weisenthal. you can follow me at The Stalwart. Follow our producers, Carmen Rodriguez at CarmenArmandDashelBennett at DashBot, Kale Brooks at KaleBrooks, and Kevin Lozano at KevinLloydLozano.
57:15Tracy Alloway:And for more OddLots content, go to Bloomberg.com slash OddLots. We have a daily newsletter and all of our episodes. And you can chat about all these topics 24-7 in our Discord, discord.gg slash OddLots. And if you enjoy OddLots, if you like it when we talk about giant wafers, then please leave us a positive review on your favorite podcast platform. And remember, if you're a Bloomberg subscriber, you can listen to all of our episodes absolutely ad-free. All you need to do is find the Bloomberg channel on Apple Podcasts and follow the instructions there. Thanks for listening.
58:19We'll see you next time.
58:33Tracy Alloway:Giving you a single layer of control, a single standard of trust. So whether an AI agent supports a single user or your entire enterprise, with Okta, you'll turn risk into opportunity. Secure every agent. Secure any agent. Okta secures AI. A business gift should do more than check a box. It should reflect your brand and show someone they're appreciated, recognized, and truly seen. 4imprint offers thousands of high-quality, customizable products, like premium apparel, drinkware, tech, and more, making it easy to create a gift that feels meaningful and on-brand. And with 4imprint's expert support and their 360-degree guarantee, you can be 4imprint certain your order will arrive exactly as intended.
59:16Tracy Alloway:Explore gifting with confidence at 4imprint.com. 4imprint. 4certain. When you're running a business, the best days are the ones where priorities stay on track. For midsize and large companies, risk can affect multiple parts of the organization at once, from property and liability to cyber and regulatory challenges. At that level, managing risk becomes an ongoing discipline. At the Hartford, the focus is on helping businesses manage risk before it turns into something more disruptive. And when losses do happen, that work is paired with insurance coverage shaped by years of underwriting, risk engineering, and claims experience.
59:52Learn more at thehartford.com slash risk mitigation. Policies provided by Hartford Fire Insurance Company and its property and casualty affiliates, Hartford, Connecticut.
From the publisher
Size is the name of the game for the AI chipmaker Cerebras: Their chips are truly massive, about the size of a dinner plate. According to Andrew Feldman, CEO and founder of Cerebras, that is about 58 times larger than the average chip. That sheer size enables blazing fast inference for AI queries. Feldman joins us on the week of his company's IPO to talk about his core product and how it fits into the AI boom. We discuss the history of the GPU, competition between open-and closed-source models, the company's relationship with with TSMC, and more.
Read more:
Nvidia Tells Skeptical Investors That AI Is Ready to Go Mainstream
Trump Set to Sign AI Cybersecurity Directive as Soon as Thursday
Only Bloomberg - Business News, Stock Markets, Finance, Breaking & World News subscribers can get the Odd Lots newsletter in their inbox each week, plus unlimited access to the site and app. Subscribe at bloomberg.com/subscriptions/oddlots
Subscribe to the Odd Lots Newsletter
Join the conversation: discord.gg/oddlots
See omnystudio.com/listener for privacy information.
