In short
Odd Lots Podcast Summary: Episode - "Two Veteran Chip Builders Have a Plan to Take On Nvidia"
Podcast Overview Hosts: Joe Weisenthal and Tracy Alloway Description: Explore intriguing topics in finance, markets, and economics.
Episode Summary In this episode, Weisenthal and Alloway discuss the current landscape of artificial intelligence (AI) chips, focusing on Nvidia's dominance and the emergence of a new startup, MatX, aimed at designing chips specifically for large language models (LLMs). Co-founders Reiner Pope and Mike Gunter share their insights on chip design, the challenges of competing with established players, and the future of AI hardware.
Key Themes and Discussions
- Nvidia's Dominance in AI Chips
- Financial Success: Nvidia is a significant player, generating substantial revenue from AI applications.
- Market Position: The hosts note that Nvidia's GPUs are synonymous with AI and are widely used beyond this scope, including gaming and cryptocurrency mining.
- Barriers to Entry in Chip Design
- Moats: Discussion on the concept of 'moats' in semiconductor manufacturing, highlighting the high costs of R&D, the need for specialized expertise, and the network effects that protect established companies like Nvidia.
- Challenges for Startups: MatX aims to create a chip exclusively for LLMs, indicating a specific market focus that could disrupt Nvidia's dominance.
- Chip Development Process
- Design Lifecycle: Pope and Gunter outline the stages of chip design, from initial architecture to manufacturing, emphasizing the complexity and time investment (3-5 years).
- Verification and Testing: A large verification team is crucial to ensure functionality before manufacturing, underlining the precision required in chip design.
- Innovations and Competitive Strategies
- LLM-Specific Chips: MatX's strategy focuses on creating chips that optimize performance for LLMs, potentially allowing for significant cost savings on compute resources.
- Nvidia's Constraints: The discussion touches on why Nvidia might struggle to pivot to LLM-specific chips due to existing commitments to its CUDA software ecosystem.
- Market Demand and Customer Needs
- Customer Expectations: Insights into what potential customers (like OpenAI and other AI labs) are looking for in terms of performance metrics, particularly flops per dollar.
- Future of AI Models: The evolving landscape of AI, particularly regarding LLMs, raises questions about the scalability of models and the corresponding hardware needed.
- Future Outlook and Industry Predictions
- Market Dynamics: Both guests express confidence in the potential for new entrants to provide competitive alternatives to Nvidia, given the increasing demand for AI capabilities.
- AGI Predictions: The episode concludes with discussions on the timelines for artificial general intelligence (AGI), with a skeptical view on the immediate feasibility of achieving AGI.
Key Takeaways
- Emerging Competition: The emergence of MatX indicates a potential shift in the AI hardware landscape, seeking to challenge Nvidia’s stronghold.
- Complex Development: Chip design is a lengthy, complex process requiring significant investment and expertise, which poses challenges for startups.
- Customer-Centric Innovation: Understanding customer requirements, particularly around cost efficiency and performance, is crucial for new chip developers to succeed.
- Industry Evolution: As AI continues to develop, the demand for specialized chips will grow, presenting opportunities for new competitors and innovations in the semiconductor space.
Conclusion The episode provides an in-depth look at the challenges and opportunities in the AI chip market, highlighting how new companies like MatX are attempting to carve out a niche in a landscape dominated by Nvidia. The discussion sheds light on the intricate processes involved in chip design and the strategic considerations necessary for success in this competitive field.
For further insights, listeners can explore past episodes related to AI and semiconductor industries on Bloomberg’s Odd Lots podcast.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00You're being sold an AI future where you're obsolete or irrelevant. That vision is wrong. At Palantir, they're building AI that helps workers and unlocks their full potential. American workers are our nation's greatest strength. AI shouldn't eliminate them. It should elevate them. Palantir is here to tell their stories. From factories to hospitals, AI is freeing people from drudgery, letting them do what humans do best. Create. Solve. Build. Palantir, making Americans irreplaceable.
1:01like small business, Hiscox Small Business Insurance.
1:07Bloomberg Audio Studios. Podcasts, radio, news.
1:24Hello and welcome to another episode of the Odd Lots podcast. I'm Joe Weisenthal. And I'm Tracy Alloway. Tracy, here's something I know about AI. I don't know much, but here's something I do know. How to log into ChatGPT? No, I'm good at that. I'm good at logging into ChatGPT and Claude, and I'm reasonably good at asking questions. Now, here's actually something about the business of AI that I know. Okay. I know that NVIDIA is making a ton of money, and the stock has gone to the moon, and that other companies would like a slice of that pie. Yes. Yes. That's a good thing to know. It's like a basic, simple thing, which is that when people think about AI chips, there's literally one company that comes to mind.
2:10I know others are involved. AMD has stuff. Intel obviously wants to play others. But there is obviously that one gigantic pile of cash that's flowing to this one company. I don't know if it's still, but at one point it was the biggest company in the world is pulled back a little bit. Well, I would say two things. One, other companies would like a piece of that pie. And B, companies that are in the business of building AI models would like to find a way to get cheaper, more efficient, less energy intensive chips so that they don't have to always pay the NVIDIA tax. Do you want to know what I know about AI and semiconductors?
2:49Let's go for it. Okay, here's the one thing that I know, which is that whenever you have this conversation about NVIDIA, the one word that always comes up is moat. Oh, yes, moat. So, like, you're either talking about, like, medieval castles or you're talking about semiconductor manufacturing. That's when you hear the word moat because over and over again, people will say it is expensive to make the chips. You need a lot of money for research and development and to set up the fabs. And you need a lot of first-person expertise in building them. And then there's also the network effect. So a company like NVIDIA has this huge moat around its business.
3:26The question, of course, is whether or not, getting back to the medieval castle analogy, it is unassailable. That's right. In fact, semiconductors seems to be moat after moat after moat because there's ASML's moat, and then there's Taiwan Semiconductor's moat, and then there's NVIDIA's moat. And so, yes, it's like there's a series of moats. And if someone could overcome these moats or find a way to build a bridge over one of these moats and enter this proverbial castle, that would be very lucrative. We know that many are trying to enter these moats, but it's incredibly costly. and capital intensive and difficult.
4:05And there are just not many people who know how to do any of this stuff. And so the question of whether these moats can be overcome. But again, there are many businesses that would love to see more robust competition in the space so that their payment is not a tax. You know, one thing I don't know, and I don't think we've ever done an episode purely on this, but I don't really understand the different designs of chips. So I I know that some chips, specifically NVIDIAs, are supposed to be better at AI. They're better at running lots of little calculations all at the same time. And I know there's basic chips that go into your refrigerator or your car or whatever.
4:46But I don't really know the difference between what a chip that was designed specifically to run a large language model would look like compared to a standard basic chip. I don't know anything about chip design. I just sort of imagine someone using some CAD software, etching little lines in the thing and drawing some sort of circuitry or placing the transistors. Actually, a chip design game would be really fun, now that I think about it, where you could just draw little things on the square. Okay, anyway. Well, we are going to learn about how chip design works. We are going to learn about what makes a chip particularly good for the task of training and running inference on these AI models.
5:30And I have to say, I really do believe we have the two perfect guests because they are both veterans in this space and they are both active in the attempt to bridge some of these moats and enter the space and bring competition to the industry. We are going to be speaking with Rainer Pope, co-founder and CEO of MedX, as well as Mike Gunter, co-founder and CTO of Maddox. It's a new company that's trying to build chips specifically for the purpose of large language models. Both of them have a lot of experience in the space. We're going to get our hands dirty, so to speak, and understand how you build the hardware for all this stuff and what makes it win and whether it's even a winnable game.
6:13Reiner and Mike, thank you so much for coming on Outlaws. Thanks. Happy to be here. Pleasure to be here. So why don't you tell us, what does a chip designer do? I know I have this completely cartoonish view in my head that cannot possibly be right of someone on a big screen using some CAD software to sort of, you know, figure out what's going to be etched in that wafer of silicon. What is the job of chip design? So maybe this is best told by what is the story of chip development from the beginning of a project to the end of it. So there's a range of different ways this can go, but there's a lot of things that are in common.
6:48So generally, a chip design team is at the low end, maybe 30 people up to many, many thousands of people at the high end. And the project typically runs for somewhere in the range of three to five years from conception to actually shipping to customers. And so over that time, what we see in the lifecycle is we tend to start with a small team of architects. If you think of designing a house, the team of architects, they're the people who decide what rooms go in here, how many bedrooms, how many bathrooms, what are the flows between them, how do people walk through the corridors, and so on. The core screen design of the chip.
7:24In the chip itself, that is what kinds of components at the high level we have. And then after that initial exploration, this moves then over to the microarchitects. These are the people who are designing the individual rooms. What are the components that go in the individual rooms? So at that point, everything we've done so far is a design stage thing. This is done in documents, spreadsheets, and it's a verbal and human communication form. But beyond that, that's when it starts to actually touch the computer in a more meaningful sense. And so the microarchitects will hand over to the logic designers.
7:56They are the people who are actually writing code. So even though you think of chips as being this very physical thing where there's wires and gates and everything, the way we transmit this information to the computer is actually writing code. We write Verilog that expresses the design of the chip. So that's what the logic designers are doing. That's an extended period of time building out all of the different matrix multiplies, memories, circuitry that connects to the outside world, and so on. And then the output of all of them is this Verilog piece of software code that gets then compiled by a computer down to a set of gates, which are logic gates and or gates and so on.
8:33And then wires that connect them together. That's the net list. This file, then there's a few more stages still coming here. This file gets handed off to physical designers who, again, work with CAD tools to convert this kind of logical description. I was right. So someone is using CAD tools. Absolutely. There's a CAD tool, but it's only a part of the job. Okay. So the physical designers are converting the sort of logical description into a physical placement. So where do each of these gates go? Now there's 200 billion logic gates on a chip. So a human is not going to be placing all of those manually.
9:04So there's a huge amount of software assistance here. But what the human is doing is providing oversight through this process and saying, I've done this a ton of times before. This placement kind of looks wrong. It doesn't match my heuristics. And so I can probably do a better job here. So that's the physical designers. And the output of their work is actually, eventually you get a polygons. So basically an image saying, here is the thing that is going to get etched onto a piece of silicon. So that file is ultimately a huge, really big image in some form, a bunch of polygons on it. It gets handed over to a manufacturing company such as TSMC.
9:41They spend maybe four or five months initially creating a mask set. So those are the templates or the stencils that will be used to stamp out many, many copies of the chip. And then stamps out many copies of the chip, you get a chip back. This is typically about two or three years after you started the project. you get chips back. And now you have a bring up team who puts this chip into a whole board and connects it to power and electricity and starts testing it. And then after another six to 12 months or maybe even more, eventually you actually can hand this over to customers. There's maybe just one or two other things which are not in that flow, but very essential to call out too, are because of this whole process taking so long, especially the manufacturing, we also have very large teams of verification people.
10:27So these are the people who before we actually send it to manufacturing and pay$20 to$30 million of manufacturing. We have a substantial team doing a lot of testing. And this is software-based testing, so writing tests in the same way a software engineer might, to make sure that the functionality actually works as intended. To underline the comparison to ordinary software, which Reiner touched on, we're writing code, but it's on super hard mode. So if you have a software that's deployed via website, you can fix a bug in 10 minutes at basically zero cost. Whereas in our case, the reason that we have a large team of people doing verification, making sure that what we've done is correct, is that it's potentially four months and$30 million for every mistake that you let through.
11:15Likewise, there is software, but it's a relatively small fraction of software that's very performance critical, where you want the code to run as fast as possible. But in some sense, every line of code that you write in hardware has an impact on the overall performance of the product because every line of code ends up getting embodied in silicon and every line of code affects the eventual performance. So it's kind of coding, but on hard mode. So I intuitively understand the importance of getting the software right. But why does placement on like the actual chip or wafer? Why does that matter? Are you trying to make it more efficient?
11:53Are you trying to reduce the rise time? Or why does it matter where the little bits and bobs are placed, to use the scientific term? Yeah, you're right that reducing the rise time is a massive issue. And, you know, fundamentally, the issue is that chips, you know, at a very abstract level are composed of, or at a somewhat concrete level, really, are composed of transistors and wires. And the placement has a dramatic effect on the length of the wires, which has a dramatic effect on both the performance of the chip and how much you can fit. But in terms of the impact that this has on the quality of chip that you produce, wires have over time not been shrinking in the same way that transistors have.
12:40And so getting the wiring right, which usually means getting the placement right, has become more and more important over time.
13:02Silicon Valley is selling you a future where you're obsolete, or worse, identical. At Palantir, they're witnessing something different and revolutionary. From re-industrializing the nation's defense base, to shipyard workers building faster, and frontline workers boosting productivity, AI is transforming work across the nation. AI is not replacing American workers or flattening them into conformity. It's unleashing what makes each one irreplaceable, their judgment, their craft, their creativity. When American workers become more powerfully themselves, they own the future. Palantir, making Americans irreplaceable.
13:43Support for the show comes from Public.com. You're thoughtful about where your money goes. You've got your core holdings, some recurring crypto buys, maybe even a few strategic option plays on the side. The point is, you're engaged with your investments, and Public gets that. That's why they built an investing platform for those who take it seriously. On Public, you can put together a multi-asset portfolio for the long haul. Stocks, bonds, options, crypto, it's all there. Plus an industry-leading 3.6 % APY, high-yield cash account. Switch to the platform built for those who take investing seriously.
14:16Go to public.com slash market and earn an uncapped 1 % bonus when you transfer your portfolio. That's public.com slash market. Paid for by Public Investing. All investing involves the risk of loss, including loss of principal. Brokered services for U.S.-listed registered securities, options, and bonds in a self-directed account are offered by Public Investing, Inc., member FINRA, and SIPC. Crypto trading provided by XeroHash. Complete disclosures available at public.com slash disclosures. Can chips be beautiful? I know code can be elegant, and some people will say certain code is beautiful, but have you ever looked at a semiconductor and been like, oh, wow, that's really nicely put together?
14:56For me, I mean, I think absolutely, yes. This is like why I work in this space is I just really like geeking out on the design of things. But to me, what beautiful for a chip means is that it kind of does exactly what it was designed to do. And no more, no less. I mean, obviously, less would be a bit of a disappointment. But often if it does more, do you think, well, maybe I designed it for slightly the wrong purpose or something like that? I think this is a good seg into getting into your business specifically. So we all know that so much of this AI is powered by these NVIDIA GPUs, but NVIDIA GPUs have been used for a long time for many things that do not have anything to do with large language models or the specific AI applications that people are excited about today in 2024.
15:42So for a while they were, well, Video games is obviously the big one for decades and decades. And then there was like five minutes where people got really excited to use them for Ethereum mining. And now everyone's really excited about their use for artificial intelligence and large language models and some of these other generative AI applications that people are excited about right now. Why don't you tell us maybe the sort of idea behind Maddox, but specifically what you were both doing when you were at Alphabet or Google, which, you know, it has its own chips. I believe it has something called TPUs.
16:17What was the project at Google? Why did Google find it necessary or a good business to start building their own chips for in-house purposes? And then why did you feel the need to then leave to build what you're building now for LLM specifically? Yeah. So what Google was seeing, and this was at this point, some time back, more than a decade ago, they were seeing that the use of artificial intelligence, LLMs were not a thing at that point, was going up. And they were worried about how much money they would have to spend on traditional, it would have been GPUs at that time. And so they built a very specialized chip to do neural nets.
17:05And that chip specialized on matrix multiplication. So they put in a structure called a systolic array, which they definitely didn't invent. It has existed from the 70s. that is especially good at doing matrix multiplication. Now, after that, NVIDIA has added a similar structure into their chips. And the initial Google TPU was an inference-focused only chip. And then they have subsequently made chips that can be used for both training and inference. And I guess now's a good point. So the very last thing that I was doing at Google, I was on the TPU team, and Reiner was on the large language model team.
17:47minute. It's probably good to have him sort of tell the story from here. So, I mean, what we were seeing, and this is what we personally were seeing, but Google was seeing more generally as well, is just large language models are a thing. There was this period of time between GPT-3 and chat GPT-3 coming out. GPT-3 came out in 2020. And so people who were very plugged into the field recognized the importance of it, or at least to some extent recognized the importance of it back then. And so there was this push to, you know, everyone wanted to create their own large language model that was better than GPT-3.
18:19And so, I mean, at the time I was on the large language model team, we helped training Google Palm, and we were using thousands of TPUs for that. And one of the things we were saying is, well, look, what does it cost to deploy this in Google search? There's quite a lot of search queries. I think it's the public estimates are about 100 ,000 of them per second. If you multiply out how much each query costs, and if you want to run that on large language models, that's a lot more expensive. And then also, if I want to train a model that's 10 times bigger than my current model or 100 times bigger, suddenly these models have just moved from costing a million dollars or$100 ,000 to train to tens of millions and hundreds of millions of dollars.
18:57And so the overall goal was, can we make it cheaper by any way possible? So of course, there's algorithmic approaches. There's a lot of opportunity on the algorithm and research side, but then the other really big lever is just making better hardware. So one of the things we were looking at was trying to make Google's TPUs better for large language models. What led us actually, I mean, this is personally about Mike and me in this case, what led us to leave Google to make Maddox was we saw that there was, we believe that there is some opportunity to make chips substantially better if you're only looking to focus on large language models.
19:29And so the chips that were designed pre-GPT3 and especially pre-ChatGPT try to do a really good job on small models as well as a really good job on large models. And so what you find is that the circuitry in those chips, there's a bit of circuitry for what you need for small models, there's a bit of circuitry for what you need for large models, also for maybe embedding lookups. There's three or four different kinds of workloads, and all of them take some of the real estate in your silicon. And so if you really want to make the best use of the real estate, you should just focus on the thing you care about most and hope that there's a big market there.
20:03So the game and what we decided to do and we see some others deciding to do as well is to really try and focus on just the one workload that seems like it's going to become a$100 billion or a trillion dollar industry. I know there's always this sort of cliche when talking about tech like, oh, Google and Facebook, they can just build this and they'll destroy your little startup because they have infinite amount of money. Except that doesn't actually seem to happen in the real world as much as people on Twitter expect it to happen. But can you just sort of give a sense of maybe the business and organizational incentives for why a company like Google doesn't say, oh, this is a$100 billion market.
20:45NVIDIA is worth$3.5 trillion or$3 trillion. Let's build our own LLM-specific chips. Why doesn't that happen at these large hyperscaler companies that presumably have all the talent and money to do it? So Google's TPUs are primarily built to serve their internal customers. And Google's revenue, for the most part, comes from Google Search, and in particular from Google Search Ads. Google Search Ads is a customer of the TPUs. It's a relatively difficult thing to say that hundreds of billions of dollars of revenue that we're making, we're going to make a chip that doesn't really support that particularly well and focuses on this at this point unproven in terms of revenue market.
21:34And it's not just ads, but there are a variety of other customers. For instance, you may have noticed how Google is pretty good at identifying good photos and doing a whole variety of other things that are supported in many cases by the TPUs. I think one of the other things too that we see in all chip companies in general or companies producing chips is because producing chips is so expensive, you end up in this place where you really want to put all your resources behind one chip effort. And so just because the thinking is that there's a huge amount of return on investment in making this one thing better rather than fragmenting your efforts.
22:12Really what you'd like to do in this situation where there's a new emerging field that might be huge or might not, but it's hard to say yet, what you'd like to do is maybe spin up a second effort on the side and have like a skunk works. Yeah, that's right. That would be my idea. Just let Ryan or just let the two of you go have your own little office somewhere else. Yeah, just organizationally that it's often challenging to do. And we see this across all companies. Every chip company really has essentially only one mainstream chip product that they're iterating on and making better and better over time.
22:43To what degree is chip design driven by the customer? And what I mean by that is, So the TPUs at Google were developed to handle Google's internal workloads. But at other chip designers, to what degree will customers come and basically do a reverse inquiry and ask for a specific chip? Or what is the dialogue between customers and the big chip designers actually look like? Yeah, it's a fun interplay of, I want my provider to do a good job, but I also don't want to leak my IP too much. so you can see this how this played out in so mike was talking about through the development of the the tpus which were publicly announced in 2016 and around the same time nvidia's first gpu with the tensor core so that was the first gpu that was really focused on matrix multiplication that was the volta generation came out at about the same time and some of this actually was a result of when google had this recognition of look matrix multiplication is so important we need to make it really better.
23:44They simultaneously worked on it themselves, but also went to NVIDIA and said, we're not telling you much, but can you do better at matrix multiplication? And so that was enough for NVIDIA to go on. The first generation, they made a pretty good attempt, but if you talk to people at NVIDIA, they'll say that actually the second generation of the Tensor Core, which was in the Ampere generation, was where they really nailed it. So when it's big enough, you sometimes see these customers coming and saying what they want, but maybe they'll try and disguise what they're asking for or not giving you the absolute minimum amount of information to help a vendor make what they want without revealing too much about their IP.
24:34Support for the show comes from Public.com. You're thoughtful about where your money goes. You've got your core holdings, some recurring crypto buys, maybe even a few strategic option plays on the side. The point is you're engaged with your investments and public gets that. That's why they built an investing platform for those who take it seriously. On public, you can put together a multi-asset portfolio for the long haul. Stocks, bonds, options, crypto, it's all there. Plus an industry leading 3.6 % APY, high yield cash account. Switch to the platform built for those who take investing seriously.
25:06Go to public.com slash market and earn an uncapped 1 % bonus when you transfer your portfolio. That's public.com slash market. Paid for by Public Investing. All investing involves the risk of loss, including loss of principal. Brokered services for U.S.-listed registered securities, options, and bonds in a self-directed account are offered by Public Investing, Inc., member FINRA, and SIPC. Crypto trading provided by XeroHash. Complete disclosures available at public.com slash disclosures. When you own your own business, you own every decision. Now own the card that rewards you for it. The Chase Sapphire Reserve for Business card brings the best Sapphire Reserve benefits to business owners who expect hardworking rewards.
Read the full transcript
25:45Designed to meet the needs of business owners at scale, this painful card elevates your travel experience and offers premium benefits and value toward business services that can take your business to the next level. Sapphire Reserve for Business provides over$2 ,500 in annual value. Fuel your business and maximize rewards with 8x points on all purchases through Chase Travel, 3x points on social media and search engine advertising, annual partnership credits, and more. Make every journey more rewarding with a$300 annual travel credit and access to a network of airport lounges, whether you're looking for pre-flight productivity or time to rest and recharge.
26:23Chase Sapphire Reserve for Business. With over$2 ,500 in annual value, it's the car that gives back all you put in. Learn more at chase.com forward slash reserve business. Chase for business. Make more of what's yours. Accounts subject to credit approval. Restrictions and limitations apply. Cards are issued by JPMorgan Chase Bank and a member FDIC. Let's get to Maddox. Tell us the product that you're designing and how it fundamentally will differ from the offerings on the market, most notably from NVIDIA. Yeah. So we make chips and, in fact, racks and clusters for large language models. Okay. So when you look at NVIDIA's GPUs, you already talked about all of this.
27:07The original background in gaming, this brief moment in Ethereum, and then even within AI, they're doing small models and large models. Yeah. So what that translates to in, you can think of it as the rooms of the house or something, they have a different room for each of those different use cases. So different circuitry in the chip for all of these use cases. And the fundamental bet is that if you say, look, I don't care about that. I'm going to do a lousy job if you try and run a game on me, or I'm going to do a lousy job if you want to run a convolutional network on me. But if you give me a large model with very large matrices, I'm going to crush it.
27:42That's the bet that we're making at Maddox. So we spend as much of our silicon as we can on making this work. There's a lot of detail in making all of this work out because you need not just the matrix multiplication, but all of the memory bandwidths and communication bandwidths and the actual engineering things to make it pan out. But that's the core bit. And why can't NVIDIA do this? So, you know, NVIDIA has a lot of resources. It has that big moat, as we were discussing in the intro, and it has the GPUs that are already in production and working on new ones. But why couldn't it start designing an LLM-focused chip from scratch?
28:18Right. So you talked about NVIDIA's moat, and that moat has two components. One component is that they build the very best hardware. And I think that is the result of having a very large team that executes extremely well and making good choices about how to serve their market. They also have a tremendous software moat. And both of these moats are important to different sets of customers. So they're a tremendous software moat. They have a very broad, deep software ecosystem based on CUDA that allows it. Oh, yeah. I remember this came up in our discussion with CoreWeave. Yeah. Yeah. And so that allows customers who are not very sophisticated, who don't have gigantic engineering budgets themselves, to use those chips and use NVIDIA's chips and be efficient at that.
29:11So the thing about Emote is not only does it in some sense keep other people out, it also keeps you in. So insofar as they want to keep their software mode, their CUDA mode, they have to remain compatible with CUDA. And compatibility with that software mode, compatibility with CUDA requires certain hardware structures. So NVIDIA has lots and lots of threads. They have a very flexible memory system. These things are great for being able to flexibly address a whole bunch of different types of neural net problems. but they all cost in terms of hardware. And they're not necessarily, the choices to have those sorts of things are not necessarily the choices, in fact, not the choices that you would want to make if you were aiming specifically at an LLM.
30:01So in order to be fully competitive with a chip that's specialized for LLMs, they would have to give up all of that. And Jensen himself has said that the one non-negotiable rule in our company is that we have to be compatible with CUDA. This is interesting. So the challenge for them of spinning out something totally different is that it would be outside the family. And so it's outside the CUDA family, so to speak. And meanwhile, you already have like Pytorch and Triton waiting in the wings, I guess. So why don't you tell us a little bit more about the business of LLM chips specifically? Because there's a lot of questions like, you know, one question is you have all these people in Silicon Valley who seem motivated by the idea of like AGI, that that's the goal, that we're going to have super intelligence one day, maybe thousands IQs into hundreds of thousands one day that'll make us all seem very dumb, etc.
31:00Are you implicitly making a bet by your company that it'll be LLMs that will get there? Because as you mentioned, there are other algorithmic ideas. There are other ideas for how you might be able to expand intelligence. How much of your company's bet is the idea that the future of generative AI, or as we know it, is going to be along the LLM pathway? One of the core things, I think there's two core ingredients of the LLM pathway. Yeah. One so far is the transformer architecture, which is a model architecture and was substantially better than the things that came before. But the other one, and that actually has a much longer history, is the scaling hypothesis in general.
31:42So that's the, there's a general observation, which has been widely recognized for a decade or more that if I am, sorry, I'm training a neural net or some kind of AI model. If I want to make its quality better and make it bigger. And so what does bigger mean? Bigger means I have to spend more compute training it. Bigger means I have more neurons. Those are the loosely analogous to the processing power in a human brain, although the analogy is weak. If I make my model bigger, I get better quality. That's a simple qualitative thing to say. And that's been true for a really long time in these models.
32:18So the advantage of that, or the thing that we've seen really recently is we've seen this turned up to 11. So around the time when GPT-3 came out, so in 2020, a paper was published called Scaling Laws. And so this took this qualitative observation and made it quantitative and said, actually, we can even fit an equation to it. And so that gave people a lot more conviction to it. And this is what led to the people saying, well, if I have a better model, I can solve more problems with AI than I could before. And so every time I spend 10 times as much training on it, I unlock new use cases. And so that's what led to this craze.
32:56And the remarkable thing is that while there are these diminishing returns, I have to spend 10 times as much computing power to get some improvement. Beyond that sort of logarithmic scale, we don't see as yet any plateauing. And so it seems like there continues to be opportunity here. So the key thing is this scaling hypothesis or scaling laws in general that are causing these models to grow. And then, I mean, as a hardware provider, what you might look at is you might say, that's the thing I really want to bet on. I want to bet on the growth of models. And I mean, now it's a little more in the details, but the thing you actually have to bet on is the growth of matrix size, which is very strongly correlated with the growth of models.
33:33Just to hammer this point home, if more AI was learning from stuff like self-play or synthetic data rather than scraping the internet, would the design of the chips have to take that into account? How would the chips vary between those different learning styles? Yeah. So in general, when you're building a chip, you have to make it programmable because you're going to make this chip and you'll ship a new version every two years. But what people want to do with the chip is going to change every month or so. So it has to be programmable to some extent. So that's true for all of the chips that anyone ships.
34:10And so there's different scales of programmability and what kinds of changes you need to adapt to. So changes in kind of the way you feed it data, that's maybe on the very, very outer layers of it doesn't affect much of the core of the chip. And so those kinds of changes tend to be some of the easier changes to adapt to. The things that then become a little harder to adapt to is if I'm substantially changing my model architecture. So a small change might be maybe I change the number of layers or I reorder some of the layers in my model, or maybe I use the same ingredients, but shuffle them around in some way.
34:44A bigger change would be that say, okay, I'm actually going to throw out all of these ingredients and use a completely different set of primitives. And that's often the, that's, that last step is the one that, that would really kill you if you're betting very much on a particular set of ingredients. So an example of a potential different set of primitives that are used in other models that aren't used in LLMs are, we made mention of these embedding things that are used in recommender and ad models. So Facebook has talked about building special purpose hardware to support inference on those kind of models.
35:16Those are, they have much less emphasis, relative emphasis, particularly on matrix multiply. Another possible direction that model architecture could go that would be different and bad for a chip designed for current LLMs would be instead of having very large matrices in about 100 layers. You could have much smaller matrices, but 10 ,000 layers. And that would demand a different sort of design to be good at that kind of model. So a bet that looks good, given the modern history of neural nets, is that matrices will get larger over time. You know, you're talking about scaling laws. And so everyone talks about, okay, computation power, energy efficiency, et cetera.
36:04And I never know if they're true, But then sometimes you read these stories that like Sam Altman wants to go around the world and raise like five trillion dollars to build his own semiconductor fabs and have the entire architecture because that's what it's going to take. What about the data side? Because this is another thing people talk about, the data wall, that there's only one Internet to scrape. And then after that, what if you're not there at AGI yet? Again, I know you're solving for the hardware side, but when you think about risks going forward along the LLM pathway, what's your perspective on, well, what happens when we've just, we've ingested all the data?
36:44So there's two ways you can make a model better. One of them is by training on more data. And the other one is making a bigger model. And these two effects work in a really complementary way. So you can think of it as like having a bigger brain and then practicing more. And so both of these are going to help to some extent. So there's a risk that we hit a data wall. In general, there's been a long history of people predicting walls in different kinds of walls in AI training and then ingenuity overcoming this. And so I would bet that there's a fairly large amount of mileage to continue here. Tracy mentioned self-training and generating new data.
37:23That's the vibe in the industry is that this is a promising direction for sure. But even if you don't bet on that, there's mileage and it's less attractive mileage, but there is mileage in making the models bigger. So I believe, and I think this is shared by many people inside us in the industry as well, is that there's at least a few more orders of magnitude available here before we run out of easy engineering knobs to turn. But of course, one of the limiting factors here is just the dollars you spend. So you have some amount of budget that I'm willing to spend. And I mean, maybe Sam Altman can raise$5 trillion.
37:59I don't think necessarily everyone else can raise that amount of money to train a model. And so if you've got a fixed amount of dollars that you want to spend and you want to train the best model, you want to make the best use of the multiplies. You want to make the best use of the dollars you spend. And so that means fundamentally what you're paying for is the flops, which flops is a floating point operation. So the number of multiplies you can do. And then every time I increase my model size or increase the amount of training data I've got, I'm spending more flops. And so flops converts into intelligence.
38:28And then if I've got a fixed budget, really what I want to maximize is my flops per dollar. I find this so fascinating because there are so many different directions that you could theoretically go in and so many decisions that need to be made. Do you go after that scale? How do you tailor the design for different methods of data input? Although, as you said earlier, maybe that's one of the easiest things to respond to. But then there are other trade-offs that you have to think about between speed and power consumption and I guess area utilization or the placement of all the bits and bobs that we were discussing earlier and cost effectiveness too.
39:09How do you balance all those elements and are there particular things that you're willing to sacrifice for others? So different people can choose different targets to go after in the market. And so one target, which you could argue NVIDIA is winning on currently, and one of the reasons that their chips, their products are so popular is, as Reiner said, just the amount of flops you can get out of a chip. And if all the chips are roughly the same to make, that translates into flops per dollar. So another target you could also go after would be the time to respond to one user. So to get the answer back, one approach is maximizing the throughput that you can have and other is minimizing the latency.
39:59So kind of the difference between a 747 flying a group of passengers across the country versus an SR-71 getting there in a couple hours, but only bringing one or two people. Let's talk about the business itself. So, you know, the old, you know, 10 years ago, someone starting a tech startup, they, you know, get three or four people in an office and then they write something up and then they have a code and it doesn't, maybe they don't even have to raise any money to do it. And they certainly don't have to depend on whether Taiwan Semiconductor has any capacity at their fab or anything like this.
40:34Walk us through the sort of nuts and bolts of what it actually takes to build a chip business from the ground up, both in terms of costs and time and what you have to rely on. You know, we've talked about some of the design element. What are the business side requirements and what will it take to actually succeed? So fortunately, and we've kind of referred to this in multiple places, there's a huge ecosystem around designing chips. So there's a portion you have to do yourself and there's a portion that you can buy. So the placement of Tracy's bits and bobs and also the testing that we've talked about, there are EDA, electronic design automation companies that build those tools.
41:20Likewise, there are companies that do just manufacturing. So TSMC and their suppliers. And then there are many other companies. So most companies don't go directly to TSMC. So very sophisticated companies like Apple or NVIDIA interface directly with them. But most other companies go through ASIC vendors. And so the most prominent companies in that space are Broadcom and Marvell. And then there are a bunch of smaller companies. A couple that are close to TSMC are Allchip and GUC. and so they'll do a lot of the work of taking your code and actually getting it placed on the chip. That's often a very good thing to outsource because the work is somewhat seasonal.
42:12You're only ready to do that placement when you're near the end of this three-year project and so you kind of don't have work unless you're a massive company for people the whole time. So while that ecosystem means that you don't have to hire a ton of, you know, a huge number of people yourself. All of those people have to get paid. Right. And so you do have to raise a fair bit of money. And another big element of actually thing that you end up spending money on is there are parts of the chip that are very special, difficult to design and take multiple iterations of taping things out and seeing if they work.
42:49So the very high speed interconnect that connects together chips is an example of that. So those are designed by yet another set of companies. And the design is difficult and fairly expensive because of the need to do multiple tape outs. And so it's very fairly expensive to buy that IP. So when you add up the cost of the IP, the cost of the ASIC vendors services, and then the mask fees that TSMC charges using ASML's mask creation software, you're talking about tens of millions of dollars to bring a state-of-the-art chip to market. The numbers are much lower for a simpler chip without the very high-speed IOs and on an older node.
43:37But for an advanced node, it's a pretty expensive process. When do you think you'll be able to bring your chips to market? Generally, we see these projects taking three to five years across most companies. We started on this seriously at the beginning of 24. So about three years from there is likely for us. Tell us about what customers, because I've heard this, you know, we're all trying to find some alternative to NVIDIA, whether it's to reduce energy costs or just reduce costs in general, or being able to even access chips at all, since not everyone can get them because there are only so many chips getting made.
44:14But when you talk to like theoretical customers, A, who do you imagine as your customers? Is it the open AIs of the world? Is it the metas of the world? Is it labs that we haven't heard of yet that could only get into this if there were sort of more focused, lower cost options? And then B, what are they asking for? What do they say like, you know what we're using in video right now, but we would really like X or Y in the ideal world? So there's a range of possible customers in the world. The way that we see, or A, way you divide them up and how we choose to do that is what is the ratio of engineering time they're putting into their work versus the amount of compute spent that they're putting in.
44:56So the ideal customer in general for a hardware vendor who's trying to make the absolute best but not necessarily easiest to use hardware is a company that is spending a lot more on their computing power than they are spending on the engineering time. Because then that makes a really good trade off of maybe I can spend a bit more engineering time to make your hardware work, but I get a big saving on my computing costs. So companies like OpenAI would be obviously a slam dunk. There's many more companies as well. So the companies that meet this criteria of spending many times more on compute than on engineering, there's actually a set of maybe 10, 15 large language model labs that are not as well known as OpenAI, but you might think CharacterAI, Cohere, and many other companies like that, and Mistral.
45:43So the common thing that we hear from those companies, all of those are spending hundreds of millions of dollars on compute, is I just want better flops per dollar. That's actually the single deciding factor. And that's primarily the reason they're deciding on today, deciding on NVIDIA's products rather than some of the other products in the market is because the flops per dollar of those products is the best you can buy. But when you give them a spec sheet and the first thing they're going to look at is just what's the most floating point operations I can run on my chip? And then you can rule out 90 % of products there on the basis of, okay, it just doesn't meet that bar.
46:19But then after that, you then go through the more detailed analysis of saying, okay, well, I've got these floating point operations, but is the rest going to work out? Do I have the memory bandwidths and the interconnect? But for sure, the number one criteria is that top line, flops. When we talk about delivering more flops per dollar, what are you aiming for? What is current benchmark flops per dollar? And then are we talking like, can it be done like 90 % cheaper? What do you think is realistic in terms of coming to market with something meaningfully better on that metric? So NVIDIA's Blackwell in their FP4 format offers 10 petaflops in their chip and that chip sells for ballpark 30 to 50 ,000 depends on many factors.
47:04That is about a factor of two to four better than a previous generation NVIDIA chip, which was the Hopper chip. So part of that factor is coming from going to lower precision, going from 8-bit precision to 4-bit precision. In general, precision is one of the best ways to improve the flops you can pack into a certain amount of silicon. And then some of it is also coming from other factors, such as cost reductions that NVIDIA has been deploying. So that's a benchmark for where NVIDIA is at now. you need to be at least a few integer multiples better than that in order to compete with the incumbent.
47:37So at least, you know, two or three times better on that metric, we would say. But then, of course, if you're designing for the future, you have to compete against the next generation after that too. And so you want to be many times better than the future chip, which isn't out yet. And so that's the thing you aim for. Is there anything else that we should sort of understand about this business that we haven't touched on that you think is important? One thing, given that this is odd lots, that I think the reason that Sam Altman is going around the world talking about trillions of dollars of spend is that he wants to move the expectations of all of the suppliers up.
48:12So as we've observed in the semiconductor shortage, if the suppliers are preparing for a certain amount of demand and demand, in the case of famously of the auto manufacturers as as a result of COVID, canceled their orders, and then they found that demand was much, much, much larger than they expected. It took a very long time to catch up. A similar thing happened with the NVIDIA's H100. So TSMC was actually perfectly capable of keeping up with demand for the chips themselves. But the chips for these AI products use a very special kind of packaging, which puts the compute chips very close to the memory chips and hence allows them to communicate very quickly, called COOS.
49:01And the capacity for COOS was limited because TSMC built with a particular expectation of demand. And when H100 became such a monster product, their COOS capacity wasn't able to keep pace with demand. So, you know, supply chain tends to be really good if you predict accurately. Yeah. And if you predict badly, you know, on the low side, then you end up with these shortages. But on the other hand, these companies, because the manufacturing companies have very high CapEx, they're fairly low to predict badly on the high side because that leads them to having spent a bunch of money on capital CapEx that they're unable to recover.
49:47Yeah, this is very interesting, this idea that in some part it's a signal. We're not slowing down. We have more and more that we want to do. So if you're anywhere along the semiconductor supply chain, don't start curbing your expectations or curbing your production because we want to build a lot more. I'm curious, one last question, I guess, for both of you. You hear a lot of people in the industry say, like, we might just be three or four years away from AGI or superintelligence, however that's defined. And then you get into a lot of these philosophical questions and ethical questions about, you know, whatever is the AI going to, what is going to be the role for humans or is it going to kill us all or whatever, you know, fear scenario you want.
50:31But the two of you, like, how do you see that question? Like, could we hit it in just a few short years where we have something that people agree is, oh, this is AGI? Is it short runway or just a couple of years away from this? Or does it feel like, no, that's still quite a few years out, if ever? I think what we have now is incredibly valuable. Sorry. Approximately zero, to be blunt. Okay. Thank you. My P, great things. I mean, I think we kind of already have great things and we've just gotten in the models of this level of quality recently, and we're learning how to use them and the quality is going up.
51:12The fact that we can get a computer to write code pretty well is fairly amazing to me. That you can ask it to tell a good joke in the style of a particular person and it can do that is also amazing. Yeah. Well, I'm glad your odds of total doom and annihilation are zero. That makes me feel a little bit better. Ryan and Mike, thank you so much for coming on OddLots. I learned a ton from that conversation. It was a pleasure.
51:53Tracy, there was obviously a ton that was really interesting in that conversation, but I particularly liked the part about incentives of large legacy incumbents about entering a totally new business. So for a company like Google, the primary purpose of their chips is going to be serving an in-house business purpose. And even with all the money that they have, and even with the engineering talent, there's still a sort of trade-off question involved of how much do we want to build chips for some other purpose, for some sort of external service? Yeah. And I also thought the point about why Sam Altman is going around talking about how, you know, how many billions he's going to spend was really interesting.
52:37And it kind of makes sense in the aftermath of the pandemic. And semiconductors, I'm sure you remember this, I think that was actually where we first learned about the bullwhip effect and this idea that very small changes in one end of the supply chain, which would be customer demand, can end up reverberating, you know, all the way through the supply chain. And so when you had carmakers start to cut back on their orders that had a much bigger and longer impact than you might have anticipated. And so it's interesting to see companies coming at it from the other end and saying like, no, we have all this money and we're going to be here for a long time.
53:15We are not slowing down. We are going to AGI. And so if you think like, oh, we're going to come out with GPT-5 and then we're going to focus on just like commercializing that and selling it to airlines to do customer support after that and just go into glide mode and take business. They want to signal that they're building more and more and more. I thought that was interesting. I thought it was interesting, the point about NVIDIA and CUDA and the idea that, okay, yes, the CUDA software ecosystem is perceived to be this moat that makes it harder for other semiconductor companies to break into the same business, but it's also constraining from an NVIDIA perspective, the idea that, OK, if they want everything to be CUDA compatible or be within the same family of software usage, then that also constrains the potential sidelines that they might get into.
54:08Right. And opens up space for competitors. But I don't know why I haven't really internalized this lesson before, because it comes up in every conversation we do on semiconductors. But I think there's still a perception, or at least maybe I still have this perception, that the moat around NVIDIA is like the actual hardware. Yes. But it's not. It's the software. It's CUDA. It seems like it's both. Well, yeah, but I think I'm starting to appreciate how much of it is CUDA is what I'm saying. It certainly seems to come up over and over again. How much the fact that this is what people use, that it's the software that makes it easy for less sophisticated customers to use the applications, it seems extremely powerful.
54:53It's also interesting to hear about the ecosystem of businesses around semiconductor design. And he mentioned Broadcom. Reiner mentioned Broadcom, which is a company that I don't think we've ever really talked about very much on the show. But if you look at that stock, I mean, it looks kind of like you're looking at a chart of NVIDIA. That has been a gigantic winner over the last few years. Back in 2020, it was a$31 stock. Now it's$146 stock. Okay, it's only a five-bagger, so maybe not quite NVIDIA returns. And this idea that— I like how NVIDIA has just skewed what's expected of every stock. It's like—it's on a different plane.
55:36And this idea that a semiconductor startup doesn't necessarily interface directly with TSMC, like that's only for the most sophisticated advance, and then there are some of these companies in the middle. I thought that was extremely interesting. You know what, Joe? I asked ChatGPT what the most beautiful semiconductor is. Yeah? It says gallium arsenide is considered beautiful for several reasons. Its crystal structure is often admired for its clarity and elegance. Oh, wow. So I guess semiconductors made of gallium arsenide for the most beautiful. So at the pure, there's beauty at the molecular level.
56:13Yeah. But actually, I thought, you know, I thought when you asked that question, it's like, oh, it's just sort of a, you know, philosophical, you know, fun, whimsical question. But this idea of like doing the minimum required or not building a bunch of extra rooms in the house that you don't really need. And as we know, I mean, it's just objectively true that even if NVIDIA chips are the best in the world for AI, they do other stuff beyond AI. and they do Ethereum mining, or they used to, and that was based on proof of work back in the old days. And of course, they're for video games. But if you really just want a computer, or if you really just want a model that can speak in English or write code, or can just think without doing video games and chip mining, then perhaps there are a bunch of rooms in the house that are totally unnecessary.
57:03Yeah. And I mean, there's efficiency costs to that. Efficiency costs. Yeah. You're trying to streamline it as much as possible. All right. Shall we leave it there? Let's leave it there. This has been another episode of the All Thoughts podcast. I'm Tracy Alloway. You can follow me at Tracy Alloway. And I'm Jill Weisenthal. You can follow me at The Stalwart. Follow our guests, Reiner Pope. He's at Reiner Pope. And Mike Gunter, he's MikeGunter underscore. Follow our producers, Kerman Rodriguez at Kermanerman, Dash O 'Bennett at Dashbot, and Kale Brooks at Kale Brooks. Thank you to our producer, Moses Andam.
57:35For more OddLots content, go to Bloomberg.com slash OddLots, where we have transcripts, a blog, and a newsletter. And you can chat about all of these topics 24-7 in the Discord, discord.gg slash OddLots. There's even a semiconductor room in there, so you can just go there and just talk about chips all day if you want. And if you enjoy OddLots, if you like it when we talk about what the most beautiful semiconductor is, then please leave us a positive review on your favorite podcast platform. And remember, if you're a Bloomberg subscriber, you can listen to all of our episodes absolutely ad-free.
58:10All you need to do is connect your Bloomberg account with Apple Podcasts. In order to do that, just find the Bloomberg channel on the platform and follow the instructions there. Thanks for listening.
58:30Thank you.
59:00from accountants and architects to photographers and yoga instructors, look to Hiscox Insurance for protection. Find flexible coverage that adapts to the needs of your small business with a fast, easy online quote at Hiscox.com. That's H-I-S-C-O-X dot com. There's no business like small business. Hiscox Small Business Insurance. In business, a thoughtful gift does more than say thank you. It recognizes achievement, builds loyalty, and shows someone they are genuinely valued. With 4imprint, you can choose from thousands of high-quality products, apparel, drinkware, tech, and more, designed to leave a lasting impression.
59:37And with expert support, dependable service, and their 360-degree guarantee, your gift will arrive exactly as intended, on time and on brand. Explore gifting with purpose at 4imprint.com. 4imprint, for certain.
From the publisher
When it comes to chips for artificial intelligence, obviously the name that automatically comes to mind is Nvidia. The company is making a fortune selling semiconductors used for hot AI applications like large language models, and stock investors have rewarded it handsomely for doing so. But of course, Nvidia's GPUs can be used for more than just AI. They're also used for video games, graphics, cryptocurrency mining and more. But a new startup called MatX is aiming to build the ultimate chip just for LLMs. Co-founders Reiner Pope and Mike Gunter spent years at Alphabet, which has its own internal semiconductor operations, and now they've stalked out on their own to create a new chip company from scratch. We talk about how they're going about their job, what it takes to actually design and build a chip, and what it will take to get customers to switch over from the industry leader.
For more episodes on this topic:
Coreweave's CSO on What It Really Takes to Build an AI Datacenter
Inside the Battle for Chips That Will Power Artificial Intelligence
Only Bloomberg.com subscribers can get the Odd Lots newsletter in their inbox each week, plus unlimited access to the site and app. Subscribe here:
See omnystudio.com/listener for privacy information.
