20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML

24 Feb 2025 · 1 h 13 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Notes: The Twenty Minute VC (20VC) - Episode with Steeve Morin | AI Infrastructure Insights

Episode Overview Host: Harry Stebbings Guest: Steeve Morin, Founder & CEO @ ZML Focus: The future of AI, the competition in chip manufacturing, and market dynamics in inference and training.

---

Key Topics Discussed

  1. The Changing Landscape of Inference (04:17)
  2. Inference is projected to dominate the AI landscape in the next five years, shifting towards a 95% inference to 5% training ratio.
  3. Companies need to focus on three critical elements to succeed: Products, Data, and Compute.
  4. Google, with its varied product ecosystem (e.g., Android, Google Docs), is positioned as a strong competitor in the AI space.
  1. AI Hardware Innovations and Challenges (09:17)
  2. Challenges include reliance on Nvidia's CUDA which ties many models to specific hardware, limiting the efficiency of alternatives like AMD.
  3. PyTorch: The leading ML framework, heavily tied to Nvidia, complicates the transition to other hardware due to compatibility issues.
  1. The Economics of AI Compute (15:38)
  2. Nvidia's high margins (around 90%) raise questions about sustainability, especially if companies seek alternatives.
  3. The cost efficiency of AMD GPUs is significant, offering 4x better efficiency for certain models but facing supply issues.
  1. Training vs. Inference: Infrastructure Needs (18:01)
  2. Training: Requires extensive infrastructure and interconnectivity; more is better for rapid iteration cycles.
  3. Inference: Demands reliability and speed with less complexity; the focus shifts towards minimizing interconnects and optimizing the production environment.
  1. Future of AI Chips and Market Dynamics (25:08)
  2. The AI hardware market is becoming more competitive, with companies like AMD and Google pushing for alternatives to Nvidia.
  3. Innovations such as compute-in-memory are on the horizon, aimed at increasing efficiency and lowering costs.
  1. Nvidia's Market Position and Competitors (34:43)
  2. Nvidia's hold on the market is challenged by its high pricing and competition from new entrants like AMD and specialized chips from Google.
  3. Dependency on Nvidia's infrastructure puts various companies at a competitive disadvantage.
  1. Zero Buy-In Strategy (39:12)
  2. Emphasizes the importance of flexibility in using multiple compute providers without heavy investment, to avoid vendor lock-in.
  1. Challenges of Incremental Gains in the Market (38:18)
  2. Many companies face high switching costs and are often hesitant to break from established solutions, even if alternatives promise better efficiency.
  1. The Role of Microsoft and Google (40:40)
  2. Microsoft’s strategy involving AMD is seen as pivotal for their AI infrastructure, indicating a potential shift in competitive dynamics.
  1. Scaling Laws and Model Efficiency (46:40)
  2. As AI models become larger, efficiency in training and inference is paramount to cope with increasing demands without unnecessary cost.
  1. Retrieval Augmented Generation (RAG) (57:08)
  2. RAG represents a shift in how models utilize data, enhancing their performance by embedding context directly rather than relying solely on pre-trained models.

---

Key Takeaways

  • Market Dynamics: The AI landscape is rapidly evolving, with inference becoming more central to business models. Companies must leverage their strengths in products, data, and compute to stay competitive.
  • Hardware Innovation: New AI chips designed for efficiency and performance are entering the market, challenging Nvidia's dominance.
  • Scalability and Cost: Understanding the true costs associated with AI infrastructure is critical, as market pressures may force companies to rethink their strategies regarding compute resources.
  • Future Directions: The next few years will see a push for more specialized architectures and the integration of AI technologies that minimize the reliance on traditional GPU models.

---

Conclusion Steeve Morin provides a comprehensive look at the intricacies of AI infrastructure, the competitive landscape of chip manufacturing, and the profound implications these developments have for the future of AI. The discussion emphasizes the need for a strategic approach to harnessing products, data, and compute effectively in an increasingly competitive environment.

For more insights, listeners can explore further episodes of The Twenty Minute VC on their [official website](https://www.20vc.com).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00The thing with Nvidia is that they spend a lot of energy making you care about stuff you shouldn't care about. And they were very successful, like who gives a shit about CUDA? OpenAI is amazing, but it's not their compute. Ultimately, if you don't own your compute, you're starting with something at your ankle. In five years, I would say 95 % inference, 5 % training. You have the products, the data, and the compute. Who has all three? Google has like Android, Google Docs, they have everything, they can sprinkle everywhere. This is the sleeping giant in my mind. This is 20VC with me Harry Stebings and I'll show you with Jonathan Ross at Grook once so well last week.

0:36But I had so many more questions on two things, the future of chips and the future of inference. So stay with dig deep on both and there's no one better join me than Steve Moron. Steve is the founder of ZML, a next generation inference engine enabling peak performance on a wide range of chips. Literally, the perfect speaker for this topic. And this was a super nerdy show. It was probably the most information Dan's episode we've done in a long time. So do slow it down, pause it, get a notebook out, but wow, there is so much gold in this one. But before we dive in today, turning your back of a napkin idea into a billion dollar start -up requires countless hours of collaboration and teamwork.

1:19It can be really difficult to build a team that's aligned on everything from values to workflow. But that's exactly what Coda was made to do. Coda is an all -in -one collaborative workspace that started as a napkin sketch. Now, just five years since launching in beta, Coda has helped 50 ,000 teams all over the world get on the same page. Now at 20VC, we've used Coda to bring structure to our content planning and episode prep, and it's made a huge difference. Instead of bouncing between different tools, we can keep everything from guest research to scheduling and notes all in one place, which saves us so much time.

1:55With Cody, you get the flexibility of docs, the structure of spreadsheets and the power of applications, all built for enterprise, and it's got the intelligence of AI which makes it even more awesome. If you're a startup team looking to increase alignment and agility, Cody can help you move from planning to execution in record time. To try it for yourself, go to coder .io -2 -0 -VC today and get six free months of the team plan for startups, that's coder .io, slash 20VC to get started for free and get six free months of the team plan. Now that your team is aligned in collaborating, let's tackle those messy expense reports.

2:32You know, those receipts that seem to multiply like rabbits in your wallet, the endless email chains asking, can you approve this? Don't even get me started on a month and planet when you realize you have to reconcile it all. Or Plio offers smart company cards, physical, virtual and vendor -specific, so teams can buy what they need while finance stays in control. Automate your expense reports, process invoices seamlessly and manage reimbursements effortlessly, all in one platform. With integrations to tools like Zero, QuickBooks and NetSuite, PLEO fits right into your workflow, saving time and giving you full visibility over every entity, payment and subscription.

3:10Join over 37 ,000 companies already using PLEO to streamline their finances, try PLEO today. It's like magic, but with fewer rabbits, find out more at pleo .io forward slash 20vc. And don't forget to revolutionize how your team works together. Rome, a company of tomorrow runs at hyperspeed with quick drop -in meetings. A company of tomorrow is globally distributed and fully digitized. A company of tomorrow instantly connects human and AI workers. A company of tomorrow is in a Rome virtual office. See a visualization of your whole company, the live presence, the drop in meetings, the AI summaries, the chat.

3:48It's an incredible view to see. Rome is a breakthrough workplace experience, loved by over 500 companies, of tomorrow, for a fraction of the cost of Zoom and Slack. Visit Rome .OR .AM for an instant demo of Rome today. Nobody knows what the future holds, but I do know this. It's going to be built in a Rome virtual office, hopefully by you. That's roym .ro .am for an instant demo. You have now arrived at your destination. Steve, I am so grateful to you for joining me today. I've wanted to make this one happen for a while, but when we were discussing who'd be best for this topic, I was like, we've got to have Steve on.

4:26So thank you for joining me, Steve. Well, thank you. I feel humbled. I appreciate it. Thank you. Dude, I want to start. Can you just give us a quick overview of ZML and specifically your role in the infrastructure strategy today and why you sit. So at the very bottom of things, ZML is an ML framework that runs any models on any hardware. We sit ultimately at the infrastructure layer. We enable anybody to run their model better, faster, more reliably, but on any compute whatsoever. Doesn't really matter. It could be Nvidia, it can be AMD, it could be TPU, and whatnot. And we do all that without compromise.

5:07That's the key point because if there's a compromise then it's not really you know, agnostic right? Can I ask you then if we think about sitting between any model and any provider there in terms of AMD and video Do you think then we will be existing in a world where people are using multiple models simultaneously and that is Concurrently running. Yes, you actually can see it. It's been happening for a while. Mothers now are not the rate of distractions at least If you look at close source model, they're not really models. They're more like back end. And there are a lot of tricks that you feel like you're talking to one model, but ultimately you're talking to a constellation and assembly of back ends that produces a response.

5:49Probably the number one, I would say, obviously, would be that if you ask a model to generate an image, then it will switch to a diffusion model, right? Not the NLM. And there's many, many more tricks, the turbo models and OpenAI do that. that there's a lot of tricks. So definitely models in the sense of getting, you know, weights and running them is something that is ultimately going away because, you know, in favor of like full blown back end, right? You feel like you're talking to a model, but ultimately you're talking to an API. The thing is that API will be running locally in your own, you know, cloud, you know, instances and so on.

6:28So we will have a world where we're switching between models and that's kind of this trickery around. Okay, perfect. We've got that at the top, then we've got that at the middle, and then you said, and then on any hardware. So will we be using multiple hardware providers at the same time, or will we be more rigid in our hardware usage? No, absolutely. You can get, or do, like, probably in order of magnitude, more efficiency, depending on the hardware you run on. That is substantial. Not a lot of people have that problem at the moment. Things are getting built as we speak, but a simple example is if you switch from NVIDIA to AMD on a 70B model, you can get four times better efficiency in terms of spin.

7:10So that is substantial. That is very much substantial. Now the problem is getting some AMD GPUs. I'm really sorry. If there is such a cost efficiency four times, why does everyone not do that? So there's a few reason. And probably the most important one is the PyTorch Cuda, I would say Duo, and that's very, very hard to break. These two are very much intertwined. Can you just explain to us what PyTorch is? Oh, yes, absolutely. PyTorch is the ML framework that people use to build actually train models, right? You can do in France with it. But by far the most successful framework for training is PyTorch.

7:51and PyTorch was very much built on top of CUDA, which is Nvidia software. Let's just say the strengths of PyTorch make it ultimately very, very bound to CUDA. Of course, it runs on AMD, it runs on Apple and so on. But there was always the tens of little details that not exactly run like you would expect, and there's work involved, but then also there's supply. So probably there's the number one thing. The second thing is there's a lot of GPUs on the market pretty much all of them are Nvidia. The reason being that if you think in layers and you say, all right, I'm going to buy, let's say GPUs and I'm going to sell them to folks to maybe not even do training, right?

8:38Just do inference. Then most likely if you look at it that way you'll end up buying Nvidia because everybody will want to run an Nvidia because nobody knows really how to do whatever and they've trained on Nvidia so they're like, I can just reuse my code and so on. So there's like this self perpetuating circle of people just buy Nvidia because they want to resell and people just use Nvidia because it's there, right? But it's by far not the not the most efficient platform and arguably even in terms of software where it's not the best software platform. So that is probably two of the most, I would, I would wait for the most important reasons.

9:17Can I, before, you know, are we chatting about Nvidia and AMD when deep seek obviously happened and the stock crash that happened? Why did Nvidia rebound, you think, in a way that AMD didn't? Because the chips are there. There's a lot of things, but in my opinion, there's going to be a need for inference. Very hard to say whether it will be worth, you know, everybody's money to do it on H100. That is a bubble that I think will blow some time. I'm kind of afraid of that, to be honest. Why do you think that's a bubble that will blow some time? Why is that not legitimate? Because it was built on the A100, I would say financial model, which was a generation zero, we do training.

10:01But when it's last generation, we do inference. And it worked beautifully, right? for a 100 then H 100 comes along and inference is it's worth five times the price and it maybe runs twice In terms of performance on inference that is on training. It's a lot better, but on inference It's like maybe twice as fast when it actually when it came out it rained at the same speed than the a 100 So there's a money gap that's going to have to you know be bridged on time, right and And the part that worries me is that I see, you know, a mortization plans on, like, you know, six, seven years, right? With the GPUs at the Coleto, and I'm like, well, I'm not sure how it's going to work because, at least, when they came out, they were five times the price, and they're just two times, you know, faster.

10:47Something has got to give. It's speed of development, trumping chip development speeds, where it's now becoming a real problem, where, as we say, models are far outpacing the speed of chip deployment. Not much, ultimately. The two things that could really very much shake the industry, the chip industry, in my opinion, is our agents and reasoning. Number one, agents. Why does that change the chip? I think this is where Nvidia can be attacked. I mean, why agents and why reasoning? The differences for agents and reasoning, you need to wait until the end of the request to get whatever it is you came for.

11:28If you don't really care about the speed at which the text outputs, which is what you want in a chat, right? You only care about how much time does it take between the beginning of my request and the end. And so that fundamentally changes the incentives from throughput bound to latency bound. And so GPUs, let's say you're running a GPUs, I'd let's say 10 ,000 tokens per second. you very much like to do it, you know, a hundred times, one hundred, right? And they can do that. But they cannot do, they cannot give you 10 ,000 tokens per second only on you, per stream, what we say. But in terms of, you know, agents are reasoning, this is exactly what you want.

12:09Because you don't want to wait like, you know, 50 seconds for whatever thinking, right? And agents, it's the same. So these two, I think, are the shot that might make Nvidia change its course, I mean, they're not idiots, right? How should agents change in videos strategy? Hard to say, because Nvidia has a very, very vertical approach. They do more of more, right? Like if you look at Blackwell, it's actually crazy what they did for Blackwell. They assembled two chips, but the surface was so big that the chips started to bend a bit, which further perpetuated the problem because it then didn't make contact with with the heat sink and so on.

12:54So they are very much, and the power envelope, they push it to a thousand watts, it requires liquid cooling and so on. So they are very much in a very vertical foot to the pedal in terms of GPU scaling, but the thing is GPUs are a good trick for AI, but they're not built for AI. It's not a specialized chip, it is a specialization of a GPU, but it is not an AI chip. Forgive me for continuing to see asking stupid questions. Why a GPU is not built for AI? And if not, what is better? So the way it worked is that you can think of a screen as a matrix. And if you have to render pixels on a screen, there's a lot of pixels, and everything has to happen in parallel, right?

13:38So that you don't waste time. Turns out, you know, matrices are very important thing in AI. So there was this cool trick in which we essentially tricked the GPU into... back that was like probably 20 years ago, we would trick the GPU into believing it was doing graphics rendering where actually we would making it do parallel work, right? It was called GPGPU at the time, right? So it was always a cool trick, but it was not dedicated for this. The pioneers probably were of course Google with GPU, which are very much more advanced on their architectural level, but essentially the way they work, it kind of works for AI, But for LLMs, that starts to crack because they're so big and there's a lot of memory transfers and so on.

14:26Actually, that's why GROC achieves not GROC but a lot of GROC, C -ReeBress and all these folks, they achieve very high performance single -stream. It's because the data is right in the chip that I don't have to get it from memory, which is slow, which GP you have to do. So there's a lot of these things that ultimately make it a good trick, but not I would say dedicated solution per se. That said though, the reason probably in video one, at least in the training space is because of melanox, right, not because of the raw compute, because you need to run, you know, lots of DGPUs in parallel. So the interconnect between them is ultimately, what matters, right, so how fast can they exchange data?

15:11Because remember, when you do a matrix multiplication, Let's say you read the matrix is read like hundreds of times during the multiplication So there's a lot of transfers going on and so far Milan ox with you know infinity band at The best technology so that's why you know a lot of people and when you do training by the way It is the name of the game the interconnect when you do inference? Not so much you don't care when you do inference before we move to inference, I do just want to stay on chips and just say, okay, so we have TPUs, we have Nvidia, we have AMD. It's this in terms of distribution of gains.

15:50Is this a winner -take -all market? Is this cloud where you have several providers who are dominant? What is the distribution of gains that look like in the chip market? So I would divide it in two categories. Well, three categories. The GPUs you can buy or rent. The TPUs you can rent and the TPUs you can buy. This is how the market is structured today, right? Right now, if you want to go dedicated, at least in the cloud, there's two options, the TPUs and training. TPUs on Google, training on Amazon. So these are available chips you can rent them today. If you want to buy GPUs or rent GPUs, they're GPUs, we know it on the time.

16:33and there's this new wave of computing, which are dedicated chips, you can actually buy. The 10 -storant, the etched, the visorra. So I think it will be a mix of, for instance, let's say you are in Google Cloud, of course you don't want to do Nvidia. You get ripped off. Here's the dirty secret. Is that Nvidia, like a TSMC sells you at 60 % margin, Nvidia sells you at 90 % margin. And on top of that, there's Amazon that takes, let's say, a 30 % margin. So you are a very thin crust on a very big cake. It's a bit of a losing game if you go all in on one provider, you want optionality with increasing competitiveness within each of those layers.

17:18Do we not see margin reduction? Absolutely, yes. Yeah, yeah. Here's the problem, though. Let's say you are on Google Cloud and your own TPUs. Suddenly, you just remove that 90 % chunk on the spend. The problem is that for multiple software reasons, which we are solving, DML is that they're not really, I would say, a commercial success. They are very much successful inside of Google, but not much outside of Google. Amazon, same is pushing very, very hard for their training chips. So the future I see is that you use whatever your provider has because you don't want to pay 90 % outrageous margin and try to make a profit out of that.

18:01Okay, so when we move to actually inference and training, everyone's focused so much on training. I'd love to understand what are the fundamental differences in infrastructure needs when we think about training versus inference. So these two obey fundamentally different, I would say, tectonic forces. So in training, more is better. You want more of everything essentially and the recipe for success is a speed of iteration You change stuff you see how it works and you do it again Hopefully it converges and it's like you know changing the wheel of a moving car so to speak So that is training on inference.

18:40This is a complete reverse less is better You want less headaches? You don't want to be working up at night because inference is production You could say that training is research and inference is production and it's fundamentally different in terms of of infra, probably the number one thing that is, the number one difference between these two is the need for interconnect. So if you do production, if you can avoid to have interconnect between a cluster of GPUs, of course you will avoid that if you can. And this is why models have the sizes they have. So that people can run them without the need to connect multiple machines together.

19:22It's very constraining in terms of the environment. So that is probably the fundamental difference, the need for interconnect. And number two is, ultimately, do you really care about what your model is running on as long as it's outputting whatever you want it to output? Can you just help me understand, sorry? Why is training more is more and not great and in inference less is more? Why do we have that difference? Think of it like doing a painting and doing a million painting. The tools you will use, the process you will do, if you do one painting, what you favor is the speed at which you can do a stroke and do some iteration.

20:03If you do a million, what you want is a process, a process that is reliable that can deliver you efficiently a million paintings. So that is the same for training versus inference. If you run around millions of instances of a model, you cannot hack your way to do that. By the way, people do hack their way today, but this is probably the fundamental difference. How did people then put inference in production today? You know, we've seen with training, that's really what Nvidia have dominated so heavily. How do people put inference in production? There's a lot of duct tape. Here's also probably one of the problem is that training on first principle is actually two passes Forward and backward right it's called forward pass and backward pass right in front is running only the forward pass So that's how things are today There are people who are trying to specialize a bit because you know at some point duct tape doesn't really work out and when you run big scales that makes a problem and it's a problem that's growing because a lot of people are coming on a market with the needs for inference.

21:10That wasn't the case, you know, a year and a half ago or a year ago. Open AI had this problem, right? Maybe anthropic at this problem, but it wasn't a universal problem yet and now it's becoming an universal problem. Can you articulate what problem did open AI and anthropic have with regards to inference? So for instance, probably the number one thing depending on how you deploy, but if you're deploying in -front, the number one thing that will get you is what's called auto scaling. So as your systems get more and more loaded, you want to provision because these things are tremendously expensive.

21:45You want to provision them as you scale, right? So you want to say, I have a thousand GPUs, 24 hours, even if there's nobody on the production, I will pay for them. Which is, mind you, what people are doing today. This is crazy. So what you want to do is you want to, you know, provision, compute as you grow your needs, right? And you want to do it up and you want to do down. Probably the number one thing that, you know, gives you a lot of efficiency in terms of spend. Like we're talking, you know, multiples, like, you know, five, you know, sometimes 10x, you know, improvement. The thing is, this is a problem, at least in that, in, I would say, regular backend engineering, this is a problem everybody knows, right?

22:27Everybody is doing it because the savings are so huge, but on AI, nobody really had the problem. So now they're coming up to it. So this is one example. So the problem is that they're not doing provisioning, they're paying a shit ton more because they are fully in production all the time versus provisioned as needed. That's one example. Yeah. Another one is choosing the right compute. It's like kind of, I would say, a visual circle because provisioning compute is very hard. So if you lose compute, it's very bad. You are essentially incentivized to over buy in the case of, you know, Amazon or Google that would be buying reserved compute, which you're not going to use because if you buy it on the menu, you will get tremendously ripped off.

23:13So that creates this like face carcity of compute because that people buy preventively because they rate shit down money and they're not using it. So this is a major problem too. When you buy compute preemptively, does it not become outdated by the time you use it? It might well be yes. We are being spared a bit because Blackwell is late and all the others are getting canceled and so H series I would say are still you know in the active but yes absolutely but you know what choice do you have? This is the I think, well, we have a moment in time where there is this massive overhang or over -supply of compute, which we've proactively bought ahead of time, but then actually the hyperscalers go, would rather just burn it and buy fresh and we have the money to do that.

24:04So I might tell you that I think you already started. I'm getting cold emails for discounts from services I never heard about. And I started getting these emails probably around October, November. Some people are left with a lot of capex that they don't know what to do with. You know, it's a different thing to build a cluster and run a training run that it is to build literally a cloud provider, a hyper -scaler or whatever you want to call it. So there are a lot of people who do their training runs on the regular providers, but then move to regular hyper -scaler when they do production. So I very much worry there will be an over supply of these chips.

24:48The problem is is that remember the chips are the collateral. So somewhere in the US or whatever there's going to be a data center with like a thousand GPUs that people may buy 30 cents on the dollar. This is what might happen. What is the timeframe for that might happen? Probably this year. Janssen has made it very clear that inference opens up more avenue opportunity for Nvidia. He said that 40 % of their revenues today comes from inference. Right. To what extent is that correct? Or actually, as Jonathan at Grox said in the show, Nvidia is not meant for inference, definitely not. And actually, that market won't be won by Nvidia.

25:32Technically speaking is right, but realistically speaking, I'm not sure I agree. the thing is these chips are on the market. They're here. Al Tab on Chrome and get one. That is something that I don't take lightly. Availability, that is, right? I think Nvidia is just to stay at least if not for the H100 bubble bust because these chips are going to be on the market and people will buy them and do inference with them. Remains to see the OPEX and the electricity and etc. but The thing is the only chips that are really frontier on that sense are probably TPUs and then the upcoming chips. But the thing is, they're great chips, but they're not on the market.

26:18Or like there are outrageous prices, like millions of dollars to run a model. So what chips are great and why aren't they on the market? Let's say for instance, cerebrus, incredible technology, incredibly expensive. So how will the market value the premium of having single stream very high tokens per second? There is a value into that, right? As we saw with Mistral and Piaplexity, but I think that was done at the loss. I don't know. I don't have the details, but I think it was done at the loss that Syribras put it out. So today there's three actors on the market that can deliver this. I think this will be, I would say, the pushing force for change in the inference, landscape, agents and reasoning, so that is very high tokens per second only for you.

27:05What is forcing the price of a syriprous to be so high? And then you had Jonathan at Groc on the show say that had an 80 % cheaper than Nvidia. So there's this trick, because here's the thing, there's no magic. This little trick is called SRAM. SRAM is memory on the chip directly, so that is very, very fast memory. But here's the problem with SRAM. Is that SRAM consumes surface on the chip, which makes it a bigger chip, which is very hard in terms of yield, right? Because the chances of problems are higher and so on. So SRAM is, I would say, very, very, very fast memory, which gives you a lot of advantage when you do very, very high inference, but it's terribly expensive.

27:51And if you look at, for instance, Grog, they have on their generation, this generation, they have 200 and 30 megabytes of SRAM per chip. A 70B model is 140 gigabytes. So you do the math, right? Syriebres has a 44 gigabytes of SRAM into what they're called their wafer scale engine, which is a chip, the size of a wafer. I mean, most likely it's interconnected, but it's huge, right? And it has to be watercool. They have a copper, you know, like we'd say needles that touch the chip. to it's crazy stuff, very, very impressive technology by mind you, but very, very expensive. So my bet is, I think there will be chips on the market that do that at much lower price.

Read the full transcript

28:35And there's two companies I see going in that direction. One is called etched and the other one is called the V -sore. That's the two I see. Because if you can deliver this at the, I would say, the price that is comparable to GPUs, you've won. is minimizing SRAM the only way to reduce unit cost on these chips. Really. It's hard to say. I mean, you need some SRAM, but if you can, you know, have a smaller process node, but if you can hook yourself with external memory, then yes, you can do that a little better. But the thing is, if you go like full blown SRAM, then, you know, there's no magic. You will have to pay the price.

29:14I'm so enjoying this. I'm also loving my notes here are just expanding by the day. If that's today, how do you think the inference market evolves over the next three to five years? Pushed by reasoning. So reasoning, not in the sense that you see on deep seek and whatever, right? Reasoning and what's called latent space reasoning. Latent space reasoning and agents will push the market towards a different types of compute. What's latent space reasoning? So the way model's reason today is the reason it was in tokens. So it's as if if you think to yourself you would say out loud what you're thinking.

29:53So yes, it works, but it's a bit inefficient, right? And you lose information doing this. Latent space reasoning is this without going, I would say to English or whatever, right? So staying, you know, and what's called the latent space, which is where all the information of an LLM, Let's say, an LLM. An LLM lives. So this is very much how we work as humans. And we move toward what Jan LeCancel's energy -based model, in which we have different types of longer or shorter, I would say, thinking times, if you will. So that fundamentally, GPUs cannot deliver this. Plain and simple. At scale. Why can't GPUs deliver it?

30:37Because the access to external memory prevents it. So HBM is all the rage, right? But HBM compared to S -RAM is absolutely, you know, dark slow. So this is the problem you get. So HBM is like the best we can do, but it's still slow versus versus S -RAM. So when I had Jonathan on, he was like, Ashi, in video, I have such a stronghold because they're one of the only buyers of HBM, and that gives them this unique position. Ashi is being a sole buyer of HBM irrelevant. if the world needs that RAM instead. No, you want HBM to be clear. No, SRAM, this will not deliver. It's a dead end in terms of scaling SRAM, mean scaling the surface, mean you get depreciating problems.

31:24It explodes everywhere, right? So you need some SRAM, right? So we'll have bigger amounts of SRAM into chips. And of course, bigger was called external memory into chips. The issue with HBM is that it's still slow and yes, maybe Nvidia has a strong hold and they can prevent you from getting some. So that would be like, I call it the Nutella situation in which Nutella owns 80 % of the azoneats market, right? So yes, you can do a competitor, but who with your buy the nuts from, right? So there will be a need for HBM, there will be a need for SRAM. I would say better or more dedicated architecture will be able to deliver these things.

32:04And then there's like the next frontier after that, which is called compute in memory. There's two companies that are on that market. One is called Rain, Rain .ai. Samaltman is one of the investors. There's no surprise. The other one is called Fractile. So this is an X -Rontier. The idea is that instead of transferring the data between external memory and the CPU and do the compute there, you actually bring the CPU to the memory and you do everything. it's crazy stuff, but it's coming. How does it? Maybe not this year, but... How does that change the situation? It makes it much more efficient, but what does that actually mean in reality?

32:44It means you get maybe not a SRAM level performance, but you get very a lot faster performance in terms of compute. So, and if you translate that to LLM's, let's say, you get much, much higher tokens per second. In ACE, InGhostream, which is exactly what you want, when you go into reasoning. You want your model to maybe think, let's say, for like half a second and then boom. You don't want to wait 50 seconds and you know, context switch to some other thing, which is the problem everybody has today, mind you. So yeah, I think inference will be pushed, the compute landscape will be pushed to change because of these two constraints.

33:20I know I'm working on it. If you would describe value between training and inference out of a pool of 100, is it 80 inference 20 training? What does that look like? In five years I would say 95 % inference, 5 % training. Do you think Nvidia owns both of those markets in five years time? Depends on the supply. I think that there is a shot that they don't. Because here's the thing, even if we take same amount of, let's imagine we have a new chip from Amazon, right? That is the same amount. Oh, wait, we do. It's called Tranium. You know, why would I pay 90 % margin of Nvidia if I can freely change to Tranium?

34:04My old production is runs on AWS anyways. And like if you run on the cloud and you're running on Nvidia, you're getting, you know, squeezed out of your money, right? So if you're on production on dedicated chips, of course, you know, so maybe through commoditization, but you know, Hey, I'm on AWS, I can just click and boom, it runs on AWS's chips. Who cares, right? I just run my model like I did two minutes ago. With that realization, do you think we'll see Nvidia move up, stash, and also move into the cloud model? They are. They have a protocol name sort of does that. The thing with Nvidia is that they spend a lot of energy making you care about stuff you shouldn't care about.

34:49And they were very successful. Like who gives a shit about CUDA? I'm sorry, but I don't want to care about that, right? I want to do my stuff and Nvidia got me into saying, hey, you should care about this because there's nothing else on the market Well, that's not true, but ultimately this is the GPU I have in my machine. So you know off I go if tomorrow that changes Why would I pay 90 % margin on my compute? That's insane. This is why I believe it ultimately goes through the software Because the software like if my entry this is my entry point to the ecosystem So if the software abstracts away those in your synchroces as they do on CPUs, right, then the providers will compete on specs and not on fake modes or circumstantial modes, right?

35:35So this is where I think the market is going and of course there's the availability problem. There is, if you piss off Jensen, you might need to kiss the ring to get back in line, right? But I mean ultimately this is I don't see this as being sustainable from we chatted before you said about AMD And I said hey, I bought Nvidia and I bought AMD and Nvidia thanks to Jensen I made a ton of money and AMD I'm up 1 % versus the 20 % gain You said that Nvidia AMD basically sold everything to Microsoft to matter and had a GTM problem Can you just unpack that for me? So all I would say chipmakers have a GTM problem all of them whether you know It's Google whether it's AMD whether it's 10 store the problem is is that there's I would say Probably two fundamental problems the the number one is if you maintaining multiple stacks today is very very very hard So you don't so let's say I buy you know AMD I want to buy AMD, right?

36:41That means I'm going to abandon Nvidia. Oh crap, you know, I have a six year amountization plan on that. Oh man, what do I do? So do I need to support both stacks? And clear maybe and until AMD tells me, hey, you know, you have, I don't know, let's say a thousand and video GPUs You're about to buy a hundred thousand of AMD. I mean, come on, right? And I'm like, okay, that is you know makes it worse my while, right? But that is ultimately the fundamental problem is that the steps are very high, right? I need to have a lot of incentives to buy into the ecosystem. So I need to buy a lot of them. So if you're AMD, that is already a problem.

37:19But then Microsoft comes along and buys it all, makes by the way, OpenAI, or at least on the inference side, puts OpenAI in the green because of the efficiency gains. I'm just showing on the sunset. You're saying the switching costs are really high from one provider to another. Oh, yeah, absolutely. Which is why you don't or are you saying that to get into one of these bi -processes, you have to buy so much that it prohibits you. It's actually both. The buy -in is very high. So to make it worth it, you have to buy a lot. And if you buy a lot, this is what we talk to all of them. They always have the same questions and it's completely understandable.

37:57They say, this is great. But who's the customer? Because on the other side, let's take Amazon, for instance, with Ranium. Apple just came and said, hey, we're gonna buy 100 ,000 of them. So you want to buy, you know, 10 ,000, you feel like the big shot, right? Yeah, but go back to the queue because there's Apple before you, right? So they have to have very high commitments. You cannot be incrementally better. It's very hard, right? And also very hard. I can give you a, I can give you one metric if you want. I know for a fact that being seven times better and And whatever, take whatever metric you want, whether it's spend, whether it's whatever, is not enough to get people to switch.

38:37People will choose nothing over something. So this is a very hard market to enter into because you cannot also compete of incremental gains. It's very hard, right? You have to convince a lot of people. Maybe you can go the Middle East route in which they sprinkle everything and they evaluate everything. That's not, you know, very sustainable, I would say, strategy in the long term, but at least in the long term. What is the right sustainable strategy then? You don't want to go so heavy that you can't ever get out and you have that switching course. Right. But you also don't want to sprinkle it around and do, as you said, absolutely.

39:12The right approach to me is making the buying zero. If the buying is zero, you don't worry about this. You just buy whatever is best today. How do you do that by renting? Oh, because this is what we do. This is our promise. Artisus is that if the buy -in is zero, you completely unlock that value. Because you're free. When you say the buy -in is zero, what does that actually mean? It means that you can freely switch to compute freely. You just say, hey, now it's AMD, boom it runs. You just say, oh, it's 10 store and boom it runs. How do you need it then? Do you have agreements with all the different providers?

39:50Oh, yeah, yeah, yeah. Not agreements, but we work with them. to support dirt, dirt chips. But the thing is, my at least as, you know, I would say a user myself of, you know, our Varteck is that if it's free for me to switch or to choose whichever provider I want in terms of compute, right? AMD and VD, whatever. Then I can take whatever is best today and I can take whatever is best tomorrow and I can run both. I can run three different platforms at the same time. I don't care. I only run, you know, what is good at the moment. And that inlocks to me a very cool thing, which is incremental, you know, improvement.

40:26If you are 30 % better, I'll switch to you. So are you taking the risk on those on that hardware then? If you're the one providing them to turn off and on, on demand, provisioning, you name it, you take the risk. This is actually a great question. I think that if you are doing it bottom up and for two applications, you will lose because nobody will care. As they don't today, right? If you look at TPUs, they're available, they're great. Nobody cares. Why does nobody care about TPUs? Because the cost of buying, it's always the same, right? You have to spend six months of engineering to switch to TPUs.

41:02And mind you, TPUs do training. They're the only ones with training now. But AMD can do training, but it's also, but in terms of maturity, the by far the most mature software and compute is TPUs and then it's Nvidia. Right? So the buy -in is so high that people are like, we'll see, right? I'm not on Google Cloud, I have to sign up. Oh my God, right? So these are tremendous chips, these are tremendous assets. Now in terms of the risk, I think if you want to do it, you have to do it top to bottom. You have to start with whatever it is you're going to build and then permeate downwards into the infrastructure.

41:41take for example, Microsoft with OpenAI, they just bought all of AMD supply and they run you know, ChadGPT on it. That's it and that puts them in the green. That's actually what make them, you know, preferable on inference. So, or at least let's say not lose money, right? I'm sorry, how does Microsoft buying all of AMD supply make them not lose money on inference? Just help me understand. Because I can give you like actual numbers. If you run 8, 800. You can put two 7 TB models on them because of the RAM. That's number one. Number two is, if you go from one GPU to two, you don't get twice the performance.

42:19Maybe you get 10 % better performance. Yeah, that's the dirty secret nobody talks about. I'm talking inference, right? So you go from let's say 100 to 110 by doubling the amount of GPUs. That is insane. So you'd rather have two by one, then one by two. So with one machine of 8, 800, you can run 270b's model. If you do four GPUs and four GPUs, that's number one. If you run on AMD, well, there's enough memory inside the GPU to run one model per cal. So you get AGPUs eight times the throughput. When on the other end you get AGPUs to maybe two and a half times the throughput. So that is, you know, a Forex right there.

43:04Just, you know, by virtue of this, of this. So that is, you know, the the compute part. But if you look at all of these things, there are tremendous amount of, you know, we talk to companies who have chips of coming with almost 300 gigabytes of memory on it, right? So that is more, you know, a model, like one chip per model. This is the best thing you want. If you want 70b's, right? So which is what, But I would say not the state of the art, but this is the regular stuff people will use for serving. So if you look top to bottom and you know what you're going to build with them, then it's a lot better to do the efficiency gains because four times is a big deal.

43:42And bind you, these chips are 30 % cheaper than Nvidia's. It's like a no -brainer. But if you go bottom up and say I'm going to rent them out, people will not rent them. Simple. So that's why, you know, I think it's a good way to attack it from the software because ultimately, do you really care about that your MacBook, let's say, is an M2 or an M3? It's like, oh, it's the better one. And that's it, right? And imagine if you had to care about these things, that would be insane. When I listen to you now, I'm like, shit, I should sell my Nvidia and buy more AM. Okay. If you were forced to buy one, I'm not saying sell the other, I'm not saying like this on the other, buy one, which would you buy and why?

44:27Stock? Yeah. I used to think the market was efficient, so probably I would go today at least, I would go within video still, because the supply. But you know, if we play our cards right, we ship our stuff, hopefully I will come back and tell you to buy AMD as much as you can, or a 10 store, you know, if they go public or whoever else. These chips are amazing, by the way. What does everyone think they know about inference that they actually don't? Or what does everyone get wrong about inference? Probably not a lot of people are accustomed to what it entails to run production. So that inference is production and production is hard.

45:09Somebody has to wake up at night. And I used to be that guy, right? I don't want to do it again. So production is hard. Thankfully, we have a lot of software nowadays to do that a lot better But there's not a lot of reuse because the AI field at least is not really a custom to that yet It's changing, but you know the discussions I had you know a year ago and the discussions I had today are not the same They're going into the right direction, but they're not there exactly yet So it's probably that would be the number one thing that is only you know training code running only for what best right?

45:45This is not what it is. Can I say how do you evaluate the data center investment we're seeing being made? You know when you look at Facebook doing 60 to 65 Microsoft doing 80 and some of the intense CapEx expenditure that you're seeing how do you think about that on the data center side? I mean, there's still going after training. So there's still this frontier probably. It's why also and VDIS is the better buy right now. Because on the Nvidia side, if you do training, it's incremental. If you have bought a thousand Nvidia GPUs and you buy a thousand new Nvidia GPUs that gives you 2000 GPUs, right?

46:19But if you buy a thousand and a thousand AMD that gives you twice a thousand, right? It's a bit different. So they're still going after training, definitely, and they're very pragmatic in doing so. But I mean, they have the KAPX to spend. They're not making their money out of it probably. The only one, by the way, that owns their compute are Google. There's like this triangle of, I would say, of wind that I, this is my mental model, mind you. You have the products, the data, and the compute. Who has all three? And you get everything flows from there. Product is data compute. He has all three. Google, Amazon.

46:55Amazon, they don't have products. They have Amazon, right? They have AWS, but they don't have actual products. Google, that's like, you know, Android, Google Docs, whatever they have, everything, they can sprinkle everywhere. This is the sleeping giant in my mind. If they're not busy doing a rig, they might... It's fascinating, because everyone, if you are a shallow thinker, you think the OpenAI challenges that Golden Goose, which is such a Googleist threat all the never -know. Well, I mean, OpenAI is amazing, but it's not their compute. It is Microsoft's compute. And if you own your compute, you own your margin is essentially what you're saying.

47:32Yeah. Microsoft, even Microsoft, they bought, when they weren't running Nvidia, they bought Nvidia, you know, some outrageous margins. I talk to a lot of people that build the dissenters and I tell them, you know, they're mind you, these people like by tens of thousands of GPUs. And I ask them, hey, do you get at least a discount or something? And they're like, no, the only thing we get is the supply. So, I mean, ultimately, if you don't own your compute, you're starting with, you know, something at your ankle, definitely. And so, this is why I like to think in this, like, this triangle product data compute.

48:07And you can see where everybody sits and their weaknesses and their strengths. Can I ask you if we move a little bit? You said, it's totally rational that everyone's focusing on training still. When we think about that, it's rational if you think that efficiency and scaling will continue to continue play such emphasis on it. How do you think about model scaling and scaling more coming into place? How do you think about it? There's like a brute force approach to this. It is a very American approach. More and more but the thing is you look at for instance the XAI cluster. It's not a hundred thousand GPUs.

48:43It is four times twenty five thousand. You're starting, you know, see some because in Finni band and Nikis Rocky, which is anyways, the technology they used to bridge their GPUs together, you have a profound, right? At some point you're fighting physics. So you can push, it's like, you know, trying to get to the speed of light. As you approach it, the amount of energy you need is a lot higher and a lot higher and it grows and grows. So there's two, I would say, counter to that would be that number one is we still scale, but there's a lot of waste and and access, you know, spending on the engineering side, which is the deep seek approach, right?

49:21Very successful at that, mind you. They said, yeah, if we do this and this differently, then we get, you know, multiple sometimes, right? So virtually you increase your compute capacity because you're more efficient. And the other approach is Jan's, Jan LeCanel's approach, which is, this is not scaling, and at some point, you need, we need to look at the problem in the face and do something better, right? So of course we push and push, push, because there's capital still, but I'm more of these two approaches. I think you can do more with less. At what point do we stop and say, hey, there is a lot of wastage and we could do more better.

49:58I think until somebody does it, DeepSeek was a good wake -up call, right? Suddenly, efficiency is in. That's number one and number two is until there's a new architecture that comes out and changes the game. So in the case of a LLM, for instance, you have these what's called non -transformer models that changes fundamentally the compute requirements. So that might be a frontier that completely absolutes the transformers. And if the transformers are the, I would say, the building block by which current model work, right? So the way they work is that for each token or a syllable, if you will, the model will look at everything behind it.

50:36So you can see that as you add more text, you have more work to do. So there are these new architectures that do not require this that might change, you know, these things and probably shift the amount of compute needed to do training or to do inference. And then there's the new thing which is Yann's thesis, which is the word model. As in LLM's are at the end, what we need is something that understands the world fundamentally, and this is it's JEPA thesis, it's called. I'm very bullish on this, but it's very frontier. Why you bullish on it and why is it so frontier? Because it's young look on.

51:14It's hard to. He's no bullshit, right? So he explained to me how it worked and I was blown away. But it makes a lot of sense. We are creeped out because the machine talks back to us. But it's not a new thing, right? It used to, you know, this is not new technology when it came out, like when it exploded, it was a new technology. But suddenly it was talking back and that freaked us out and we got crazy on it, right? But language is one form of communication, but it is ultimately a very narrow window into the world. We use it to describe the world, arguably with some loss, right? And so the JIPA approach is, long story short, is that you have essentially two things you want to do and you try and minimize the energy to do them.

51:59And from this understanding emerges, physics emerges, etc. Because you're trying to minimize the amount of energy to go from one state to the other. And that actually makes sense. Like if you try and pick this air pod case, I'm not going to go around trip around the block to get it, right? I just get it. And in my brain it's wired to just do the thing. If I go and talk to myself out loud, put the hand down, move to the left and whatever, that feels very inefficient. So probably this will be something that changes. And in the case of the LLMS, there's good work also on what's called diffusion -based LLMS, which means like instead of thinking what's called autoregressively, that means you get a new token, you re -inject, and you redo, et cetera.

52:46They think more like what we do, which is in patches, right? Imagine a paragraph of text and words appear until it's done. Is distillation wrong? And if we're all progressively moving towards a better future for humanity, more efficient models, it's dilation, not effectively open -source in another round -per. I think it's fair game to be honest. I will not share the tier. It's fair game. If you, there were some people who tried to ask, I don't remember if it was an OpenAI model, so a diffusion model image, right? They asked it to generate an image from a Star Wars movie at whatever time stamp and it came out with the Star Wars movie, you know, screenshot.

53:30Obviously it was strained with it. I think it's fair game because there's no free lunch, right? It was strained with data. You had a good ride, somebody was sneaky and took it, but you took it from the beginning too. So let's just accept it's fair game. And you can also learn from their advancements. Absolutely, absolutely. I take my cup and enjoy it very much that movie every single day. You mentioned the training there. Obviously data and data quality dictates a lot of training ability. When you think about the future of data that feeds into training, how do you think about how that will be between synthetic data versus real data?

54:15I'm a bit split on this. There's a part of me that said that if you re -inject data into the system, the system deteriorates. That feels a bit, I would say, intuitive, but if you look at AlphaGo for instance, the moment it's ramped up in its skills is when they started generating games, synthetic games, right? So I'm a bit, you know, split, but there are some verticals that very much benefit from this code, LLM, for instance. We can run code, right? So this is the poolside thesis. Just so I understand why it is it work for coding and not for other things. Because you don't use the mod, the AI model to generate output.

54:53You use the machine, you just run the code. And you see what it makes and you run all these code and you create data out of it. Whereas if you run the LLM and you say to an LLM, or I generate me two trillion tokens of text, it will do it with it's, you know, you may inject and stuff. So there's a lot of tricks, but ultimately my guts tell me that it feels wrong, right? Because you re -inject Data that I was there and so it will detail you right there's you know, there's loss So yeah, I'm a bit bullish. I'm not sure exactly on what vertical code is one We'll see distillation is in some sense a bit like that you create synthetic data from a bigger model into a smaller one Probably the most, I would say mind blowing thing about distillation is that sometimes the smaller models become better than the bigger model through distillation.

55:45So all the models become better than bigger models purely because of the quality of the data that's inputted through them. The one theory is that the smaller model is better at generating output that you would want it to generate essentially. It's not better in the general sense. It's better at the task at which you were measuring it. This is what it learned to to imitate. How do you think about the future in terms of large monolithic models versus more dynamic architectures smaller models? Sometimes it's wasteful to run big models a lot of times. It's actually wasteful to run big models I think there's going to be a lot of smaller models for efficiency reasons, but there's a there's a but which is you talk to people at DeepMine and they don't even find tune anymore.

56:35Because they have such what's called big context window, which is what the data at the model you inject at runtime, that nowadays they just dump data into it and just say, do whatever that data tells you to do instead of fine -tuning as we used to do. So if the efficiency gains were not there yet, but if the efficiency gains, I would say pass that threshold, we'll just do it at runtime. We'll just have a great model that we'll just specialize at each request. But that's not for tomorrow, I think. What is retrieval augmented generation fast? It's a very clever trick. What you do is you represent knowledge into what's called the vector space or latent space.

57:23And what you do is through what's called vector search. So imagine you have, Let's say a 3D space that represents all knowledge, all of everything. Let's say a cat sit here, a dog sit close because it's an animal, but it's far from some other property and so on. What you do is you run the user's request through this same system. It's called an embedding. That will give you a vector and you will take whatever is closer to you. What's called symmetrically close. And then it's actually very clever. You actually insert those pieces of text before the request. So it's as if you would say knowing the following and you give the data, let's say it's law or whatever, please answer my request.

58:12And that's it. So that's a bit of a clever trick. It's a bit dirty because of course, you know, you are limited by the amount of data you can input. right? So there's this problem in which how do you chunk, you know, the data that you input? A lot of things we do not retrieve relevant generation, and when we say here's a link, summarize it to the key points. Is that not right? Because we're inputting the data. It is. It is. It depends on how it works, but yes, sometimes it is. But think of it as in, it's like a preamble to your question. Knowing the following, and the following is a tiny window So into the content, please answer my question.

58:52And of course, as you talk more and more, it will forget because that window is fixed. And how does that shift the movement from large generalized model to smaller, more advanced models? What pushes smaller models are efficiency, roughly speed. You know, less is better. So if we can do it less, then less it is. Simple as this, right? In terms of RAG, the key frontier is what we call attention level search. But this is something we're working on. You have the exclusivity now, I'm putting it out there. It doesn't push, I would say, model sizes. What really pushes model sizes are the efficiency rather than specializing.

59:29Meaning that if you can do the same performance with a smaller model that is fine -tuned, with a rag or whatever, then you'll do it with a smaller because, again, less is better. Kashi, before we move into a great far -run, I just want to ask you, when we had deep seek as we mentioned, And to what extent were you surprised that such innovation, I would argue, and I think many would agree with me, came from a Chinese competitor, not from a Western competitor. Oh, I love it. Constraint is the mother of innovation. Yes, we can, we can, you know, troll a bit about, you know, the Singapore, you know, grey market and all of these things.

1:00:07But ultimately, like, they had no choice. Here's the thing. If you can buy more, why would you give a damn? You can just buy more. So if you are pushed to efficiency, then you will deliver efficiency. These are very, very skilled people. This is the coolest thing to me about AI, honestly, is the geography doesn't matter anymore. You can just do things. You appear out of nowhere, boom. You're on the map. So I'm very, very glad that they did. I found the reaction very entertaining, to be honest. So yeah, I mean, constraint is a very good driver of efficiency. Do you think it is a meaningful threat to open AI and chat GPT?

1:00:45Bluntly, they still have the consumer loyalty, the consumer brand. To what extent is it actually a long -term threat? I'm not sure who is a threat to open AI at the moment. Here's why. You look at the numbers. I mean, we live in a bubble. We follow every new episode, the whatever new model, whatever who said who it and so on. But I go to my mother and I ask her, you know, do you know chat GPT? and she says, yes, and you know, I don't know. I don't want to done anybody, but do you want to know some other model? And she says, what it is? What is it? Right? Even Gemini, right? Like a Google, right?

1:01:18So they have a strong brand. They have a strong product, but there's a balance between the product and the models, honestly. So this is Gary from fruit stack, actually, who told me that is mental model in terms of model providers. It'll be like car makers. There's no winner tickle. Everybody will have their own, because ultimately, also human knowledge is Everybody has everything. So we're converging. But I liked the analogy. Yes, deep seek made a very good, you know, made waves, but it was it was, you know, waves that were amplified by the media and the narrative and the drama. Do you think export regulations inhibit China's ability to compete in any way?

1:01:58Today, maybe, tomorrow, I'm not sure. They're a bit late in terms of, you know, AZIC. They are like A100 level, but they have probably, I would say, one of their unfair advantage is that it's like, you know, when you do exercise in the water, right? It's like this, so this is their state. They are constrained, so they are bound to do better. They can just not buy their way into better compute. So I think it hinders their success, Yes, but I think it's short term, do you think that way? Are you fearful that Europe are going to regulate ourselves into constraints in a world of AI? No, I don't care.

1:02:40This is something that makes me wonder sometimes, I understand the narrative and so on, but I am absolutely not fearful. Let's be successful first and then we'll talk about the politics. I have so far, but again, I'm not mistral. I'm not, you know, I'm not building gigawatt that is the tensenters and so on. So if you build gigawatt that the centers you're running through these problems, but maybe you're running through these problems. But the thing is, if you're successful, everything flows from there. From there. Steve, I'm being direct here, but I'm asking you for the pros. Everyone says, Mr.

1:03:12Al, just doesn't have enough money to compete. That is kind of word on the street. To what extent is that fair? There are very competent. I think it's easy to spread further. There's a lot of fun going around, especially about regulation and everything. But here's the thing, I look around me and I don't see what I read. So I am hardly convinced about everybody was saying that they were dead and boom, they came out with their release and it was insane. So what I know is that I hope they don't have too much money. That's for sure. You want to be clever, right? Final one if we do a quick five, so enjoyed this Steve.

1:03:49Final one if we do it. where Stalgate was a $500 billion dollar announcement. How did you evaluate that? My first impression was that I don't buy it. I would say, you know, American start, right? You start with the claim and we'll figure it out later. I don't buy it. And ultimately, ultimately, I'm not sure I cared that much about it. Let's imagine it's true, right? Congratulations, amazing. But it is more of the same. It is a vertical scaling. And as you know, my days are spent on efficiency. So I look at these things as being like, all right, this is a bigger, you know, this is an American car of AI It's big it consumes a lot of gas, but ultimately, you know, it's not a good car I think there has to be you know sufficient capital, but at some point I'm not sure it is really a differentiator That was piled to deep seek then deep seek came there was always you know my thesis, but you know you need money You need infrastructure, you need, but what is ultimately the probably the two limiting factor today is talent and energy?

1:04:54That's it. The rest, you know, yes, of course you can buy 500 billion of GPUs. I by the way, 90 % margin. So if we work on that margin, we can you know, shrink that number probably. So I'm not easily entertained by these numbers. I've seen how the sausage is made way too many times. Dude, I want to do a quick fire with you. So I say a short statement, you give me your immediate. Sure. If you had to bet on one major shift in AI infrastructure over the next five years, what would it be? Oh, yeah, latency, reasoning, definitely. This year, what does that mean? So the shift from throughput, so how speed my answers to how long it takes for my answer complete to appear.

1:05:38That is probably one of the fundamental, like this year, right? Longer term, I'm very rooting for non -transformer models that will change the compute also landscape. And of course, you know, world models, right? Yes. And or energy -based models. What's one piece of advice you give to AI startups navigating the changing landscape of training inference in hardware? Probably the number one thing I would say is do not resell compute if you can. A lot of AI startups that are building on top of AI are trying to make a margin on top of a very big cake and Ultimately what they sell is compute if you look at the dollar of spend, you know for one dollar of spend Maybe 98 % of it goes to somebody else's margin So if you do AI as much as you can try to verticalize on on the product But not on the compute if you are you know if your business model implies buying a lot of tokens It's a very hard circle to square to put that into $20 a month.

1:06:42I always say, please look at it from that angle. If you can, try and avoid it. What's the biggest challenge that Jensen Huang faces today? The highs are very high, but they don't last forever. So probably it's how to navigate the downslope. Blackwell is probably something that keeps him awake at night. Why would that keep him awake at night? with that not re -energize him, more orders, new enthusiasm, new product, baby. Because orders are getting canceled. Why are they getting canceled? They have a lot of problems with this chips. So a lot of people are canceling their orders. These chips are like on the frontier of scaling.

1:07:23And so they were supposed to come out last summer. But that heat dissipation and matter bending problem, it used to be called the people who are very privy to Silicon told me, this is what we call a pretty big fricking problem. And quote, probably how to navigate the downslope. Maybe you don't know, but the supply of H100 was actually a smooth out over the year so that they decided that they didn't have like a big spike in deliveries and then a quarter or less, right? Which pissed a lot of people, mind you, who bought a lot of them. Some of them even haven't received their order from last year.

1:08:00And they already see like the new chip, the B200, and then the one after, you know, and they're super pissed. There will be a downslope at some point. The question is, you know, when, how, like if there's like the H100 bubble, of course, it will impact Nvidia. But Blackwell is, I'm probably going to get a lot of flag for this, but, you know, I've seen some very worrying numbers about it and varying testimonies about people who operate these things, right? So that ride will stop. or at least, you know, slow down. Steve, I'm not sure I've ever learned quite as much in one episode. Seriously, we said before.

1:08:38Oh, wow. No, I love what I do because I'm able to ask anything to the smartest people in their business. And I so appreciate you unpacking so much with me today, man. I'm thrilled to say that I actually finally get what you do after years. That's the thing. That's the thing. But you've been a star. So thank you, man. Thank you, appreciated. Thank you. I mean, I said it there. I think I learned more in that episode than I have done in the last 1000 when it comes to technical specs and the future of AI. Steve was incredible. If you want to watch the episode you can find it on YouTube by searching for 20VC, that's 20VC.

1:09:17But before we leave you today, turning your back of a napkin idea into a billion -dollar start -up requires countless hours of collaboration and teamwork. It can be really difficult to build a team that's aligned on everything from values to workflow. But that's exactly what Coda was made to do. Coda is an all -in -one collaborative workspace that started as a napkin sketch. Now, just five years since launching in beta, Coda has helped 50 ,000 teams all over the world get on the same page. Now at 20VC, we've used Coda to bring structure to our content planning and episode prep, and it's made a huge difference.

1:09:53Instead of bouncing between different tools, we can keep everything from guest research to scheduling and notes all in one place, which saves us so much time. With Cody you get the flexibility of docs, the structure of spreadsheets and the power of applications, all built for enterprise, and it's got the intelligence of AI which makes it even more awesome. If you restart up team looking to increase alignment and agility, Cody can help you move from planning to execution in record time. To try it for yourself, go to coder .io slash 2 -0 VC today and get 6 free months of the team plan for startups.

1:10:27That's coder .io slash 2 -0 VC to get started for free and get 6 free months of the team plan. Now that your team is aligned in collaborating, let's tackle those messy expense reports. You know, those receipts that seem to multiply like rabbits in your wallet, the endless email chains asking, can you approve this? Don't even get me started on a month and planet when you realise you have to reconcile it all. Or Plio offers smart company cards, physical, virtual and vendor specific, so teams can buy what they need while finance stays in control. Automate your expense reports, process invoices seamlessly and manage reimbursements effortlessly, all in one platform.

1:11:06With integrations to tools like Zero, QuickBooks and NetSuite, Plio fits right into your workflow, saving time and giving you full visibility over every entity, payment and subscription. Join over 37 ,000 companies already using PLEO to streamline their finances. Try PLEO today. It's like magic, but with fewer rabbits, find out more at pleo .io -4 -slash20vc. And don't forget to revolutionize how your team works together. Rome, a company of tomorrow runs at hyper speed. With quick drop -in meetings, a company of tomorrow is globally distributed and fully digitized. A company of tomorrow instantly connects human and AI workers.

1:11:46A company of tomorrow is in a Rome virtual office. See a visualization of your whole company, the live presence, the drop -in meetings, the AI summaries, the chats. It's an incredible view to see. Rome is a breakthrough workplace experience loved by over 500 companies of tomorrow for a fraction of the cost of Zoom and Slack. Visit Rome, that's O -R .am for an instant demo of Rome today. Nobody knows what the future holds, but I do know this. It's going to be built in a Rome virtual office, hopefully by you. That's Rome -R -O .am for an instant demo. As always, I so appreciate all your support and stay tuned for an incredible episode coming on Wednesday with Oscar, found at Glovo, on turning Glovo into a $2 billion business.

From the publisher

Steeve Morin is the Founder & CEO @ ZML, a next-generation inference engine enabling peak performance on a wide range of chips. Prior to founding ZML, Steeve was the VP Engineering at Zenly for 7 years leading eng to millions of users and an acquisition by Snap. 

In Today’s Episode We Discuss:

04:17 How Will Inference Change and Evolve Over the Next 5 Years

09:17 Challenges and Innovations in AI Hardware

15:38 The Economics of AI Compute

18:01 Training vs. Inference: Infrastructure Needs

25:08 The Future of AI Chips and Market Dynamics

34:43 Nvidia's Market Position and Competitors

38:18 Challenges of Incremental Gains in the Market

39:12 The Zero Buy-In Strategy

39:34 Switching Between Compute Providers

40:40 The Importance of a Top-Down Strategy for Microsoft and Google

41:42 Microsoft's Strategy with AMD

45:50 Data Center Investments and Training

46:40 How to Succeed in AI: The Triangle of Products, Data, and Compute

48:25 Scaling Laws and Model Efficiency

49:52 Future of AI Models and Architectures

57:08 Retrieval Augmented Generation (RAG)

01:00:52 Why OpenAI’s Position is Not as Strong as People Think

01:06:47 Challenges in AI Hardware Supply

 

More from The Twenty Minute VC (20VC): Venture Capital | Startup Funding | The Pitch

All 521 episodes
20VC: Why Google Will Win the AI Arms Race & OpenAI Will NotThe Twenty Minute VC (20VC): Venture Capital | Startup Funding | The Pitch · 1 h 13 min
Listen in VO