The "Android Moment" for AI Infrastructure: Why Modular Just Raised $250M

26 Nov 2025 · 1 h 1 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

The Neuron Podcast: Episode Summary

Episode Title

The "Android Moment" for AI Infrastructure: Why Modular Just Raised $250M

Hosts

  • Grant Harvey
  • Corey Noles

Guest

  • Tim Davis, Co-Founder & President of Modular

Episode Overview

In this episode, Tim Davis discusses the recent $250 million funding round for Modular, a company focused on revolutionizing AI infrastructure. The conversation delves into the current challenges in AI deployment related to expensive, vendor-specific hardware, and how Modular aims to create a more accessible and efficient platform for AI development akin to the Android operating system for mobile devices.

---

Key Concepts and Discussions

  1. AI Infrastructure Challenges
  2. AI models are often constrained by proprietary hardware.
  3. There is a lack of flexibility for AI engineers to deploy models across different hardware platforms (e.g., NVIDIA, AMD, Apple Silicon).
  4. Modular aims to solve these issues by creating a "hypervisor for AI" that allows code to run on any GPU.
  1. Funding and Valuation
  2. Modular raised $250 million at a $1.6 billion valuation.
  3. The funding supports their vision of democratizing AI development and making it more cost-effective.
  1. Modular's Vision
  2. A unified compute layer that abstracts hardware complexity.
  3. The goal is to allow developers to focus on throughput, latency, and cost without getting bogged down by hardware specifics.
  4. The importance of creating a programming model that caters to various hardware accelerators.
  1. Historical Context
  2. Tim shares his background in AI from his time at Google Brain and the creation of TensorFlow.
  3. The episode compares the current AI ecosystem to the evolution of mobile operating systems, particularly Android.
  1. Economic Implications
  2. Companies using Modular’s platform have reported significant cost reductions (70-80%).
  3. Lowering costs can increase technology penetration and adoption among businesses.
  1. Future of AI Development
  2. The conversation touches on the potential paths to superintelligence, questioning whether merely scaling models will lead to true intelligence.
  3. Emphasis on the need for reinforcement learning and continuous learning, rather than just scaling current models.
  1. Importance of Interpretability
  2. Tim voices concerns about the lack of understanding of AI models. He argues that it’s crucial to comprehend how these models function before deploying them widely.
  1. Modular’s Products
  2. Mojo: A new programming language designed to enable cross-platform AI model development.
  3. Max: A modeling and serving framework that integrates with various hardware types for efficient AI deployment.
  1. Call to Action
  2. Developers are encouraged to explore Modular’s platform, emphasizing that it is free to try and offers streamlined processes to get started with AI infrastructure.

---

Conclusion The episode wraps up with Tim emphasizing the potential for Modular to change the AI infrastructure landscape by providing a more flexible, accessible, and efficient platform for developers. The discussion reflects on the broader implications for innovation in AI and the necessity for understanding the technology being deployed.

Subscribe & Follow

  • [The Neuron Newsletter](https://www.theneurondaily.com/subscribe)
  • [Modular Website](https://modular.com)
  • [Getting Started with Modular](https://modular.com/get-started)

---

Key Takeaways

  • The AI infrastructure landscape is ripe for innovation.
  • Modular's approach could democratize AI development, akin to Android's impact on mobile.
  • Understanding AI model deployment and operational efficiency is crucial for the industry's future.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Most folks may not understand that when you take a photo, you're using AI. When you search for something, obviously, you're using AI. Even when you pull down and look at applications on your phone, you're using AI. Increasingly now, even when you're typing on your keyboard, you are using AI. There is pattern recognition in the predictive text that you're doing. Humans are a similar creature, and I think recommendation systems have proven that over time. All of that needed to be powered by infrastructure. The world is actually quite hard for an AI engineer. It's not like you just train a model and then amazingly, it just deploys everywhere.

0:28That's not actually how the infrastructure works.

0:39Welcome, humans, to the Neuron Podcast. I'm Corey Knowles, as always, joined by the one and only Grant Harvey, who writes the Neuron every day. Today, we're going to talk AI infrastructure and whatever else comes to mind with Tim Davis, co-founder and president of Modular. it's a company that has raised 250 million dollars to change how ai runs on hardware tim welcome to the neuron yeah thanks guys thanks thanks very much for having me congratulations

1:07Tim Davis:on the raise as well you know that's exciting stuff yeah i mean um you know i've always viewed you know so fundraising uh you know is an interesting process i'm happy to walk through it and and we have a unique methodology i think uh that we've taken in in fundraising but um but yeah in many ways it's just uh you know we're excited that we have people who believe in in our vision and what we're trying to achieve. And yeah, I'd love to talk to you both about that and other questions today. Well, I guess to start out, you're with Modular and I understand that you and your co-founder are both from Google Brain, right?

1:35Yeah. So maybe a little bit about me and a little bit about the background of the company. So I spent, yeah, it was probably six years, six or seven years at Google, a lot of that inside Google Brain. And that's actually where I met Chris Flattner, who is a very renowned engineer. And he was brought in to help bring up some of the TPU infrastructure, which is sort of Google's custom silicon for running AI models. Before that, Chris was at Apple. And then before that had done a lot of other things in his career, LLVM, which is an open source compiler foundation. And so we had been working on an infrastructure platform called TensorFlow, which was one of the original machine learning and AI frameworks at Google.

2:24And Google had released that really to, you know, provide the world with the ability to build, you know, AI models and in many ways help, you know, this was back in 2015, 2016, help begin to democratize AI. They had actually just acquired DeepMind at the time. And so, you know, I independently became very excited about AI. I had had a startup before joining Google and, you know, we had started to dive in more. What, you know, now is probably traditional machine learning. Yeah. So sort of, you know, what folks would... Sentiment analysis kind of stuff and... Yeah, I think, you know, more traditional techniques for, you know, essentially recognizing patterns in data and predicting outcomes.

3:12and then you know as deep learning began to to accelerate and so the complexity of sort of these networks that we started to build we started working on you know sort of open infrastructure for the world to help accelerate those things and so I worked on you know not only TensorFlow with an incredible group of people inside Google Brain you know I also spent a lot of time looking at edge so scaling this out to mobile phones penetrating Android devices so you know most folks may not understand that when you take a photo, you're using AI. When you search for something, obviously you're using AI.

3:46Even when you pull down and look at applications on your phone, you're using AI. Increasingly now, even when you're typing on your keyboard, you are using AI. There is pattern recognition in the predictive text that you're doing on your keyboard. Shopping on Amazon. 100%. I mean, humans are a similar creature and I think recommendation systems have proven that over time. So all of that needed to be powered by infrastructure. And when you're a company as significantly large as Google, you know, trying to standardize on a piece of infrastructure is really important. But one of the things that happened, and I think, you know, what Chris and I started to realize when we were at Google was Google really had like, and it's probably interesting for your audience, they actually had three stacks.

4:30And this is like full software stacks. They had a software stack for TensorFlow that was sort of optimized for TPUs, which is their proprietary silicon. They had a software stack that was optimized for CPUs and GPUs, which is now what the world would know as, for example, NVIDIA machines or AMD machines. And then another software stack that was optimized really for edge devices. So think of Android phones or iPhones or even smaller things like, you know, microcontrollers that run around in, you know, in all sorts of different, you know, robots and systems like that. And so, you know, what we quickly began to realize was the world is actually quite hard for an AI engineer that, you know, believes in deploying not only in large scale data centers, but all the way out to, you know, edge based environments.

5:22It's not like you just train a model and then amazingly, it just deploys everywhere, right? That's not actually how the infrastructure works. Yeah. And so I think what we realized was, well, you know, what, what we really need to try to build is a platform where, you know, maybe that future is actually possible where you could have in some ways abstract away all this hardware complexity, the idea that, you know, and I think this is a, you know, goes into everything for modular, which is a unified compute layer, which I'm happy to talk about. But I think one of the biggest sort of product lessons, one of the most interesting developer lessons from our time at Google was realizing, and this will be a controversial statement, so prepare yourself.

6:03Most developers don't care about the hardware.

6:07Tim Davis:Now that will be in a world where - Oh, that's not controversial. I feel like that's very well understood. Great. Because in a world where, you know, NVIDIA is an incredible, you know, four and a half billion dollar company, it sounds weird to say that because in some ways the construct is, well, clearly the hardware is the most important thing. Every developer that I've talked to, including myself building applications, you know, what I care about is what is the throughput? What is the latency? What is the accuracy target? And what is my cost, you know, threshold? Once I have those, they're my inputs.

6:37You know, fundamentally the hardware is sort of the output that then has to meet the requirements. And you end up having to, you know, in some ways be exposed to it because you end up not having many choices actually at the end that can meet those requirements and so we realized well you know and we can get a little bit into what we you know in our thesis on the future of super intelligence but you know we strongly believed and there was this concept at google of what you train is what you serve and in today's world that actually doesn't exist when you train a model uh you know whether you're using pytorch or or other frameworks you actually need to do a lot then to to serve that at scale in production It's not like it just instantly happens.

7:16And so we realized, well, hang on a second, if we could build a platform that not only helps abstract away some of this compute, but began to work towards this vision of what you train is what you serve, then maybe a very significant contribution to the planet and to humanity is helping to realize this sort of vision of superintelligence. All right, so if you're building anything in AI right now, whether it's models, tools, workflows, you're going to want to hear about this. Dell just dropped something that's honestly in a league all of its own, the Dell Pro Max with GB10. That GB stands for Grace Blackwell, which is NVIDIA's next generation architecture, and here's why you're going to care about that.

7:54This machine looks small, but it's a powerhouse. You're getting 128 gigs of unified LPDDR5X memory, super low latency, and the brand new NVIDIA GB10 module, which lets you run local inferencing on models all the way up to 200 billion parameters, which is insane. No cloud queues, no sky-high compute costs, just your own personal AI sitting right there on your desk waiting for you. And here's the wild part. If you need to go bigger, you can connect two Dell Pro Max units together using the ConnectX 7 SmartNIC and 200GB networking to scale your workloads bigger. It's seamless, fast, and designed for serious AI development.

8:36Everything stays local and secure, whether you're building at the edge, handling sensitive data, or just trying to push the limits of what your models can do. Plus, it comes ready with the full NVIDIA AI software stack and DGXOS, so you get right to work. If you want cutting-edge AI performance and a compact form factor, check out the Dell Pro Max with NVIDIA GB10. It's the future, and it's already here. Check out the link in the description here to go get one today.

9:02Tim Davis:I feel like a lot of software, those famous sprays, I forget who says it, is like software is abstracting away the horrors of hardware, right? Have you heard that? Okay, yeah. I was just going to say, and it's an interesting observation, right? Because I think most programming languages today in the world, you know, they're very CPU specific, right? Like it's not like, you know, Python was never truly designed to go and run on a large scale, highly parallelizable, accelerated machine. That actually wasn't why it was created there. most programming languages today were very specific on CPU-related programming because most of the world's compute was CPU.

9:41As now everything's sort of turned over and we're going through this massive upheaval of infrastructure that's spinning towards, you know, parallelizable, acceleratable compute, you know, what is the compute model that supports that, right? And I think that's where what we realized at Google was you really need not only a programming model that jumps on a CPU and can do everything that you want on a CPU, but that can then also carry that over to an accelerated device and the most famous of this today is cuda but that is very locked to nvidia's systems you know and they've made a bunch of incredible uh you know progress and design decisions as to why they did that but then it doesn't help the the developer who's out there going but hang on a second i you know i don't want to just necessarily use nvidia i also don't want to use other types of silicon there's all this new hardware being invented there's all this hardware that exists right i won't use every gpu i can get my hands on yeah yeah but like you know and it's funny cory even in a mobile phone right like you think about it um one nvidia doesn't actually get you there but two in a mobile phone you have a cpu a dpu a digital signal signal processor a dsp and then an npu so your neural processing unit oh you have four pieces of silicon um and so again you like you know to really uh unleash the power of that to really enable people to say well what are these amazing applications that we could be building on mobile phones and other devices, wearables that are out there in the world.

11:02What is the compute model? What is the platform that enables us to do that? And there isn't really a good one today until, here's a plug, until we came around, I would claim. And that's what we're building.

11:14Tim Davis:Do you think that we are hardware constrained on the path to AGI or superintelligence, or are we software constrained, actually? So stepping back, what is intelligence? I would argue today the path we're on and i you know uh i would argue i don't necessarily know this is the right path but the path we're on today is let's keep training you know according to the scaling laws let's keep training bigger and bigger models right so i think grok was trained on 5e to the 26 flops um you know now there's discussion about let's go and train models on you know 10 to the 20 28 29 flops which which ends up being gigawatt facilities running model training cycles for months and months and months.

11:57It is an enormous scale. But fundamentally, what are these models? What are these LLMs? They're autoregressive models. They're taking basically history of text and then trying to predict the next token in many ways from left to right generation process.

12:14Tim Davis:And so is that actually intelligence? like are we are we actually uh you know creating you know uh intelligent machines when we do this i mean the perception i think to to humans is well clearly they're intelligent they're spinning out you know incredible uh solutions to problems but all of that has been trained on human data uh you know this incredible massive corpus of text that now has existed on the internet but is that actually the future of super intelligence, right? Like, I would argue that, you know, when you think about what is intelligence, and how do you define intelligence, you're sort of saying, you know, there's this agent or an entity that's taking actions, receiving feedback, trying to understand the environment that it's in against some goal, right?

13:03And, you know, and this is where I think it's hard to say that the future of intelligence is going to exist in the data center medium, trained on large-scale corpuses of human text that have existed in some ways in the past, right? The future of superintelligence, in my opinion, is this sort of reinforcement learning approach. And Richard Sutton actually just did a wonderful blog post and interview on this where you really look at reinforcement learning as an approach that needs to be scaled out into the world. Well, if it's out into the world, then what is the infrastructure, one, that's going to enable us to get out into the world?

13:45This process of continual training and continual learning. Software needs to power that. Hardware needs to power that. And then you get into like the physics of heat thresholds and thermal layers and, you know, all of these, you know, just even battery capacity out in the real world to actually have a constant training and learning process that occurs, you know, in reality, not in a data center. And so, you know, I sort of feel like that today we are working towards systems that have a very strong economic reward. There's no question that what has been built so far is incredibly valuable. We see this in code generation and many of the use cases that have become commercially valuable.

14:25But I don't necessarily know that training models and continuing just to build our gigawatt facilities to continue to have this perception of intelligence in regurgitating a lot of history of human text, particularly. and even I would argue even in video generation and image generation I don't believe the models understand what they're actually doing and and hilariously you know we actually don't understand

14:55Tim Davis:what these models are doing and I think you know Dario Amodi has written some some wonderful uh you know some wonderful uh essays on you know mechanistic interoperability and and understanding you know how do these models work but if you step back like as you know as engineers what other part of engineering as a society would we be comfortable with structuring and using products when we don't understand how they work like if you go to a structural engineer and you say is that bridge going to hold up you know would you be comfortable there's a probability that it will and we certainly hope i certainly believe that it should but i don't know if it will like that i did a demo where someone drove a car over it it didn't break so therefore you know everyone can drive on it I have a question for you, Tim.

15:44I hold concern there, I guess I would say. I do think perhaps we need a different approach. And I'm not in any way suggesting that I know what that approach is. But I often wonder, are we on the right vector? And, you know, humans have a nature of just trying to go to the simplest path from things. And sometimes I think we do need some more contrarian thinking in innovation, particularly in this space. I can't help but wonder if LLMs aren't like the CPU. If you think of AI as a computer, as we're bringing in memory, as we're bringing in tool use, as we're bringing in multimodal capabilities, as we're bringing world models into the picture, as they start training on things like the mountains and mountains of corporate data in the world that is sitting locked in databases right now.

16:33And even, you know, the possibility of what kind of data are you getting back from AI wearables, from electric cars, from what's Google Maps collecting? I mean, I feel like there's probably an interesting amount of data that would be really boring to read, but might do a lot in giving these things a little more of a perspective, almost a worldview. I'm not necessarily talking about consciousness. I just mean in terms of fleshing out the bits that were originally just an LLM and going from there. Do you feel like that is a possibility or do you think that's the wrong direction? You know, I think what is likely, and I know, you know, even back at Google, I think this was a thesis that we had, you know, back in 2017, 2018, was, look, at the end of the day, you know, distillation is now, you know, an incredible process.

17:31This was, you know, something that had been invented earlier on. You have these massive models and they have this incredible sort of history of human knowledge contained in them. but but again when you when you if you just take that and you put it out into the real world it's not like it has an understanding of you know an object if it's occluded that it continues around you know the entirety of the object right that's like when when when you know when children are first born they recognize physical concepts pretty quick you know they don't need to be trained on some historical corpus of data and so you ask like um you know is it like a cpu i think it has to be more powerful than a CPU.

18:10I think there has to be some generality to the infrastructure, no question. But if you look at where we are today in AI in some ways with GPU capacity and how fast it's been scaling, I sort of in some ways believe it's reminiscent of the first sort of foray into cloud infrastructure. There was one major provider at the time that was AWS and everyone, you know, first, you know, traditional enterprise was like, ah, we would never go to the cloud. You would be crazy to put your data in the cloud. You know, you fast forward a decade. Now it's like, not only do we want our data in the cloud, now we want to be multi-cloud.

18:48And so, you know, I think your analogy is interesting, but I think the actual inferencing process and the speed at which it has to happen and the latency at which, you know, meaningful results have to return to the user. and you know there's a quoted statistic that we used to use in google search was for every 100 milliseconds of latency it was one percent in revenue and i think that's being validated also by by amazon so in the real world like if you think about deploying um you know just on a cpu for example in a in an autonomous vehicle could you imagine flying down a freeway and, I don't know, a ball floats out onto the freeway.

19:29And because the CPU just is not as powerful as a large parallel just in terms of flops and processing, the car will have hit the ball and be another two miles down the road before you get an inference response that says, watch out for the ball, right?

19:44Tim Davis:Like imagine if it had to go back to open AI in the cloud and come back before it could react. And I think that's the practical reality, though, that any future system, any future form of intelligence needs to be able to have, you know, and I think we as humans have this sort of biological process that we've inherited over generations, but it is incredible that we are able to process so fast and respond so quickly to things. And I think computationally, you know, a CPU can't get us there. So then the question is in the real world, you know, I don't know if you've been in a Waymo. Not yet, but I hope to this month.

20:20yeah i mean that they are marvels in the sense of um uh you know they they and we can talk about all the different the sensors and the different approaches on those sensors and on those vehicles but at the end of the day you know you do need a degree of processing power against some battery capacity and thermal thresholds that enable intelligence to respond quickly because everyone i think is optimizing now for a high degree of perceived intelligence at the fastest possible outcome. And I think it's actually the nature of, you know, consumer behavior over time because of internet products that have been, you know, created over the last 20 years.

21:00And I think Google is obviously a big leader there, but that's not to say that Amazon and Meta and other organizations haven't done incredible work. But what we have done is train everyone to get everything instantly. And so, you know, even now, you know, we introduced thought reasoning on these large models like Chattapay and, you know, Deep Sequel obviously was a big provider of it. Um, but now it's funny that like, you sort of sit there seeing these things think and you're like, why is this taking so long? Um, and it's just an interesting observation that, you know, everyone now wants, wants things even faster.

21:31And I think that, um, that that's fine when you're, you're connected to a high, you know, particularly in the Western world, by the way, when we're connected to high speed bandwidth, fiber connections, um, and you can get information piped back to you very quickly. Yeah. The world becomes very different, though, if you've traveled to India or other places where the network quality is poor. I mean, I would argue even, you know, I'm a dual citizen of both Australia and the United States. And there are plenty of areas in both countries, frankly, where I feel like I'm back in dial-up, you know, the original internet sound, just trying to get some data onto my mobile phone.

22:05So, you know, there are a lot of, there are still a lot of, you know, physical challenges that exist in executing AI out in the real world.

22:12Tim Davis:Yeah. And I think, you know, this just goes back to the point, I think we do need the ability to have a unified compute model, because it's not just going to be one type of silicon that gets us true super intelligence. I think even if you believe in a construct where there is a large world model that is being trained on huge data center capacity that we're rolling out, you know, both in America and other countries around the world. I think at the end of the day, you want to distill that down to some core knowledge corpus. And that model may end up being a billion parameters, two, three, four, five billion parameters that resides on an edge device that is continually training and continually learning in the environment where you are.

22:50Because if there is anything true about technology is throughout its history, it always moves closer to the user. It always has. Now, the medium we have settled on, thanks to Steve Jobs and the wonderful folks at Apple, and then obviously that's proliferated across many other countries, is the phone. But is that going to always be the device that has the highest IO threshold for us? I don't know the answer to that. Same. But what I do know is...

23:16Tim Davis:I think empirically, no. We just don't know what the next thing is, in my opinion. Right. And it's so funny because it's the fickle nature of human beings, both from a perception of what's cool and what device do I actually want around me all day long, to what provides me a high enough degree of utility, that I'm actually willing to invest in this platform. And, you know, I was one of the, you know, I had Google Glass early on. And when it was very, you know, because I worked at Google on TensorFlow Lite. And, you know, it was just a weird, it was a weird social interaction to be wearing glasses and then sort of be like looking up to the sky when you're talking to someone.

Read the full transcript

23:58And it was sort of a little bit like, what exactly are you doing? why do you keep looking up at like off to the you know off to this region in space and then coming back and talking to me again clearly that's not going to work right like that's not a social medium that works and i i think it will be interesting to see even with the evolution of glasses as a as a product you know how do you get the consumer interaction how do you get the experience good enough that people want to believe that the utility is there um to have that and you Fundamentally, my point is to have that as an input device of what is happening around a person to improve AI.

24:34Because the more that we can get AI out into the real world, I think the more that it will actually be closer towards intelligence, it will be able to perceive, it will be able to understand complex environments and then respond to some goal. And I think that's what we're lacking in today's systems. Get it walking around in robots and checking out the world? Yeah. Yeah. I mean, I think that's, you know, I think that's an interesting, that's sort of, you know, in many ways, uh, uh, uh, an interesting frontier because you aren't in that, you know, are so limited in that world and you have to make constant trade-offs between if I tune and, and, and deploy a more powerful model on a robot, then I know I'm doing that at, you know, at, at risk of using more energy, which reduces the utility to the end state consumer.

25:22That's like, again i want to plug in my you know my iphone and get five days of battery i don't want to plug in my robot and then 25 minutes later it's like sorry i've run out of power i'll finish your laundry in four hours yeah 100 why did why did i pay you know fifty thousand dollars for this robot that that i need to charge every 10 minutes um like that's not going to fly and so there is going to be a an interesting balance there and a lot of it a lot of it you know not only is it hardware innovation and silicon innovation a lot of it is like what is the software model that's going to empower that to enable developers to move quickly and iterate fast yeah and so i think that is very

25:57Tim Davis:much as modular are you are you actually building like software to work on these edge devices and robots and all this sort of stuff yeah so how we you know i can talk a little bit about our platform so so how we started was we we recognized at google that one of the biggest challenges is that most most software that people interact with today is built on top of what hardware manufacturers build for their chips. So at the end of the day, you write a program, it needs to map to the silicon that it's executing on. But the challenge is every hardware manufacturer comes out with its own stack. So, you know, NVIDIA has CUDA, AMD has ROCAM.

26:33I mean, you could go through all the different silicon providers. So where we started was to say, is there a world where we could build an independent software stack that enables us to bring up hardware without needing to rely on the vendors, you know, libraries and infrastructure. And so that was sort of the big challenge. So one of the things we built, and I say this lovingly, that Chris Lightner, my co-founder, he doesn't need much incentive to build a new programming language. He loves, you know, I mean, this has been in many ways his life's work. But so what we realized, though, was the problem was, was there a way to build an abstraction where we could essentially write programs that became portable by nature.

27:14And there's a lot to that. I could go into more detail. But fundamentally, the concept is, could you write software, particularly AI software, very low-level operations that would enable you to go across different types of hardware easily? And so we had some infrastructure. Sort of? Yeah, except it's a brand new programming language. So we actually rely on some technology we built at Google called MLIR, which stands for, and this might go too low level, but it stands for multi-level intermediate representation. So it's about building a representation of a program that can go across different types of silicon.

27:48And there's methodology behind that, and I can spend time talking about it. But fundamentally, the goal was enabling our engineering teams to be able to write high performance kernels and low-level operations that could scale when AI models execute across different types of silicon. So we started building that and we also started building it in a way that was heterogeneous. So what I mean by that is we believed from our infrastructure work at Google that the future was going to be lots of hardware interacting and different types of hardware interacting with each other. Part of this was from, like I talked about, the mobile experience.

28:24You have four different types of architectures on a mobile phone. Ideally, you want them all working together. You want them all humming together. Well, again, if you don't have a programming model that can actually program all different four different types of accelerators, that becomes really hard. And so we said, well, let's create this new programming model that makes it easier to actually program, for example, a CPU and a GPU together. If we could do that, you know, that's like stage one to getting more utilization, more efficiency out of the hardware. And the big challenge in doing that was, cool, but could we meet the performance of someone like NVIDIA on their own silicon?

29:02Is that possible? Because if we couldn't achieve that, from a commercial standpoint, from a business standpoint, no one in the world is going to be like, wow, I love your idea, but wait, I lose dollar per token per watt. Can't be great and lossy. 100%. It's just not possible. So we started by doing that, and that's low level. It typically is for more advanced programmers. But, you know, we are sort of working on that programming model evolving. And then on top of that, and so you can think of Mojo and that programming language called Mojo, you can think of that as a comparator to, say, CUDA, which is a low-level programming model for programming.

29:39GPUs, the primary difference is we can go across any type of hardware, not just, you know, not just GPUs. Then above that, you know, one of the things we realized at Google, and this is sort of a long winded story is we realized we needed to build, you know, a new, essentially a new AI framework. And we call that Max. It is a modeling and serving framework. And so really this was carried through all of the lessons and the mistakes and the learnings that we took from, you know, building TensorFlow as a high performance, you know, AI framework. You know, TensorFlow now has sort of fallen off in terms of popularity and you have PyTorch and other things.

30:17But what's interesting and what we realized was, you know, I'll go back to one of the comments I said earlier. We realized we needed to build a framework that could realize this vision of what you train is what you serve. And back in 2017, 2018 at Google, we saw the inference training flip. We saw it all the way back then. And it's a funny story because when we first went out and started pitching to get capital for Modula, you know a lot of the investors that we pitched to we're like yeah no inference is the thing you know inference scales to the size of your user base training scales to the size of your research team i assure you inference is going to be the thing everyone wants to focus on and we were you know i won't name names but we were we were told at the time oh that's wrong everyone should be doing training you should be focused on training and we were like yeah i just i think training is going to become a lot smaller it will still be a very large workload but the number of people doing it will actually regress quite significantly.

31:09And look, you know, we obviously had a thesis and other people had their theses too, but, you know, that turned out to be somewhat correct. And so what we realized was, you know, today with PyTorch, if you build a PyTorch model, you can't actually deploy that at large scale in production. And that's why open software today, like VLM, which is an open serving framework, you know, it got created to essentially help fix the problem that PyTorch can't actually serve large-scale production workloads natively, right? When we created TensorFlow, it was the other way around. TensorFlow was wonderful at large-scale inference and not that great at, particularly from a usability standpoint, not that great at training.

31:51PyTorch was incredibly great at training, but not that great at inference, right? And so what we realized was, well, there's going to be an opportunity here to build a native framework, a new framework, and it's open source, where it starts with inference, and we can do large-scale inference, and we can do large-scale inference across compute types, and then we can add training actually easily enough. And the beauty of it is it has a serving component to it natively. So you can essentially run open models, all of the popular open models that you would otherwise be aware of. You can serve them.

32:26So now you can actually serve these things on, And, you know, increasingly, we actually last month announced Mac support. So running these things locally on devices like Mac machines, but then also taking that to very large data center workloads. Right. And so at that point, what we what we realized was, hey, look, we could we have an opportunity here to really take a lot of our knowledge in AI framework design and build a new framework. And then above that, increasingly, I will say that AI is now a full stack. You know, it's a full stack challenge. I certainly believe that inference is almost in many ways more complex than training at scale now.

33:03It involves, you know, increasing the Kubernetes, particularly the Kubernetes style layer in the stack. You have, you know, this idea now of splitting out pre-fill and decode components of the transformer architecture. You have different ways of managing caching and prompt caching and how you do that across all sorts of different models and all sorts of different hardware. And so what you end up needing is, you know, essentially a solution there too. And that, you know, today we call it mammoth, but fundamentally that's going to become a cloud, you know, part of our cloud offering. And for many developers, our goal is just, look, we can serve across hardware.

33:40This idea of, you know, talking about analogies, Corey, you know, we love to frame the unified compute model really as like a hypervisor for compute, right? If you could log into a platform and you could just say, here's my throughput, here's my latency, here's how much accuracy I'm willing to reduce through methods like quantization, and here's my cost target, just make it work. When I log into Snowflake, I don't execute a query and then go, oh, well, what CPU machine did Snowflake choose to execute my query? I actually don't know. Well, AI should be similar, certainly at least in the cloud to start, similar.

34:19You should be able to log into a platform, just give your requirements and expect the best TCO for your workload. Like the total cost should be highly optimized to whatever those parameters are and software should figure it out. It should be that simple for developers then to go off and build applications. And that's, you know, Grant, hopefully that answers the question of, you know, what are we building at Modular? That is very much what we're building.

34:45Tim Davis:So if I was a developer and I was building like a robot, right, for example, or working with robots, I could use modular to run if I was using like a specific GPU or whatever on my robot, I could use modular instead of CUDA. Yeah. So today, the open infrastructure we have is the framework. So Max and Mojo, which is the programming language. You can just go to our website today, download that, start playing with it. You can write operations for custom silicon. You can write your own machine learning models and scale them on robots if you want. All of that is available today. um we do work increasingly a lot of enterprise customers that come to us are still of the belief that like we want data center specific ai at massive scale um and so of course they're you know they are looking for uh for for cost gains and performance wins and so you know one of the uh the key customers you know i mean we've we're working with a lot of customers now but but one of the you know the the sort of testimonial stories we worked with a uh you know an awesome company called inworld they're like an advanced yeah um originally gaming but now doing like an ai runtime so it sort of metricizes and and can do runtime environment for you in your in your applications and they came to us and said hey look um you know very advanced team they're all x x deep mind um and you know they said look we're using all this open infrastructure of llm and other things but we need like here's our use case and again to reiterate great, this is how people approach what they're trying to solve.

36:19We have a text-to-speech model, and we want to get the first two-second audio chunk back to a user in under 200 milliseconds is what we want to be able to achieve. We've played with their model even. Yeah, we know exactly what they're trying to solve for. And so the challenge they had was as you scale throughput, the latency explodes. So the problem is as more and more users come onto their system, those concurrent requests end up, each one ends up getting higher and higher on the latency threshold. And so, you know, it's really interesting. There's a couple other customers that I can't name yet, but there was this really interesting user observation in that, right?

36:59And what the user observation is, irrespective of how intelligent a model is, if the latency threshold goes up, you know, an everyday person interacting with AI thinks the model is dumber.

37:12Tim Davis:Yeah, I was thinking that was what it was. And so what that means is that as a consumer of technology, you're sitting there and you're like, if it's a text-to-speech model, you're saying maybe it's a nursing agent or a healthcare agent or there's something going on, right? And you say, it rings you up or it does whatever and it says, hey, how are you feeling today? And if you say, oh, well, I'm not actually feeling that great. and then 10 seconds later the model goes oh well that's no good you're like am i this the person like what what is going on with the person that that it feels very like that that every second delay with a speech model like really makes a difference when you're on the consumer end especially that's where it's really really noticeable it's like yeah please sit here while I go retrieve your information.

38:05I'll be back. Yeah. Yeah.

38:07Tim Davis:If I have to wait for ChadGBT to talk back to me, I'm like, I'm just going to close this. Like, it's not. They've come a long way. I'd rather read it, you know. I will say they've come a long way. I spend a lot of time working with those models. And even ChadGBT's, Grok's, Google's, even Claude's, they were all fast, which is a thing that's been impressive. Because when it first started, it was not that way. When they first started releasing those models, it was like. And it was mind blowing, even though it was 10 seconds. Yeah. And I think the challenge is like you get these awkward, like you sort of go to speak over it and then it sort of tries to stop.

38:43And then you're like, oh, you know, even in our conversation and, you know, I apologize in any way if I interrupt you guys. But if, you know, you go to talk and then I go to talk, I'm like, oh, no, Grant, you go. Oh, no, Corey, you go. And so there's this weird back and forth interaction.

38:58Tim Davis:And of course, there's the latency of the human brain. And then there's the latency of like, we're doing this over streaming. So there's like this, the latency of like the streaming. Yeah. The internet's still a factor. Yeah. Yeah. And so it's wonderful to hear those stories because at the end of the day, like, you know, infrastructure is a tool and, and you want to understand, but, but how is this actually impacting, you know, what you're trying to build and, and, and really understand what is the application? What is the end state? And so for them using our infrastructure on you know, they, they were deploying on a black well machines, which is, you know, incredible, incredible piece of silicon from NVIDIA.

39:29So we were actually able, you know, for them to 4X the performance over VLM. Wow. And ultimately, you know, reduce their cost base by, I think it was 60 or 70%. And so, you know, why does that matter? That matters in a few ways. One, penetration of technology, you know, fundamentally, if you can lower the cost threshold, then, you know, it becomes a lot more, it's easier for people to try it, adopt it. And I think they're now 20x cheaper than their competitors in the TTS space. But more importantly, it begins to map and make those interactions more real. And this is not to say that a TTS model is in any way intelligent and understands what it's doing.

40:17But the use cases become more powerful because I still strongly believe there's some incredible companies working on things like preventative medicine and being able to call American citizens just to check on them and to be able to say things like, hey, look, did you take your pills today? Are you feeling okay? And if those responses can become more intelligent and we can help people in their everyday life by deploying that type of technology in a real-time fashion at scale, I think that makes coming to work every day worth it.

40:48Tim Davis:I want to kind of like zoom out a little bit here because we have seen all of these really incredible deals announced recently between OpenAI and like every chip maker that I know of. They've got deals with NVIDIA, they've got deals with AMD, they've got deals with Prodcom. And NVIDIA got a deal with Intel in the middle of that. Yeah, like there's all of this stuff going on. And NVIDIA is, from my understanding, a partner with you as well. Is that correct? Yeah, they are. I mean, we're trying to partner with all the silicon providers. I think at the end of the day, what we're really trying to help unlock is, it would be wonderful to see a lot more competition in silicon fundamentally.

41:24And if software is what is restricting us, I just think there's so many problems to solve in the world and so much innovation that can happen there. And so our belief is by having a unified, independent compute platform, we can help other new startups in the chip space accelerate faster and sort of adopt our infrastructure and move quicker to meet customers' needs.

41:48Tim Davis:because my understanding is that nvidia's cuda is like kind of like the reason like yes their chips are the best right now but also cuda is really what keeps people locked in to work yeah and you know it's you can look at the history of it i mean it's now a 17 year old platform but but i would certainly say many of the decisions that were made at companies like meta and google to build these ai frameworks on top of cuda as a you know as a programming model for the for NVIDIA's acceleratable compute really acted as a massive distribution mechanism for them to get very, very mass scale, right? Like it accelerated the adoption and penetration of execution of models on NVIDIA's hardware.

42:33And if you look at other folks, like at the time those decisions were made in frameworks, you know, like TensorFlow and PyTorch and others, it was the best thing around. And so people went, well, let's just use that. But by doing that, you know, they obviously changed the, you know, the trajectory, certainly for NVIDIA. And I'm sure many, many happy stock holders at that company, but also, you know, more broadly on how other, you know, other companies can come and compete. Right. And so I think a lot of what we, you know, we fundamentally believe is not only is a unified compute model good for penetration of applications, both at the data center in the edge it's also good just for competition it makes it makes it possible for more people to compete uh and that you know there'd be more innovation um at the silicon side of it but grant you also asked you know all of these deals and and uh and all of this sort of circular money flow right you know i think there's been parallels drawn and um there's a guy uh tomas tongues uh who's a who's a former uh uh vc and he has a blog he has a wonderful blog post on this sort of comparing it back to you know when we were laying fiber back in you know in 2000s and i i think the difference of course now is i don't you know i i think that gpu utilization around the world right now most gpus uh you know are probably running reasonably reasonably hot um given that the amount of ai that's being consumed you know i think it is a very different time to uh to what uh to what was happening when we you know certainly in america when we were laying fiber and it just wasn't being utilized yet.

44:11Right. And then of course, across the arc of 20 years, that's changed. But I think, you know, you need to look at, there is, there is significant revenue consolidation on, on who is paying. You always look at where is the money coming from, right? At the end of the day, the money is coming from, you know, essentially the, the magnificent seven technology companies, which represent roughly, I believe 30 % of the S &P 500 in terms of market capitalization. Right. So, you know, if they're the ones buying all the compute, you know, you would have to look at their businesses in the long term and say, is that sustainable?

44:46I think the challenge, of course, is at least at that level, they're also building all their own chips. And so, you know, in many ways, there's this purchasing of additional capacity from companies like, you know, OpenAI. But then, you know, it's been interesting to see in the media that OpenAI is also now trying to build its own silicon. and traditionally that's what you see right like you see and i think you know this was true even at google you see you start with sort of uh models and you train models and then what you do is you go up the stack and you go up into the cloud and you say well now i need to get mass distribution for my model and prove that this is an you know one a compelling uh set of technology that that has vast user adoption and that is economically viable and then you know you prove that there's larger revenue generation and then very quickly you're like now i need to go down because i need to figure out how do i get better costs you know cost optimization across the silicon right and so i think what you and obviously open ai has been the most famous on this you could you could compare it to anthropic with with trainium and and their relationship with with aws in a similar way but you know a lot of it is just this constant belief and i i think that you know i i do personally think that we are right to challenge it, that is the future of superintelligence just continuing to scale these models larger and larger and larger and larger?

46:09Is that actually getting us on the path to some form of superintelligence? Or is it not? Because we are spending an enormous amount of capital on the belief that it is. And there may be utility there. I'm not suggesting that there is not utility in that investment. But, you know, I do agree, you know, deeply with Jan LeCun and Richard Sutton and others that have made the case for, look, auto-regressive LLMs is not actually the path to superintelligence. We need a different form of innovation. Like I was at Google when Noam and Aiden and others created the transformer and intention is not enough and, you know, changed the world.

46:49But, you know, are we convinced that that's the final architecture and that, you know, that is reflective of what we see as intelligence. You know, I personally am not convinced. And I think it would be great to be able to see a different approach. And so the sloshing of all the capital is highly, you know, is predicated on this future of just keep scaling the flops and keep building the data centers. And maybe that's a viable path, but I do think it's right to question it.

47:17Tim Davis:I wonder though, if it's not like, you know, because eventually you'll have to scale whatever the current architecture is in theory or sorry well you'll have to scale whatever the architecture is that works um if you're following the bitter the bitter lesson which for people who forget what that is is basically like if you have uh you know basically reinforcement learning scaled properly is like just all you need essentially that's a very yeah search and learning uh really the primary techniques of of being out of scale the future of of uh of these models sorry cory you were about oh yeah i said you know i even feel like something The reason the jury is still out for me is because I feel like what we're watching right now is this certain amount of concentration in smaller models.

47:59However, that's largely fed by this idea that we're going to scale this up bigger and then we're going to work to make it more efficient. We're going to work to streamline it. We're going to make it faster. We're going to make it smaller. And this idea that, you know, smaller models are getting better. And there's a big focus right now, at least it seems like to me, on better data, cleaning data, training models on increasingly better data. So I don't know, does scaling have to continue to continue being able to push these smaller models up the food chain as well? I don't know, but it's sure fascinating and a lot of fun to talk about.

48:37Yeah, and I think, you know, even if you look at humans, right? Like, I'm in no way a neurologist, but I do know that, you know, we don't execute the vast majority of our brain function when we operate, you know, in the real world. And so, you know, again, this comes back a little bit to the idea of, you know, interpretability and actually understanding why do models make certain decisions at certain points, you know, when different neurons are firing inside these model topologies. We don't understand that yet. So for all we know, much like the human brain, the vast majority of these models, you know, these models weights, which represent its pre-trained, essentially its pre-trained in-weight memory.

49:19Maybe it's just wasted. Maybe we're, you know, it was always interesting when sparsity came out and pruning and quantization techniques came out and you could chop large layers of models out. And fundamentally, it still worked very well. Yeah, and you're still at like 97 % or something. Yeah, there's a loss, but it's minimal. So, you know, are they bread columns towards, well, maybe there's a different approach here that's needed. And, you know, I think to your point, Grant, yes, you know, maybe in one school of thought, I think it's correct to say that we could keep training. But maybe in a different school of thought, you know, I think we can certainly question, well, is there a world where you have a very small but, you know, to your point, Corey, high quality model that has a incredible rich architecture that actually answers the vast majority of what we needed to answer in how we perceive intelligence to be able to scale and navigate and understand complex environments to achieve some goal.

50:19You know, I still, I'm still on that increasingly leaning to that side of the fence and I don't necessarily know how we get there. But I don't think we, you know, I think that we don't know how to get there because yet, you know, I still can't sit here and say, we even understand how these models work. So I think that - Or our brains for that matter in a lot of ways. That's right. And so I think, you know, I would love to see, and I say this in an article I wrote, I would love to see a much larger percentage of capital flow from, or certainly, you know, OPEX flow in these large frontier labs towards, you know, interoperability and research on on interpretability because I think, you know, just the basics of engineering, it should be that we understand like we're deploying a technology at mass scale and yet we don't understand that technology.

51:07And I just think when you apply that to any other construct in engineering, you know, I just think people would fall over and say, oh, that's crazy. Why would we do that? We're not going to go and build a bridge and just hope, you know, we're going to send a whole bunch of traffic over it and hope that it stays up. You know, I think most humans' perceptions of risk would be, well, I'm not driving over that bridge. And yet we're rolling out and penetrating this degree of infrastructure in society without clear answers. And I think we owe it as an industry, but we owe it as in many ways we're creating, we have the ability as a species to be creating our future incarnation of what's next.

51:47And I think we owe it to ourselves to deeply understand that.

51:49Tim Davis:So five years from now, if modular succeeds at its mission, what do you think the AI infrastructure landscape looks like? You know, with the caveat that we know that there's going to be this big kind of like shift where it's like either scaling works or we have to come up with something different. So we've established that that's a paradigm. Or we're still asking the same question five years from now. True. It could be unanswered, I guess, by then. But what does modular look like? Yeah, I think there's a couple of ways to answer that, right? We certainly hope that the open infrastructure, so Mojo and Max, gets penetrated across the development community and development ecosystem very broadly.

52:26We think that it is open source. We would love the community to rally around it. It gives people more optionality. I do see it as the future, again, in my opinion, as I understand it right now, is a lot more contingent on reinforcement learning and really continuous learning. and I think to achieve that we need software that enables us to one really take advantage of all the compute in the world I don't believe we're doing that today as crazy you know that may sound to some people we need to be able to fully saturate and utilize the compute we believe we're building the framework to do that and we would love people to adopt it and be able to realize you know go and build incredible applications and invent the future with it so that's one thing I you know I definitely think we want to see I think the second thing we want to be able to see is you As a company, we have an absolutely need to scale and are doing this now very significantly, just the commercial side of our organization.

53:20Our vision for Modular is to remain an independent company. And to do that, you need an enormous economic engine. And I think the currency of the AI industry is essentially dollar per token. We believe that we can provide efficiency gains there better than anyone else. and we believe that now a decade helping build our infrastructure, of which I will both apologize for because a lot of it has been messy. We've taken those lessons and really tried to invest in fixing those mistakes in building a new platform. I think for more, certainly at least today, and Larry Page when we were at Google had a famous saying of, develop the most sophisticated developer today because that will be everyone tomorrow.

54:05And today that's us. We are very much in the advanced user cohort, but five years from now, I would hope that that diffuses more broadly across enterprises and it's just easier and they have more optionality and they're able to do more with AI than ever before, both in the data center and at the edge. And the last thing, certainly that we would hope for is that there is more diversity in silicon. I think, I've always believed markets are enormously power or distributed. You know, I think you can look at any of the Magnificent Seven today and you could argue in their own individual right, all of them are monopolies.

54:43And I know that, you know, that's probably a scary word to say for some people, but the reality of the market is there is enormous market power in each of those verticals. And so, you know, being able to have more innovation down to the silicon level and see entrepreneurs go back to, you know, the raw atoms and physics of being able to move bits around more efficiently and be able to build silicon and tape out silicon and innovate in that space would be something that's incredible. And I think, you know, what we're hoping in our platform, particularly as, you know, more developers adopt it, is that then it's easier for hardware manufacturers to come on board and use it.

55:24And the analogy we try to apply there a little bit is similar to Android. Like Android did a wonderful job of being an open ecosystem. They helped unify a lot of the, you know, handset manufacturers to be able to have an operating system that was a different approach to Apple. but, you know, enabled them to compete in an ecosystem that then, you know, helped really now cover half the planet in terms of phone distribution. You know, we hope that there's a software layer there that makes them being able to plug into and compete quickly out of the box that, you know, that drives more innovation there too.

55:57And really what that achieves at the end is, you know, the best product should win. I always believe that, you know, if you can enable, if you can build tools and you can build infrastructure that act as a great equalizer and at the end it drives more innovation across, you know, across technology, that's a net positive for the world. And lastly, I'll say, I certainly hope that super intelligence is created on the modular platform. I mean, we would. Of course. Yeah, of course. I strongly subscribe to Jeff Bezos's regret minimization framework. And I really do believe that you can look back on your life.

56:41And I think if AI has really helped to penetrate a society and generally make it better, which I strongly believe that it will. And there's a lot of risks and there's a lot of challenges, but that's true with any new diffusion of technology as it penetrates society. and so I think our vision there is if we can help leave that small dint on the world and enable people to adopt this technology in a way that's beneficial, that would be an amazing life to be able to look back on. We still have a long way to go though. I would say that. What's the easiest way for someone to try out Modular without A, committing to a massive infrastructure migration?

57:23Yeah, it's free to try. just go to modular.com. You can click getting started, great tutorials and getting started guides. We support an enormous number of open source models. We are bringing up sort of the endpoint cloud solution that will enable developers just to start serving models really easily out of the box. But today, right now, it's been more a focus on, you know, you can download the infrastructure and play around with it yourself. We have a bunch of instruction markdown that enables you to easily punch out some code with, you know, with the help of those tools. And so We're a startup.

57:52We're like a hummingbird. We're extraordinarily responsive. And you can join a forum or a Discord or just ping us on social media and we respond as fast as we can. And we believe that we're building against what people are asking us to build. And so for us, it's just come and join the mission and help us change the world.

58:13Tim Davis:Do you think that it's easier to build on Modular now that it would be on PyTorch if you're a total noob? I think now, you know, we've done a lot of the legwork for serving. So we are still an inference platform. I think PyTorch is still wonderful as a training platform. And the team at Meta have done a really incredible job there. I believe our vision, though, is actually in the long term to make people's lives easier. Like when you deploy a container today with VLM in it, you know, you have to download one for NVIDIA and it's about 10 gigabytes. If you download it for Rock 'em, which is the AMD platform, I think it's like 20 gigabytes.

58:46But we have a single container that can execute across both those platforms now that's under a gigabyte. And the reason for that is when you rebuild something, you scoop out all the dependencies and you basically then reduce the overall footprint. And why that matters is when you think of cold start time and you're deploying on Kubernetes instances and you need to scale these containers all over the place, every bit counts. It really does. and so you know we our goal is um is to get there grant there's no question the only way we're going to get there is with a you know motivated uh group of of people but you know we you know we do have a lot of knowledge that i i think we want to impart on our platform too yeah you know at the end of the day any platform is a function of its community and and it's and the people that choose to invest in it and so you know we're very very conscious of that i i think it's something we're committed to investing in.

59:41So I would challenge folks out there if they're excited about AI infrastructure, come and give it a go. If you run into some challenges on what we're doing, be loud about it and we will help. We will respond as quickly as we can. That's the stuff you need to know, right? Exactly. Yeah. Well, Tim, thank you so much. This has been a great conversation. I really appreciate that you took the time to join us. No, it was lovely. Thanks so much, Corey and Grant. This was a great discussion. I appreciate And I apologize. We traversed the map, I feel like, a little bit in the surface area of what we touched on.

1:00:16But yeah, I really appreciate the discussion. And yeah, and thanks for having me on the show. Well, thanks so much to everyone watching. We really appreciate that you took the time to join us. Please like and subscribe to the channel so we can continue to bring you more great guests like Tim. And overall, thanks for watching. If you haven't yet, make sure to pop by the Neuron.ai and sign up for the Neuron.ai newsletter. Join about 600 ,000 others who get it every morning. And until next time, farewell humans.

From the publisher

While everyone obsesses over which AI model is smartest, a quiet revolution is happening in the infrastructure layer underneath. Modular just raised $250M at a $1.6B valuation to solve a problem most people don't know exists: AI is locked into expensive, vendor-specific hardware ecosystems. Tim Davis, Co-Founder & President of Modular, joins us to explain why his company is building the "hypervisor for AI"—making it possible to write code once and run it on any GPU, from NVIDIA to AMD to Apple Silicon. We dive into why this matters for businesses, what the Android analogy really means, how companies are seeing 70-80% cost reductions, and whether we're even on the right path to superintelligence.


Subscribe to The Neuron newsletter: https://theneuron.ai

Try Modular: https://modular.com

Getting Started Guide: https://modular.com/get-started

More from The Neuron: AI Explained

All 106 episodes
The "Android Moment" for AI Infrastructure: Why Modular Just Raised $250MThe Neuron: AI Explained · 1 h 1 min
Listen in VO