In short
Latent Space Podcast Episode Summary
Episode Information
- Podcast Title: Latent Space: The AI Engineer Podcast
- Episode Title: The Shape of Compute
- Guest: Chris Lattner of Modular
- Duration: 1 hour, 17 minutes
- Release Date: Not specified
Episode Overview This episode features Chris Lattner from Modular, who discusses the company's innovative approaches to GPU programming and their latest projects, including their programming language Mojo and the MAX inference framework. Lattner shares insights into Modular's differentiation from competitors, the business model, and his personal reflections on leadership and productivity.
Key Topics Discussed
- Introduction to Modular
- Modular is positioned as a company focused on breaking NVIDIA’s CUDA monopoly and enhancing GPU programming.
- Lattner emphasizes the need for heterogeneous compute and simplifying GPU programming complexities.
- R&D and Product Development
- Initial Phase: Modular started in a research phase for three years focused on building state-of-the-art performance without using CUDA.
- Transition to Engineering: After proving their technology's capabilities, the company shifted to engineering and product launches, resulting in significant product releases every six weeks.
- Key Technologies
- MAX Framework: This is Modular’s inference framework, designed for performance and control, particularly for generative AI applications.
- Mojo Programming Language: A new programming language that simplifies GPU programming while enhancing performance, especially for high-performance applications. Mojo is designed to be user-friendly while retaining powerful capabilities.
- Open Source Contributions
- Modular has open-sourced their MAX framework and Mojo language to foster community involvement and contributions.
- The company emphasizes the importance of community engagement and open-source contributions for growth and innovation.
- Differentiation
- Modular distinguishes itself from competitors (e.g., VLLM, SGLang) by focusing on a small number of tasks executed exceptionally well, rather than overextending across many functionalities.
- Lattner discusses the value of composable systems that can evolve and adapt over time to meet changing demands in AI technologies.
- Business Model
- Modular offers its MAX framework and Mojo programming language for free for individual use, tapping into the community for broader adoption.
- The company plans to monetize enterprise solutions, providing GPU management and support to organizations.
- Challenges and Reflections as a Founder
- Lattner shares his experiences and challenges in leading a startup, particularly in scaling a team and maintaining morale during complex development phases.
- He emphasizes the importance of clear communication and setting realistic expectations within the team.
- Daily Routine and Productivity
- Lattner provides a glimpse into his daily routine, emphasizing the importance of work-life balance, exercise, and strategic planning to manage his responsibilities as a founder.
- Future Aspirations
- Lattner expresses excitement about the future of AI and the potential for more people to engage with GPU programming through Modular’s tools and initiatives.
- He hopes for wider adoption of their technologies and anticipates future collaborations with prominent AI projects.
Key Takeaways
- Innovation in GPU Programming: Modular is at the forefront of redefining GPU programming without relying on NVIDIA’s CUDA.
- Focus on Community: Open-source contributions are pivotal for Modular’s growth strategy and community engagement.
- Emphasis on Usability: Mojo aims to make high-performance computing accessible to a broader audience, reducing the complexity typically associated with GPU programming.
- Leadership Insights: Lattner emphasizes the personal nature of startup culture and the importance of resilience in navigating challenges.
Conclusion This episode offers valuable insights into Modular's innovative technologies and approach to GPU programming. Chris Lattner’s experience and forward-thinking mindset reflect the rapidly evolving landscape of AI and computing. The discussion highlights the importance of community, open-source collaboration, and a user-centric approach in shaping the future of AI technologies.
For more detailed notes and resources, visit [Latent Space](https://latent.space).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:05Hey, everyone. Welcome to the Latent Space Podcast. This is Alessio, partner and CTO at Decibel, and I'm joined by my co-host, Swix, Fanderos Mall.ai. And we're so excited to be back in the studio with Chris Latner from Mojo Modular. Welcome back. Yeah, I'm super excited to be here. As you know, I'm a huge fan of both of you and also the podcast and a lot of what's going on in the industry. So thank you for having me. Thanks for keeping us involved in your journey. And I think you spend a lot of time writing when obviously your time is super valuable, educating the rest of the world. I just saw a two and a half hour workshop with the GPU mode guys and that was super exciting.
0:40We're decked out in your swag. I have the socks. Amazing. I love it. We'll get to the part where you are a personal human machine of just productivity and you do so much. And I think there's a lot to learn for people just on a personal level. But I think a lot of people are going to be here just for the state of modular. We're also calling it the shape of compute, I think is going to be probably the podcast title. Yeah, it's super exciting. I mean, there's so much going on in the industry with hardware and software and just innovation everywhere. Most people can catch up on the first episode that we did and we introduced Modular.
1:14I think people want to know, I think since then you sort of open sourced it, there's been a lot of updates. What would you highlight as the past year or so of updates? Yeah, so if you zoom out and say, what is Modular? We're a company that's just over three years old. Three and a quarter, so we're quite a ways in. The first three years was a very mysterious time to a lot of people because I didn't want people to really use our stuff. okay and so why why that why do you build something that you don't want people to use well it's because we're figuring it all out and so the way i explain is we were very much in a research phase and so we were trying to solve this really hard problem of how do we unlock heterogeneous compute how do we enable gp programming to be way easier how do we enable innovation that's full stack across the ai stack by making it way simpler and driving out complexity like there are these core questions and I had a lot of hypotheses, right?
2:05But building an alternate AI stack that is not as good as the existing one isn't very useful because people will always evaluate it against state-of-the-art. What I did and what the team did and what we all did together is we said, okay, well, let's get to the point where at least Chris is happy and I have very high standards and so we need to be state-of-the-art on NVIDIA GPUs, meeting NVIDIA's best on things like a Lama 3 model, which by the way is serving end to end like very high bar by the way this is like something that's a pretty well studied problem like it's not NVIDIA that's working on this it's the entire industry that's doing this exactly what is data with the art just as a rough I don't remember the number of tokens per second it's tokens per second and it's roughly between 800 to 1000 shared GPT benchmark and like industry standard kinds of things and so I said to the team like look until we can do that I don't believe it's real So, yeah, it turns out that, yeah, there's all these things.
3:02There's page attention, continuous batching. There's GPU curls. There's programming languages. There's a whole bunch of hardware stuff. And there's all this stuff. And I said, oh, by the way, we can't use CUDA. That's the whole point. That's the whole point, right? And so it's not a, let's pick the shortest, easiest path to get to a demo. This is a, let's do the hardest, most fundamental thing that everybody is telling me, as usual, that it's impossible. This can't be done, right? And so for those first three years, right, the challenge is prove that we can do something that people think is impossible and at least prove to me.
3:34And so what do you do for that? What you do is you clear the deck. You try to get enough distractions out of the way so you can focus, you can iterate, you can move quickly. You don't have to argue with committees. And so we want to be open to a certain extent because we want people to be aware of like Mojo and the things we were doing before so we can hire people. And there's very specific reasons. We don't really want design by committee. We don't want a lot of these other things. And so across that research mode, it's like the primary thing I care about is prove that it's possible and make me happy.
4:01And we transitioned right at the end of December. We had this release where we achieved that goal. And it was a super narrow release. It ran just on A100, just one model. But it had state-of-the-art performance. It was a full-stack vertically integrated thing. And my whole team is telling me, like, oh, my God, it's like they pass out and could take time off for Christmas. but then they come back and it's like, okay, well, everything sucks. It's not very good. Like we got to work. They changed their minds after they went over. No, no, it worked. We achieved the goal, but it's not very good. The code is ugly.
4:34There's technical debt. There's this and that and the other thing. It only runs an A100. It's only one model. This isn't useful. This isn't like a valuable contribution. And I built some things in the past, right? And so I said, well, that's okay. We have one thing that works end to end and we know 400 ways to make it better. Yeah. And so instant mode switch. And at that point you say, okay, well, let's just break it down like a normal engineering problem. It's not R and D anymore. It's an engineering problem. You say, okay, cool. Let's refactor these APIs. Let's deprecate this thing. Let's, oh yeah, let's add H100 support.
5:05Let's add function calling and token sampling and like all the different things you need. You can project manage that. Right. And so every six weeks we've been shipping a new release. And so we added all the function calling features and now you have agentic workflows. We have 500 models. We have H100 support. We're about to launch our AMD, MI300 and 325 support. That'll be a big deal for the industry. And as you do that in Blackwell, like all this stuff is like all now in the product. And so as this happens, now suddenly it's like, oh, okay, I get it. But this is a very fundamentally different phase for us because it works, right?
5:39And once it works, you can see it end to end. Lots of people can put the pieces together in their brain, it's not just in my brain, and a few other people that understand how all the different individual pieces work, now everybody can see it. So as we phase shift, now suddenly it's like, yeah, okay, let's open source it. Well, now we want more people involved. Now let's do each of these things. It's actually, we had a hackathon. And so we invited a hundred people to come visit and spend the day with us. And we learned GPU programming from scratch. And so we built a very fancy inference framer.
6:10The winning team for the hackathon took that four person team and one day they had not used mojo before they hadn't programmed gpus before and they had built a training system they wrote an atom optimizer a bunch of training kernels they built a simple backprop system and they actually showed that you could train a model using all the stuff we'd built for inference because it's so hackable and also because the ai coding tools are awesome but uh but this is the power of what you can do when when you're ready to scale now if we had done that six months ago or 12 months ago or something it would have been a huge mess, right?
6:41Because everything would break and there are a lot of bugs. And honestly, today it's still an early state system. There's still some bugs, but now it's useful. It can solve real world problems. And so that's the difference. And that's kind of the evolution that we've gone through as a team. I remember when we first had you, at the start, you were focused on CPU, actually optimization. How long did you work on CPU? And then how much of a jump was it to go from CPU to GPU? So the way I explain modular is if you take that first three-year R &D journey. And if you round a little bit, first year was prove compilation philosophy.
7:16And so this was writing very abstract compiler stuff and then prove that we can make a matrix multiplication go faster than Intel MKL's matrix multiplication on Intel Silicon and make it configurable and multiple D types and prove like a very narrow problem and do that by writing this MLIR compiler representation directly by hand, which was really horrible, but we proved the technology. It's a technology milestone. Year two was then say, okay, cool. I believe the fundamental approach can work, but guess what? Usability is terrible. Writing internal compiler stuff by hand sucks. And also, Matmol is a long ways from an AI framework.
7:55And so year two embarked on two paths. One is Mojo, so programming language, syntax, member of the Python family, make it much more accessible and easy to write, kernels and performance and all that kind of stuff. And then second, build an AI framework for CPUs, as you say, where you could go beat OpenVINO and things like this on Intel CPUs. End of year two, we got and we're like, foof, we've achieved this amazing thing. But you know what's cool? GPUs. And so again, two things. We said, okay, well, let's prove we could do GPUs, number one. Also, let's not just build, you know, air quotes, a CUDA replacement.
8:28Let's actually show that we could do something useful. Let's take on LLM serving. No big deal, right? And so again, two things where you say, let's go prove that we can do a thing, but then validate it against a really hard benchmark. And then that's what brought us to year three. So each of these stages is really hard and lots of interesting technical problems. Probably the biggest problem that you face is you face people that are constantly telling you it's impossible. But again, you just have to be a little bit stubborn and believe in yourself and work hard and stay focused on milestones. When they say impossible, do they mean impossible or very, very hard?
9:00Well, so I mean, it's common sense that CUDA is nearly 20 years old and Vida's got hundreds or thousands of people working on it. The entire world's been writing CUDA code for many years. It's very, very hard. And so it's, no, I mean, many people think it's impossible for a startup to do anything in the space. Like that's just common sense. All these people have thrown all this money at all these different systems. They've all failed. Why is your thing going to succeed when all these other things built by other smart people have failed? Right. And so it's conventional wisdom that change is impossible.
9:29But hey, we're an AI. You know this. Like changes all around us all the time, right? And so what you need to do is you need to map out what are the success criteria, what causes change to actually work. And across my career, like with LLVM, all the GCC people told me it was impossible. Like LLVM will fail because GCC is 20 years old and it's had hundreds of people working on it and blah, blah, blah, and spec benchmarks and whatever. Nobody told me it was impossible outside because it was secret. So that was a little bit different. But everybody inside Apple that knew about it said, no, no, no, Objective-C is fine.
9:58We should just improve Objective-C. The world doesn't need new programming languages. New programming languages never get adopted. And it's common sense that new programming languages don't go anywhere. That's conventional wisdom. You know, MLIR, super funny. MLIR is another compiler thing. And so I built this thing, brought it to the LLVM community. I said, hey, we open source this. Does LLVM want it? I know a few LLVM people, right? It was my PhD project. And all the LLVM Illuminati in the community had been working on LLVM for 15 years or something. They're like, no, no, Elvium's good enough.
10:28We don't need a new thing. Machine learning is not that important. And so, again, obviously have developed a few skills to work through this kind of challenge, but you get back to the reality that humans don't like change. And when you do have change, it takes time to diffuse into the ecosystem for people to process it. And this is where we talk about hackathon. Well, you kind of have to teach people about new things. And so this is why it's really important to take time to do that. And the blog post series and things like this are all kind of part of this education campaign because none of this stuff is actually impossible.
11:05It's just really hard work. It requires an elite team and good vision and guidance and stuff like this. But it's understandably conventional wisdom that it's impossible to do this. And it's the idea to build serving. Basically, I don't need to tell you how much better this is. I can just actually serve the models using our platform. And then you can see how much faster it is. and then obviously you'll adopt it. Yeah, well, so today you can download Max. It's available for free. You can scale it to thousands of GPUs you want for free. I'll tell you some cool things about it. So it's not as good as something like VLLM because it's missing some features and it only supports NVIDIA and AMD hardware, for example.
11:40But by the way, the container's like a gigabyte. Wow, right? Well, why is it a gigabyte? Well, it's a completely new stack. It doesn't use CUDA. You can run arbitrary PyTorch models. And if that's the case, then okay, you pull in some dependencies. but if you're running the common LLMs and the GNI models that people really care about, well, guess what? It's a completely native stack. It's super efficient, doesn't have Python in the loop for eager mode op dispatch and stuff like this. Because you don't have all that dependency, guess what? Your server starts up really fast. So if you care about horizontal autoscaling, that's actually pretty cool.
12:12If you care about reliability, it's pretty cool that you don't have all these weird things that are stacked in there. If you're wanting to do something slightly custom, guess what? You have full control over everything. and stuff's all open source and so you can go hack it. This thing's more open source than VLLM. Because VLLM depends on all these crazy binary CUDA kernels and stuff like this that are just opaque blobs from NVIDIA. And so this is a very different world. And so I don't want everybody just like overnight drop VLLM because I think it's a great project. But I think there's some interesting things here and there's specific reasons that it's interesting to certain people.
12:47Can you just maybe run people through the pieces of Macs? Because I think last time none of this, existed. Yeah, exactly. Last time has been a long time ago, both in AI and also in modular time. Yeah, so the bottom of our stack, we think about it as concentric circles. So the inside is a programming language called Mojo. Chris, why did you have to build another programming language? Well, the answer is that none of the existing ones actually solved the problem. So what's the problem? Well, the problem is that all of compute is now accelerated. You have GPUs, you have TPUs, you have all these chips, you have CPUs also.
13:20And so we need a programming language that can scale across those. And so if you go shopping and you look at that, like the best thing-ish is C++. Like there are things like OpenCL and Sickle, and there's like a million things coming out of the HPC community that have tried to scale across different kinds of hardware. Let me be the first to tell you, and I can say this now, I feel comfortable saying this, that C++ sucks. I've earned that right, having written so much. And let me also claim, AI people generally don't love C++. What do they love? They love Python, right? And so what we decided to do is say, okay, well, even within the space of C++, there isn't an actually good unifying thing that can talk to tensor cores and things like this.
14:02And so we said, okay, I want, again, Chris being unreasonable, I want something that can expose the full power of the hardware, not just one chip coming from one vendor, but the full power of any chip, it has to support the next level up very fancy compilers and other stuff like this with graph compilers and stuff like that. It needs to be portable. So portable across vendors and enable portable code. So yeah, it turns out that an H100 and an AMD chip are actually quite different. That is true. But there's a lot that you can reuse across that. And so the more common code that you can have, the better.
14:38The other piece is usability. And so we want something people can actually use and actually learn. And so it's not good enough just to have Python syntax, but you want to have performance and control and, again, full power of the hardware. And so this is where Mojo came from. And so today, Mojo is really useful for two things. And so Mojo will grow over time, but today I recommend you use it for places you care about performance. So something running on a GPU, for example, or really high performance doing continuous batching within a web server or something like this where you have to do fancy hashing.
15:09And if you care about performance, Mojo is a good thing. Other cool thing that we're about to ship, so stay tuned, is it's the best way to extend Python. And so if you have a big blob of Python code, you care about performance, you want to get the performance part out of Python, we make it super easy to move that to Mojo. Mojo is not just a little faster than Python, it's faster than Rust. So it's like tens of thousands of times faster than Python, and it's in the Python family. And so you can literally just rip out some for loops, put it in Mojo, and now you get performance wins. and then you can start improving the code in place.
15:44You can offload it to GPU. You can do this stuff. And all the packaging is super simple and it's just a beautiful way to make Python code go fast. Just a double click. You said you were about to ship this? It's technically in our nightlies, but we haven't actually announced it yet. Isn't it? Okay. We do a lot of that, by the way, because we're very developer-centric. And so if you join our Discord or Discord... Nightlies is great. It's like the best program. Yeah. And so we have a ton of stuff that's kind of unannounced, but well-known in the community. Okay. I thought this was already released.
16:11Yeah. And when you say rip out, is this, you have a Python binding to then run the mojo, like you have C bindings? This is the cool thing is it's binding free. So think about like, so today, again, sorry, I get excited about this. I forget how tragic the world is outside of our walls. The thing we're competing with is if you have a big blob of Python code, which a lot of us do, you build it, build it, build it, build it. Performance becomes a problem. Okay, what do you do? Well, you have a couple of different things. You can say, I'm going to rewrite my entire application in Rust or something, right?
16:46Some people do that. The other thing you can do is you can say, okay, I'm going to use PyBind or Nanobind or some Rust thingy and rewrite a part of my module, the performance critical part. And now I have to have this binding logic and this build system goop and all this complexity around, I have Rust code over here and I have Python code over here. Oh, by the way, you now have to hire people who can work on Rust and Python and like that. this fragments your team it's very difficult to hire Rust people I love them but there's just too few of them and so what we're doing is we're saying okay well let's keep the language basically the same so Mojo doesn't have all the features that Python does notably Mojo doesn't have classes but it has functions you can use arbitrary Python objects you can get this binding free experience and it's very similar so you can look at this as being a super fast slightly more limited Python and now it's cohesive within the ecosystem And so I think this is useful way beyond just the AI ecosystem.
17:43I think this is a pretty cool thing for Python in general. But now you bring in GPUs, and you say, okay, well, I can take your code, and you can make it real fast. Because CPUs have lots of fancy features, like their own Tensor Cores and SIMD and all this stuff. Then you can put it on a GPU. And then you can put it on eight GPUs. And so you can keep going, and this is something that Rust can't do. You can't put Rust on a GPU, really, and things like this. and this is part of this new wave of technology that we're bringing into the world. And also, I'm sorry, I get excited about the stuff we've shipped and what we're about to ship, but this fall is going to be even more fun.
18:20So stay tuned. There may be more hardware beyond AMD and NVIDIA. Nice. So that was Mojo, the first concentric circle. Yeah, thank you for keeping me on track. So the inner circle is the programming language, right? And so it's a good way to make stuff go fast. The next level out is you say, okay, well, you know what's cool? AI. Have I convinced you? And so if you get into the world of AI, you start thinking about models. And beyond models, now we have Gen AI. And so you have things like pipelines, right? And the entire pipeline where you have KV cache orchestration and you have stateful batching.
18:53And I mean, you guys are the experts, agentic everything and all this kind of stuff. And so next level out is a very simple Gen AI inference-focused framework we call Max. And so Max has a serving component. Is it Max Engine? Yeah, okay. We just call it Max. we got too complicated with sub-brands this is also part of our R &D on branding each way has the same problem yeah exactly who named LLVM? honestly it's short it's Googleable, not the worst yeah and VLLM came and decided to mess with me so Max, the way to think about it it's not a PyTorch that's not what it wants to be but it's really focused on inference it's really focused on performance and control and latency And if you want to be able to write something that then gets Python out of the loop of the model logic, like it's really good for control.
19:44And so it dovetails and is designed to work directly with Mojo. And so within a lot of LLM applications, as you all know, there's a lot of very customized GPU kernels. And so you have a lot of crazy forms of attention, like the DeepSeq things that just came out. And like all this stuff is always changing. And so a lot of those are custom kernels. But then you have a graph level that's outside of it. And the way that has always worked is you have, for example, CUDA or things like TritonLang or things like this on the inside. And then you have Python on the outside. We've embraced that model, like don't fix what ain't broken.
20:16So we use straight Python for the model level. And so we have an API. It's very simple. It's not designed to be fancy, but it feels kind of like a very simple PyTorch. And you can say, give me an attention block. Give me these things. I can figure out these ops. But it directly integrates with Mojo. And so now you get full integration in a way that you can't get because none of the other frameworks and things like this can see into the code that you're running. And so this means that you get things like automatic kernel fusion. What's that? Well, that's a very fancy compiler technology that allows you to say, okay, you'll write one version of Flash Attention, and then, cool, we can autofuse in Sulu and the other activation functions that you might want to use, and you don't have to write all the permutations of these kernels.
20:58Well, that just means you're more productive. That means you get better performance. It just drives down complexity in the system. And so you shouldn't have to know there's a fancy compiler. Everybody should hate compilers. The only reason people should know about compilers is if they're breaking. And so it just feels like a very nice, very ergonomic and efficient way to build custom models and customize existing models and things like this. And so with Max, we have 500, 600 very common model families and implemented in that. You can see builds.modular.com. We have a whole bunch of models and you can scroll through them and you can get source code and play with them and do that.
21:33That's actually really great for people that care about serving and research and all this kind of stuff. Next layer out is you say, okay, well, you have a very fancy way to do serving on a single node. That's pretty useful and pretty important, but you know, it's actually cool. Large scale deployment. And so we have a next level out cluster level. And so that's the, okay, cool. I have a Kubernetes cluster. I've got a platform team. They've got a three-year commit on 300 GPUs. And now I have product teams. I want to throw workloads against this shared pool of compute. And the folks carrying the pagers want the product teams to behave.
22:10And so they want to keep track of what's actually happening. And so that's the cluster level that goes out. And so very fancy prefix caching on a per node basis. And then you have intelligent routing. And there's a whole bunch of disaggregated prefill. And a whole bunch of cool technologies at each of these layers. but the cool thing about it is that they're all co-designed. And because the inside is heterogeneous, you can say, hey, I have some in AMD, I have some in Vida. Hey, I throw a model on there, run around the best architecture. Well, this actually simplifies a lot of the management problems.
22:41And so a lot of the complexity that we've all internalized as being inherent AI is actually really a consequence of these systems that are being designed together. And so to me, my number one goal right now is to drive out complexity, both of our stack, because we do have some tech debt that we're still fixing, but of AI in general. And AI in general has way more tech debt than it deserves. Well, yeah, I mean, there's only so much that most people, excluding you, can do to comprehend their part of the stack and optimize it. I'm curious if you have any views or insider takes on what's happening with VLM versus SGLang and everything coming out of Berkeley.
23:17Yeah, I don't. I have outsider perspective. I don't have any inside knowledge. SGLang seems, to me, as an outsider, so I'm not directly involved with either community. I just don't have time. So I can be a fanboy without participating. But SGLang... They have a beef going on. So I see. I don't know the politics and drama. But SGLang, to me, seems like a very focused team that has specific objectives and things they want to prove, and they're executing really hard towards a specific set of goals, much like the modular team. So I think there's some kindred spirits there. VLM seems much more like a massive community with a lot of stakeholders, a lot of stuff going on, and it's kind of a hot mess.
Read the full transcript
23:59But they want to be the default inference platform that everyone kind of benchmarks against. Well, and I mean, as far as I know, it's like crushing TRT-LLM and some of the other older systems and things like that. So I think metrics are good probably for that. And so I can't speak to their ambition, but it seems structurally like they're very different approaches. One is say, let's be really good at a small number of things that are really important. Yeah, that's our approach too. One is saying, let's just say yes to lots of things, and then we'll have lots of things, and some of them will work and some of them don't.
24:31One of the challenges I hear constantly about VLLM, for what it's worth, is if you go to the webpage, they say, yeah, we support all of this hardware. We support Google TPUs and Infantria and AMD and NVIDIA, obviously, and CPUs and this and that and the other thing. So they have this huge list of stuff that they support. But then you go and you try to follow any demo on some random web page saying, here's how you do something with VLM. And if you do it on a non-NVIDIA piece of hardware, it fails. And so what is the value of something like VLM? Well, the value is you want to build on top of it generally.
25:03You don't want to be obsessing about the internals of it. If you have something that advertises that it works and then you pick it up and it doesn't work and now you have to debug it, it's kind of a betrayal of the goal. And so to me, again, I know how hard it is. I know the fundamental reason why they're trying to build on top of a bunch of stuff below them that they didn't invent that doesn't work very well, honestly. Because we know that the hardware is hard and software for hardware is even harder these days. So I understand why that is. But that approach of saying, here's all the things we can do.
25:36And then having a sparse matrix of things that actually work is very different than the conservative approach, which is saying, okay, well, we do a small number of things really well. and you can rely on it. And so I can't say to which one's better. I mean, I appreciate them both and they both have great ideas. And the fact that there's the competition is good for everybody in the industry. And so - They're in the benchmark wars state of the way that this industry evolves. That's right. But I also hear just on the enterprise side of things that they're all so good. They don't want to follow the drama, right?
26:06And so like, this is where having less chaos and more I can work with somebody is actually really valuable these days. and you need state-of-the-art. You need performance and things like this, but actually having somebody that is executing well and can work with you is also super important. Have we rounded out the offerings? You talked about Max. Yeah. I realized it impresses me that you named the company Modular and you've designed all these things in a very modular way. Yeah. I'm wondering if there are any sacrifices to that or is Modular Everything the right approach? There's no trade-offs.
26:43If you allow me to show my old age, my obsession with modular design came from back when creating LLVM. So at the time, there was GCC. GCC is, again, I love it as a C and C++ compiler and stuff, and I have tons of respect for it, but it's a monolith. It was designed from the old school Unix, like compilers are Unix pipe, everything is global variables, C code, it was written in KNRC, if you even know what that is anymore. and so it was from a different epoch of software. And LLVM came and said, you know what everybody teaches you in school is you have a front-end, an optimizer, and a code generator.
27:20Let's make that clean, right? And so at the time, people told me, again, first of all, it's impossible to replace GCC because it's so established, et cetera, et cetera. But then also that you can't have clean design because look at what GCC did and it's successful and therefore you can't have performance or good support for hardware or whatever it is without that. But they might be right in an infinitely perfect universe, but we don't live in an infinitely perfect universe. Instead, what we live in is a universe where you have teams. Like you have people that are really good at writing parsers, people that are really good at writing optimizers, people that are really good at writing a code generator for x86 or something like that, right?
27:59And so these are very different skill sets. Also, like we see in AI, the requirements change all the time. And so if you write a monolith that is super hard code and hack today, well, two years from now, is it going to be relevant? And this is why what we see a lot in AI is we see these really cool and very promising systems that rise and grow very rapidly, and then they end up falling. It's because they're almost disposable frameworks. This concept really comes from each of these systems end up being different. But what I believe in very strongly is that in AI, we want progress to go faster, not slower.
28:37That's controversial because it's already going so fast, right? But if we can accelerate things, we get more product value into our lives. We get more impact. We get all the good things that come with AI. I'm a maximalist, by the way. But with that comes the reality that everything will change and break. And so if you have modularity, if you have clean architecture, you can evolve and change the design. It's not over-specialized for a specific use case. The challenge is you have to set your baseline and metrics right. This is why it's like compare against the best in the industry. And so you can't say, or at least it's unfulfilling to me, to say let's build something that's 80 % as good as the best.
29:12But it's got this other benefit. I want to be best of in the categories we care about. And when it comes to using like a prefix caching, page attention, how do you decide what you want to innovate on versus like, hey, you are actually the team that has built the best thing. We're just going to go ahead and use that. Yeah, I'm very shameless about using good ideas from wherever they come. Like everything's a remix, right? And so if somebody, I don't care who it is, if it's NVIDIA, if it's SGLang or VLM, and if somebody has a good idea, let's pull it together. But the key thing is make it composable and orthogonal and flexible and expressive.
29:46To me, what I look at is not just the things that people have done and put into VLM, for example, but the continuous stream of archive papers, right? And so I follow, you know, there's a very vibrant industry around inference research. It used to be it was just training research, right? And much of that never gets into these standard frameworks, right? And the reason for that is you have to write massively hand-coded CUDA kernels and all this stuff for any new thing. And I would want one new D-type and I have to change everything because nothing composes. And so this is, again, where if you get some of these software architecture things, right, I mean, which admittedly requires you to invent a new programming language.
30:26There's a few hard parts to this problem, but the cool thing about that is you can move way faster. And so I'll give you an example of that. This is fully public because not only do we open source our thing, we open source all the version control history. And so you can go back in time and say, like, Chris likes open source software, by the way. Let me convince you of this. I don't like people meddling in my stuff too early. But I like open source software. So you can go look at how we brought up H100, built Flash Attention from scratch in a few weeks, built all the stuff. We're beating the tree DAO reference implementation that everybody uses, for example, written fully in Mojo.
31:03Again, all of our GPU kernels are written in Mojo. You can go see the history of the team building this, and it was done in just a few weeks. And so we brought up H100, entirely new GPU architecture. If you're familiar, it has very fancy asynchronous features and tensor memory accelerator things and all this goofy stuff they added. which is really critical for performance. And again, our goal is meet and beat QDNN and TRT-LM and these things. And so it's not, you can't just do a quick path to success. You have to get everything right and everything line up because any little thing being wrong nerfs performance.
31:36And we did that in, I think it was less than two months. Public on GitHub, right? And so like that velocity is just incredible. I think it took nine months or 12 months to invent flash attention. And it took another six months to get into VLM. And this is just, like, that latter part is just integration work, right? And so now you're talking about, like, building all this stuff from scratch in a composable way that scales against other architectures and has these advantages. It's just a different level. So anyways, our stuff's still early in many ways, and so we're missing features, but if you're interested in what we're doing, you can totally follow the changelogs, and you can, like, follow them nightly where we publish all the cool stuff and all the kernels are public, and so you can see and contribute to this as well if you're interested.
32:18Yeah, do you have any requests for projects that people should take on outside of modular that you don't want to bring in? Absolutely. So we're a small team. I mean, we're over 100 people, but compared to the size of the problem we're taking on. The size of your ambitions. Yeah, we're infinitesimal compared to the size of the AI industry, right? And so this is where, for example, we didn't really care about Turing support. Turing's an older GPU architecture, and so somebody on the community is like, okay, I'll enable a new GPU architecture too. So they contributed Turing support. Now you can use Macs and Colab for free.
32:51That's pretty cool. There's a bunch of operators. So we're very focused on AI and Gen AI and things like this. By the way, our stuff isn't AI specific. So people have written ray tracers, and there's people doing flight safety and all kinds of weird things. Flight simulation? At our hackathon, somebody made a demo of looking at, I think it was a voice transcript, the black box type traffic, and then predicting when the pilot had made a mistake. And the airplane was going to have a big problem and predicting that with high confidence. So it's like your car, you know, you're driving your car until it starts beeping when you're not slowing down.
33:25Yeah, yeah, yeah. That kind of thing for the FAA and stuff like this, right? And I know nothing about that. Like, this is not my demand, trust me. This is the power of, there's so many people in our industry that are almost infinitely smart, it feels like. They're way smarter than I am in most ways. And you give them tools and you enable things. and also you have AI coding tools and things like this that help bridge some of the gap, I think we're going to have so many more products. This is really what motivates me, right? And this is where, I think we talked about last time, one of the things that really frustrated me years ago and inspired me to start Modulant in the first place is that I saw what the trillion-dollar companies could do.
34:03You look at the biggest labs with all the smart people that built all the stuff vertically top to bottom, and they could do incredible product and research and other work. that nobody with a five-person team or a startup or something could afford to do. And here we're not even talking about the compute. We're just talking about the talent. Well, and the reason for that is just all the complexity. It only worked, for example, at Google because you could walk, proverbially, walk down the hallway, tap on somebody's shoulder and say, hey, your stuff doesn't work. How do I get it to work? I'm off of the happy path.
34:32Like, how do I get this new thing to work? And they'll say, I'll hack it for you, right? Well, that doesn't work at scale. Like, we need things that compose right, that are simple, that you can understand, that you can cross boundaries. And so much of AI is a team sport, right? And we want it so more people can participate and grow and learn. And if you do that, then I think you get, again, more product innovation and less just like gigantic AI will solve all problem kind of solutions and more fine grained, purpose built, product integrated AI. I think one way of phrasing what you're doing is, you know, you have this line about it's, you want AI to accelerate and that's contrarian because it's already fast and people are already uncomfortable.
35:11But I think you're more, you're accelerating distribution. I have mixed feelings about the words democratizing, but that's really what you're doing. Well, it used to be, democratizing AI used to be the cool thing back in 2017. I know, it's not cool anymore. Yeah, I know, but it used to be the cool thing. and what it meant and what it came to mean is democratizing model training, right? And it's super interesting, again, as veterans, back in 2017, AI was about the research because nobody knew both what the product applications were, but we didn't know how to train models. And so things like PyTorch came on the scene, and I think PyTorch gets all credit for democratizing model training, right?
35:49It's taught to pretty much every computer science student that graduates. That's a huge deal, but nobody democratized inference. inference always remained a black art, right? And so this is why we have things like VLM and SGLang because they're the black box that you can just hopefully build on top of and not have to know how any of that scary stuff works is because we haven't taught the industry how to do this stuff. And, you know, it'll take time, but I think that that's not an inherent problem. I think it's just that we don't have like a PyTorch for inference or something like that. And so as we start making this stuff easier and breaking it down, we can get a lot more diffusion of good ideas.
36:23I want to double click a little bit more on some technical things, but just to sidetrack on, you know, VLM is an open source project led by academics and all that. And I think a lot of the other inference team, because effectively every team is a startup, like the Fireworks together, you know, all those guys. Your business model is very different from them. And I want to spend a little bit of time on that. Happy to talk about it. You intentionally, well, you believe in open source, but you know, it's not just that you just choose not to make money from a lot of the normal sort of hosted cloud offerings that everyone else does.
36:56Yeah, there's a philosophical reason and differentiation reason. There's a whole bunch of reasons. Yeah, so maybe remind people of what that is. Basically, how do they pay you money and what they get for that and why you pick that. So again, we're doing this on hard mode. We took a long path to product. We're not just write a few computer kernels and have some alpha and go buy a GPU reserve thing and resell our GPUs. That path has been picked by many companies and they're really good at it. So that's not a contribution that I'm very good at, and I'm not going to go build a data center for you.
37:29There's people that are way better than that. And so all the best luck. Literally Crusoe walking the Stargate grounds. I want those people to use Macs. I'm pretty good at this. That's a Crusoe, yeah. I'm pretty good at this software thing, and so you can handle all the compute. Moreover, if you get out of startups, you have a lot of people that are struggling with GPUs and cloud. right and so GPUs and cloud fundamentally are a different thing than CPUs and cloud and a lot of people walk up to it and say it's all just cloud right but let me convince you that's not true and so first of all CPUs and cloud why was that awesome well all the workloads were stateless they all could horizontally auto scale CPUs are comparatively cheap and so you get elasticity that's really cool you can load them pretty quickly it's not like gigabytes of weights Yeah.
38:20And so it turns out, what business do you know knows what they're doing in two and a half years? Nobody. Nobody, right? And so cloud for CPUs is incredibly valuable because you don't have to capacity plan that far out, right? Now you fast forward to GPUs. Well, now you have to get a three-year commit. A three-year commit on a piece of hardware that Jensen's going to make obsolete in a year, right? And so now you get this thing, and so you make some big commit, and what do you do with it? Well, you have to overcommit because you don't know what your needs are going to be and you're not ready to do this.
38:53Also, all the tech is super complicated and scary. Also, GPU workloads are stateful. And so you talk about the fancy agentic stuff and you all know all this stuff. Yeah, it's stateful. And so now you don't get horizontal autoscaling. You don't get stateless elasticity. And so you get a huge management problem. And so what we think we can do and we can help with people is say, okay, well, let's give you power over your compute. And a lot of people have different systems, and there's very simple systems that go into this, but you can get like 5x performance TCO benefit by doing intelligent routing.
39:27That's actually a big deal. For a platform team, they don't like to have to deal with this. They want hardware optionality to get to AMD, and they want this kind of power and technology. And so we're very happy to work with those folks. The way I explain it in a simple way is that a lot of the endpoint companies, and there's a lot of them out there, and so you can't make, they're not all one thing. They obviously have pros and cons and trade-offs, but generally the value prop of an endpoint is to say, look, AI and AI software and applications and workloads, it's all a hot mess. It's too complicated.
39:59Don't worry your little head about it. I'll take AI off your plate so that you don't have to worry about it. We'll take care of all the complexity for you. And it'll be easy. Just talk to our endpoint. Our approach is to say, okay, well, guess what? It's all a hot mess. Yes, 100%. Like, it's horrible. It's worse than you probably even know. and get to tomorrow, it's going to be even worse because everything keeps changing. We'll make it easy. We'll give you power. We'll give you superpowers for your enterprise and your team because every CEO that I talk to not only wants to have AI on their products, they want their team to upskill in AI.
40:31And so we don't take AI away from the enterprise. We give power over it back to their team and allow them to both have an easy experience to get started because a lot of people do want to run standard commodity models and you do want stuff to just work as table stakes. But then when they say, hey, I want to actually fine-tune it. Well, I don't want to give my proprietary data to some other startup, right? Or even some big startups out there and you know whose these are. That's my proprietary IP. And then you get to people who say, hey, I have a fancy data model. I actually have data scientists.
41:00I actually have a few GPUs. I'm going to train my model. Cool. That got democratized. Now how do I deploy it? Right? Well, again, you get back into hacking the internals of VLM and PyTorch isn't really designed for KVCache optimizations and all the modern transform features and things like this. And so suddenly you fall off this complexity cliff if you care about it being good. And so we say, okay, well, yeah, this is another step in complexity, but you can own this and you can scale this. And so we can help you with that. And so it's a different trade-off in the space, but I will admit that there are time to market and revenue growth and stuff like that has been much faster because they didn't have to build an entire replacement for CUDA to get there.
41:37Nice. And when it comes to charging, people are buying this as a platform. It's not tied to token, inferred, anything like that. We have two things going on. So the Max framework and the Mojo language, free to use on NVIDIA and CPUs, any scale, go nuts, do whatever you want. Please send patches. It's free, right? Why is this? Well, it turns out CUDA is already free. NVIDIA is already dominant in here. We want the technology to go far and wide, use it for free. It would be great if you send us patches, but you don't even have to do that, right? We do ask you to allow us to use your logo on our webpage.
42:11and so send us an email and say hey you're using it and you're scanning on 10 ,000 GPUs that'd be awesome but that's the only requirement if you want cluster management and you want enterprise support and you want things like this then you can pay on a per GPU basis and you can contact our sales team and then we can work out a deal and we can work with you directly and so that's how we break this down and also let me say one thing I would love to see again it's still early days but I would love to see PyTorch adopt Macs I'd love to see VLM adopt Macs I'd love to see SGLang adopt a Macs like we have our own little serving thing, but go look at it.
42:44It's really simple. Like it would be amazing. And again, we're in the phase now where I do want people to actually use our stuff. And so we have historically been in the mode of like, no, our stuff's closed, stay away. But we're phase shifting right now. And so you'll see much more of this being announced. Yeah. I don't know how to make this happen, but like I think you win when Mistral, Meta, DeepSeek, and Quinn adopt you and ship you natively, right? How does that happen? I don't want to talk that far in the future. But we may have... I don't think it's that far. We may have an industry-leading state-of-the-art model launching first on Mac soon.
43:24I'll stay tuned for that, but also let us know. Yeah, but that hasn't happened, so assume it doesn't happen. I think that's when it really tips because then everyone's like, okay, it's good enough for them, it's good enough for us, right? And then you get the rest of the industry. Yeah, but again, I mean, I'm in it for a long game, right? And I realize, again, the stuff I work on, it takes time for people to process. And so what we need to do and what I want us to do and what I ask the team to do is keep making things better and better and better and better and better. And there's like an S-curve of technology adoption.
43:53And so I think it's great that there's a small number of crazy early adopters that were using our stuff in February before it was open source. And it made no rational sense. It barely worked. But it was amazing. And I'm very thankful for those people. And then, of course, we open source it and we start teaching people and you get a much bigger adoption curve. You make it free, go adopt and go. And then, as you say, there's more validation that will be coming soon. And each of these things is the knee in the curve. But what it also does, it gives us the ability to fix bugs and improve things and add more features and roll out new capabilities.
44:26How does this feel rolling this out as compared to Swift? Oh, well, so let me reinterpret your question. of given you've done a few interesting things in the past, what have you learned and what are you not doing again? That actually is a better question than the one I asked. Because Swift is too narrow almost. The character of Swift, because I assume most people don't know about this, the character of Swift was, I started as a nights and weekends project in 2010, hacked on it alone nights and weekends for a year and a half, eventually told management at Apple about it. Their heads exploded. Like, why do we need something?
45:02Objective-C is good enough. why do you need this? Got approval to have a couple more people get involved in it. You were on a fellowship or an internship at Apple at the time? No, I was leading the developer tools team. I think you're doing something. Yeah, I was leading a huge team. Let's just say this was not my day job. But so it started in 2010. It launched publicly by Apple in 2014. And by the time it launched in 2014, only about 250 people in the world knew about it. Most of whom were in my team. about 200 and something of them were in my team. And then it was senior execs, marketing, Tim Cook, et cetera, right?
45:38And this was the category of people that knew about Swift. So we had built it in secret, literally in NDA, within Apple, to know about it, right? When we launched it, part of the requirement was that you had to be able to submit apps to the App Store in Swift, right? That was the requirement put on me. And so it's like, cool, that sounds great. And so we launched it and said, it's a 1.0. So you're launching a 1.0, brand new programming language, Nobody has ever seen it before. No internal user, like one demo app. It was a freaking nightmare. So it was a nightmare for the community because, I mean, fortunately, a lot of people were excited and wanted to adopt it.
46:16And a lot of people did adopt it right away. But it was not battle-hardened. It had tons of bugs. We should have launched it as a 0.5, right? And so it took another year for it to become pretty good and then two years for it to become quite good, in my opinion. Also, none of the software engineers at Apple knew about it. And so their heads exploded. They said, wait a second, why are you replacing Objective-C? I joined Apple because I love Objective-C. Why didn't you ask me my opinions about the new programming language? And so there's that whole dynamic. Oh, was there a company mandate that they had to write Swift from now on?
46:47No, but still, it's like, wait a second, this isn't the company I thought I joined, right? And stuff like this, right? And so there's this huge amount of turmoil and drama and nonsense that came out of that. And so, okay, fast forward to Mojo, lessons learned. Hey, one, don't have a hot start. And so we launched Mojo a long time ago, before it even made sense, and we called it a 0.1. And so how's that honesty in advertisement, right? It's like, this is 0.1, please don't use it, but if you're interested, we'd love your feedback, right? And so soft start, go. Second thing that's very different is that in Swift, we had one demo app.
47:23And so you have a very, I think, high-powered team building a language that had done lots of credible stuff, but was building a language for iOS developers and the compiler was written in C++. And so, yeah, there's sympathy for the user, but not a lot of understanding and not a lot of learning internally when we launched. In the case of Mojo, guess what? Modular is Mojo's first customer. We have more Mojo code in our repository than any other language. That's awesome. And it's open source. And we open source 650 ,000 lines of Mojo code. right? This is a lot. And so we suffer and we drive the features and the improvements based on our needs.
48:03We also appreciate the community and we have a whole bunch of contributions coming in and somebody just optimized my string, you know, to get rid of a bit out of my string implementation, which was suboptimal. And so that was super awesome. But driving it that way, make sure it's real. It's grounded. It's on the use case. It's not, we're trying not to over promise. Like even when you're asking me what it's useful for earlier, right? I didn't say it's a replacement for Python. I said it's a go-fast language. Someday it may be a pretty credible Python alternative, but for right now, it's good at a specific class of use cases.
48:36And if you're interested in those use cases, like making GPUs go burr, Mojo's awesome. But if you want a replacement for Rust end-to-end, then give us six months. Yeah, yeah. I mean, you're a force of nature. I think there's a lot of mystery around what is going on with Apple's AI initiatives. And I think the consumers suffer. At the end of the day, the end users are waiting for this. Unfortunately, anything I know is massively out of date. They've changed and re-org'd and grown and culture. And it's a very successful company. And so I think that they probably feel success and they're having trouble adapting to changes in the industry.
49:15And that's pretty typical of a lot of big companies. And so I can't speak to the specific causes. Speaking of Google, obviously I think they were one of the earliest. What are we talking about? The invented transformers. A lot of other things. TensorFlow, remember that? Yeah, exactly. Huge. Yeah, NTPs. I credit Google with making AI open source. Well, they did not have to open source TensorFlow. Yeah. That was an incredible decision. Full kudo to Jeff Dean and many of the other people that were involved in this, because they said, you know what's actually the most important thing for Google? It's for AI to go faster.
49:48How do we do that? We open source TensorFlow rather than making it some proprietary internal thing, which they had a previous system called disbelief. And so that is a huge moment that set the stage for PyTorch to be open source and for the research to be open and for all of these things because they decided the value system was AI go faster. The transformer paper being published, like so many contributions from Google came from that. I don't think Google gets enough credit for that. Yeah, like why is it better for Google for AI to be open source rather than Google owns it? Well, so I can't tell you if the bet worked.
50:24But I can tell you that that was a bet. But from my outsider now perspective, because I haven't been at Google for over five years, somehow time flies, the bet makes sense when you have an amazing team of researchers and SWEs that can go incorporate this into your products. And so Google does have billions of users. It has all the product services. It has all the different applications. And it has an incredible density of talent. And I think that Google's recent announcements, so just after Google I.O. and things like this. Yeah, we were both there. Yeah, it's like Google's actually working, I think.
50:57It's pretty impressive. And for a while, they were dealing with organizational drama and Google Brain versus DeepMind and some of this stuff. And I can't speak to what they've done, but it seems like they're a much more unified team. They're executing well. They're getting research into product. And so it feels like a different Google to me. Yeah, it totally does. It used to be that there was just two of everything in Google. and you didn't know which one to use. Killed by Google. They all deprecated in a year. So yeah, I think they've gotten the memo. Yeah, and the other thing that's super impressive to me about them and this me, fanboying Google, right?
51:28After railing, it's a trillion dollar companies. But the things that they announced that are actually shipping. Yeah. So much in AI is... This is more Apple shade. I wasn't specifically saying Apple. This is very common in AI. Modular's done this in the past too. This is why. So I gave this very deep tech talk. I sent you a link to the GPU mode talk. And the slide two was warning. You can actually use this. This is not vaporware. Everything here you can reproduce. These claims you can download. This is actually real. Like, here's links to the source code. I was wondering why you stressed that so much.
52:02I'm like, who hurt you, you know? Like, I know. I mean, there's so many claims. No one knows what is real anymore, right? And I mean, there's literally been product demos where, you know, it's like some electric semi rolled downhill instead of working under its own power. Nobody knows what's real. I knew they started to work. There's a WhatsApp chat with playing soccer at the Google field during the week. And about six months ago, the admin posted, it's like, hey, not enough people are showing up anymore to play soccer at lunch. What is going on? And I think that's when people started working again.
52:34I mean, Sergey Brin was at I.O. and he's working again. It's awesome. I have mad respect for that. And so my values are aligned with people who ship stuff. Yeah. Because that's what impacts the world. Let's talk about open source a little more. There's the more recent maybe open source thing, which is DeepSeq, obviously. And I think specifically in your case, they worked at the PTX layer of the GPU, which is even lower and more proprietary than CUDA. I'm curious how, both in terms of, obviously the impact was huge, but maybe impact on how much people should actually try and move away from this proprietary thing.
53:11Because now the next, from my understanding, is like the next set of chips, all the code is like useless. Well, so it's not widely known, but Blackwell is not compatible with Hopper. Hopper kernels don't always run on Blackwell, for example, right? But so your question is like, what does it mean for the industry? Well, it's like, why is it so important? Like why the DeepSeq, the DeepSeq example is like so important of like they need to navigate all this like proprietary stuff just to make it work. So I'll give you my lived experience because DeepSeq came out in December, which is when I, and probably you noticed it, right?
53:43But then the world had a big wake-up call and video stock price went down and all that stuff like a month later. So here's my explanation of what happened. What the DeepSeek team did was really impressive research. They pushed MLA forward, which is a form of attention. And they pushed low-precision training forward. They pushed a whole bunch of stuff forward. They reverse-engineered some PTX instructions that weren't well-known at the time. And so a lot of people were just like, and it was a Chinese team, right, which put Americanism, threatened Americanism, right, and things like this. And so what I found really exciting about it was they pushed the research forward and they did this incredible thing.
54:20They showed the world that it was possible. And they opened it. And they published it. And they actually taught the world about it because I don't know why they chose to do that, but it's because they believe in openness and AI moving forward, right? And so the thing that I found striking is that the world's reaction to that was more striking to me than the actual world. or whatever. I don't know. Because, so first of all, there's the Chinese-American drama, which geopolitics is not exactly my strong point. So I get it, but I push that aside, right? But the other thing I found really interesting is that people said, wow, okay, only DeepSeek is able to go down to the PTX level.
54:58But that is standard. That is what all of the leading teams do. In the case of Modular, we go literally, like we only work at that level because we get rid of all of Kudo, right? And so we've replaced the entire stack. And so we only do that. But a lot of teams that care about performance will actually go down and use the PTX, TensorCore instruction, Foo, which isn't really documented. You kind of have to figure it out and look at Cutlass code. NVIDIA doesn't make it as easy as they should to do this. When do I? Well, I think it's because they're breaking in Blackwell, and so they didn't want people to actually use it, and so they had their own issues.
55:33But good luck with that. There's a lot of smart people out in the world. And so to me, I thought it was really interesting because there wasn't a lot of awareness of how that level of the stack worked. For me, I thought it was a great wake-up call where people were like, oh, wow, if you work at this level of the stack, you know, level of stack I have inhabited for decades now, you have power over compute. You can understand and solve problems. You can drive research forward in ways that nobody else can do. And this is what the trillion dollar companies do. It's not just DeepSeq. but I think DeepSeq was a huge wake-up call and it drew attention to that layer of the stack because it wasn't just like throwing layers on top of PyTorch or VLM or it was like doing that fun amount of work and I thought that was incredible.
56:11Now, the challenge with it and the challenge with the way DeepSeq did it and the way that everybody else does all this work is that it's completely specific to one GPU. It's not just that you're working at the PTX level, it's that you're writing code that really can only work on that one GPU and that means that when Blackwell comes out, you have to throw it away and write a new one, right? But here's the open secret. That's what VLM is. Go look at VLM. They have different kernels for A100 than H100. They're now trying to catch up with Blackwell. And so state of the art for these systems since Gen AI, before that there were fancy AI compilers like XLA and that kind of stuff in the Trat AI world provided some scalability.
56:47But in Gen AI, it's a rewrite all the things when a new piece of hardware comes out. And so this is what Mojo is solving, right? And this is where, again, we can't turn the incremental cost of a new piece of hardware to zero, we can massively reduce it. And so this is a really exciting time, I think, for us that we've demonstrated now, but also what it means for the future. I still also wonder their hiring or internal training that they managed to have a small team that does this in the same way that you do. I think they have a pretty significant team. I have no insider information, of course, but it's not like a five-person team.
57:19It's like hundreds of people. Okay, yeah. Yeah, you know, it's not a thousand, right? It's amazing, especially like, I mean, presumably there's some language barrier, but even ignoring that, just getting that amount of talent density in one company is not a significant task. I agree with that. Building a company is hard. I have no visibility in how they built DeepSeek, but all I can say is thank you to DeepSeek for publishing their work. Because they didn't have to do that. And I think it left temporarily a lot of people flat-footed. And it was kind of embarrassing for certain groups. And I think a lot of people paused and were like, ooh, crap, what do I do about this?
57:53But I think that what it did is it pulled forward progress in AI by like six months. Yeah. They didn't just do GPU level stuff. They also like had a file system. Oh yeah. If you've like looked at, like do you see, obviously that's not modular as bread and butter, but like, do you see a potential there for someone else to, I don't know, take that and run with it? I've been very obsessed with the inference problem for a couple of years. And so I don't know the best way to solve that problem for training. Like Google internally has a great, right? And so like, where's that for the rest of the world?
58:24Yeah, I'm honestly just not the right person to answer that question because I have my own obsessions. While talking about inference, reasoning models and inference time compute, very big topic. Does anything change or does nothing change because it's just more inference? It depends on change from what, right? So when we started Modular over three years ago, we made the ridiculously weird bet at the time to focus on inference instead of training. Yep. And again, at the time, I'm used to this. people are like, what's wrong with you? Everybody obviously knows that training is the thing and training, training, training, and people building these massive clusters and all the spend is on training, et cetera, et cetera, et cetera.
59:02And I said, well, yeah, I understand why you see that, but I've lived this at Google. Google's like five years ahead of the rest of the industry in many ways. Training scales the size of your research team. Inference scales the size of your customer base. And so the thing that happens between those is research gets into production. And so the gap is research game production. Once you do that, suddenly it scales like crazy. So I didn't plan for Gen AI. I did not plan for inference time compute and things like this. But that bet on the production use case, because it scales with the number of applications of AI, not just the hot research team, right?
59:38Which, you know, is pretty important. It's just not something I've focused on for the last few years. That was controversial. And so I think now the whole world's flipped and I think the world gets it, right? But that was really because of lived experience. And so when you come to these new techniques, right, another controversial thing going back to the CPU thing is why are you starting with CPUs? The answer at the time was pre-processing, post-processing, full system integration, networking. You need a CPU to feed the GPU, like very standard things. But now you say KV caches. Like your eviction policy runs on a CPU.
1:00:10Like that radix hashing algorithm and block hashing and all that stuff happens like primarily CPU. that's really important for performance because if you have latency in these steps you're not keeping your GPU utilized and again this comes back to the rewrite and rust and things like this all the agentic stuff and things like this I didn't predict that but I'm not surprised and I think we're going to see more and so tight integration optimization across boundaries these are things I believe in and I think this is how you move the world forward it's amazing how basically every all the smart programmers I know always focus on where the bottleneck is And it always leads you to the right answer.
1:00:48Like if you just are very clear-eyed about that. And I don't know, it's nice to see that happening. The other talking point I'll mention for you because you didn't bring it up, but it's something that other people are talking about is that actually now because of the requirement for train of thought and reasoning models, inference is now part of training. Okay, yeah. Because you need to inference. RL has always done that. To RL, yeah. Yeah. And so that actually is a plus in favor of what you're doing anyway. Yeah, absolutely. Because it's the same code. So my experience with the RL systems were like at scale, DeepMind style, AlphaGo and things like this.
1:01:25So it's been a couple of years ago, but setting those things up was incredibly complex because now you're dealing with cluster scale orchestration, you're batching across all these agents. And again, none of the systems are set up for that, right? You've got PyTorch if you want to train a model, but nothing is set up to do this. And so, again, I can't speak to all RL systems everywhere, but they ended up being like duct tape and bailing wire and super crazy stuff. And it was incredible. Again, it's incredible what a team of experts who knows the full stack and what they can achieve, but it shouldn't be that hard.
1:01:57And so we're just not focused on solving that problem yet. Maybe we'll get there. But you have the building blocks. We have the building blocks. And again, we're not focused on solving training yet. Maybe we'll get there. I have some ideas on logical steps to do that, but I want to make sure we're grounded and we solve things all the way end to end. We make people happy. People talk about our stuff from their voice. It's not just me talking about our stuff. And so we're in that phase where you'll hear people talking about our stuff soon. I think people already are, but they will even more. And I think I appreciate you coming on the podcast to talk about that more.
1:02:29We want to turn to some personal stuff. So I think that was a great modular story. I think now on the personal side, there's a couple of things. So when you first joined us, I think I saw it was September 2023. So that was a year and a half into building the company, something like that. Now three years and a quarter. And what are like personal learnings, both from obviously being a leader in a company, having a growing team that you're managing, having a lot of responsibilities that are not technical anymore? Kind of run us through some of that. Yeah. So for me, I've built large teams from scratch before, but they've all been established companies.
1:03:04I've worked at startup before, but it was somebody else's startup. And so this is the first startup that I've founded and then built a significant team and built a product that takes years to build. So a lot of the lessons I learned were super valuable and allowed me to achieve some of the stuff. But it's also very different. And so one of the things that's very different is it's very personal. And so I've lost people from our team before, like many times. I've had to fire people and people have quit and gone Apple to Google or whatever. But that was never personal in the same way it is at Modular.
1:03:36And so this is something where the intellectual side of my brain knows, like if somebody leaves, it makes sense. They had a life change. Like I don't want to get in the way of their family or whatever. I mean, I intellectually know that. But on the other hand, it hurts a little bit. And so I think I'm getting better at handling some of that kind of stuff. The other thing is that when you're growing a team from like zero to 100 people at a big company, you have all of the infrastructure and the infrastructure is mature. right and so you're even if you grow to a team of 100 people you're still tiny you know compared to an apple or google or a company like this like you're still tiny in proportion to the size of the overall scale and so they've already got all the recruiting and all the other stuff and all the legal and finance and all that kind of stuff going they've got the manager training they got all this stuff going and so in a startup it's sometimes getting some hot messes where it's like okay well i need to do reorganization and things like this and so that's been um good learnings.
1:04:27I think we've scaled into that well. And another thing coming back to the people telling you it's impossible. Actually, here's probably the most important thing is that I'm used to being told that things are impossible. And then I'm pretty bullheaded and I have a formula. And so like, there's a path to success that I can explain if you want, but the, uh, I'm used to the feeling when it's, you know, that year one, and it's just a completely skunk works. Let's like prove a thing phase. I'm used to that. Okay, cool. Now we're telling people about it, but it's still not good enough. And people tell you all your stuff sucks.
1:04:58I'm used to that. Like, okay, well you get into that window where all the stuff is almost there. People are sweating. It's really hard. There's like all these like seemingly impossible things. It's a significant team, but the pieces don't line up yet. And so we went through some of that last fall and people like, oh my gosh, maybe it will never work. And you get this like anxiety. And I think the thing that I didn't appreciate going through that, and this is also why I'm so excited to be in this phase, that actually impacts other people's, like, their thought processes because they haven't been through it.
1:05:29And they're like, trust me, we'll get there. Well, when? Like, exactly what? Like, you get these, like, super analytical engineers who want to understand everything, and they're really expert in their part of the stack, and they don't really have that established trust with the other department and the other department and the other department are all lining up. And so definitely some warnings from that. But this is where, again, you get out of that R &D phase and you get into the execution phase, and it's like, okay, well, engineers are really good at taking an imperfect thing that works end-to-end and making it better and better and better and better and better.
1:05:59And so it just feels fundamentally different at Modular now than it did in any of those previous phases. Yeah, what's your day-to-day? You have a lot of meetings, you have just a few. Yeah, I have a lot of... You're talking about my lived life. A normal weekday for me is I wake up about 7.15, my wife and I get the kids out the door and she drives them to school at about 8. I usually do a half hour, 45 minute walk with my dogs, get exercise, strenuous up and down a hill, heart rate goes up, which is good. Listen to podcasts, for example, yours. So that's why I have time to actually follow exciting, cool things that other people are doing.
1:06:35Get to work at nine. And so work nine to six or seven or something like that. And most of that is meetings. And so that's me trying to solve whatever the problems are of the day, get home, dinner with the kids. I insist on eating with my family, hang out with them until bedtime and then crush through for like two, three more hours until I pass out and do that regularly. The second work day. And then on the weekend, it's amazing because I get a lot of time to work, not all day because I do things with kids and stuff, but I get a lot of time and there's no meetings. So that's amazing. That's when stuff actually gets done.
1:07:08Yeah, exactly. I think the key here is some kind of strategic review time that you lock off because like, I think people often say, you know, there's times when you're working in the business and then there's times you're working on the business where you step out for a bit. Often for founders it's when you do a board meeting. But I wonder if there's anything that's very meaningful for you. Do you have a coach? Do you have something like that? Yeah, so I guess there's two people I really owe a lot to. Or two categories. So one is my co-founder Tim. So Tim and I formed the company together. We walk every week, every Friday, catch up and it's like that zoom out and try not to make it tactical.
1:07:48And so make sure that we can bounce crazy ideas off. And I have a lot of crazy ideas, believe it or not. He does too. But then ground ourselves on execution and go. The other thing is this combination between my wife, who's a sounding board, and the executive team. And I, I'll admit, get a little bit crazy and want to solve the industry problem. And the exec team pulls it back to, okay, next quarter, let's make sure we have a plan. Let's make sure we can communicate. Let's decide what the actual priorities are. Okay, we can have three priorities max for the whole team, not 50 or something. Otherwise they're not priorities.
1:08:21Yeah, exactly. Everything can't be the top priority, otherwise nothing else, right? And things like this. And then my wife, who keeps me sane and is an amazing life coach. How much do you get your wife involved on the actual work? Zero. Yeah, she has her own thing going on. My wife runs the LVM Foundation, and so she's got a bunch of things going on, plus kids and everything else, too. She's like, I don't want to hear about this kernel. Well, she's great for helping me. so I'm more IQ than EQ. And so working with humans is not, it's an acquired skill, not a natural skill. And so I think this is something where once, you know, I often end up at this place where some weird thing is happening.
1:09:00What the heck is going on? It's like, Chris, it's obvious. Like they're saying this, but this is what they actually mean. I'm like, oh, I never thought about that. You mentioned coding agents. Yeah. What do you guys use internally? What do you like? Yeah, so I mean, as of this recording, I personally use Cursor. And so Cursor is great for, I mean, it's the best thing I've seen. And I don't spend a lot of time dabbling with things. And so, but Cursor is, so I write a lot of C++ and Mojo code. The key thing for working on Mojo code in AI coding tools is make sure you have a lot of code in the context window.
1:09:31And so Cursor and these tools can index really well. I was thinking you open source Mojo code. Actually, that's pretty good. That's one of the reasons we did this. Yeah. And one of the many reasons we did this. And so this is what we saw at our hackathon is people could go zero to hero with Mojo because you could just put, you just index this entire huge code base and it's phenomenal. And then learning a new language is actually easy when the AI is doing a lot of the mechanical stuff for you. And also it looks like Python so you could read it. But it just like massively scales the on-ramp. And so for new language adoptions, AI is, let me convince you AI is cool.
1:10:01I would just go ahead and like, actually like ask the datasets people at the big labs what they need. And then just like feed it to them. Yeah, just take the code, please. It's on GitHub. It's Apache 2. Just go. Yeah, we're adding markdown files. There's a cloud.md, and so we're doing some of the basic stuff. People within the company are dabbling with cloud code and some of the stuff, and so I don't have personal experience with that. But for me, I found that it's mostly, it is very useful, but it's really about boilerplate. And so it's not about inventing new algorithms and stuff like this, and it's probably because I'm not building a React component.
1:10:33Yeah. I mean, so the labs are touting that they are training or benchmarking their models for writing CUDA kernels, right? I wonder if you see a noticeable performance improvement when you change model to model. I don't know if you, you probably don't benchmark them. I haven't looked at that specific use case, but people often ask me, hey, Chris, why are you building a new programming language when AI is going to write all the code? Similarly, if you look at a lot of this, like let's generate a CUDA kernel, you start to ask and wonder, or at least I do a lot, what is the purpose of code? What I've reflected on the way I currently think about it, subject to change, obviously, because we're all learning, I take a step back and I say, well, code isn't really about telling the computer what to do.
1:11:18Code is about humans being able to understand what code does. And so someday when we get AGI or ASI or something like this, maybe it can be completely opaque and I really don't have to know. We're not there yet. So, and I don't know when that will happen, But in the meantime, I want to be able to look at what the code actually does. And I live in a world of constraints. I need to know I have a product which has all these features. If I go add another feature, what happens? Is it going to hit my latency budget? Is it going to crash around memory? Is it going to cost too much? I need to be able to reason about this.
1:11:50And so to me, I look at a lot of these coding tools, scaled beyond where they currently are now, but into the foreseeable future is saying like, okay, well, it's like hiring another engineer onto your code base or into your team. And fundamentally, coding and software engineering is a team sport. You have product managers, you have engineers, you have a lot of things. And if you automate all the engineering of code, maybe you get to product managers only and marketing only or something, theoretically. But you still want to reason about what the code does. And so in that... It's like an interpretability argument.
1:12:20And so in that world, what is actually the most important thing? The most important thing is you can express everything the hardware can do because you don't want to have some capabilities some cost or some boundary you can't penetrate. Mojo does that. The second is you want readable code that you can actually understand, right? And so you don't want assembly language or something like that. You want something high level and expressive and easy to understand. And so like this is where I think Mojo is really unique. And then the AI coding tools I see is a straight value add for adoption. Because I mean, I've already seen what people that have never touched any of this stuff can do.
1:12:52And it's just incredible. You put somebody that's intelligent, they know the use case, and now you put these tools in their hands and they can do amazing things already. And so I find it super empowering. And again, come back to me wanting to see humans being empowered with new technology and being able to upskill and be able to learn. You're not a GPU programmer today, but tomorrow, hopefully there'll be 10 times as many people programming GPUs. I think that'll make the world a better place. And so I don't think we're going back to the world of CPUs. I think GPUs will only get more important.
1:13:22And if we can help people do that, I think it's great. You mentioned some of the research on inference and the archive papers. How do you keep up with Archive? Yeah, we're a Slack shop, and so we have a paper channel. I think there's, I don't know, somewhere between three to ten papers a day that come through. And so we have amazing smart people that do that. I also follow the Reddit communities and things like this. I'm an old-school person that uses RSS still. And so if you have an RSS feed, then it's way easier for me to follow you. Which reader? And I use Feedly. Feedly. Yeah, but Archive has the ability to follow specific groups, and so I do that for various groups, and so that's another technique.
1:14:04Any notable papers come to mind that you want to give a mention to, or authors that you're really watching? Anytime they publish, you're like, I'm reading that one. I can't give you a well-considered answer. Just this morning, a person from Microsoft published a paper. forget his name offhand, but he just published a really cool paper about auto-generating flash attentions and doing Blocker or something. It had some cute name with Block in it. And so anyways, I mean, there's a ton of things going on. Okay, so that kind of, yeah, you know, it's interesting to just see what you pay attention to.
1:14:37Last but not least, the question that everybody wants to answer to. Have you finished building the Lego robotics table for your kids that you mentioned last time? And what's the, do you have a new project that you're working on? massively forgot about this. So we still use the Lego robotics table. So this is a big 4x8 sheet plywood with a bunch of 2x3s around the edge, but a 4x8 sheet applied was pretty hard to work with, and so it breaks apart into 3D composable modular sections, and so it's super great, and the kids are getting way better at programming and Legos and all this kind of stuff. Gosh, what's my most recent project?
1:15:09I was just building swords in the shop with kids on a bandsaw. And so you take a piece of wood, you have a bandsaw, Give it to a kid, and you say, don't cut your fingers off. Turns out that a bandsaw, I mean, I'm obviously joking a little bit. They get a lot of oversight, but a bandsaw is actually a very safe tool. And so the reason for that is that a bandsaw, which if you, probably a lot of people have never seen a bandsaw, you can do a Google search for it. You have two wheels, and then you have a blade that goes around the wheels, and you've got a table. And the cool thing about a bandsaw is that you can crank down the opening towards the blade, so it's just as big as a piece of wood.
1:15:46And also it pulls the wood into the table. And so because it pulls the wood into the table, the risk of something called kickback, and there's a lot of other things like this, is very low. And so you can basically tell a kid, look, you see your fingers, keep them away from the blade. And there's no sudden jerking or other things that if you use the wrong technique, you can get yourself into real trouble. And if you keep the guard all the way down, then you can't get an arm in there or something like that. Did you make your whole garage woodworking station? I'm a cars outside kind of guy. So that's, again, thank you to my wife for tolerating my odd behaviors.
1:16:21I mean, I love building things, right? And so this is fundamentally when you talk about what makes me tick, because I love the joy of discovery, right? And so whether it's building an amazing team or building a new table, you know, built like dining room table or things like this. You can look at my website. I'm not a very good web designer. I have some woodworking projects on there, but I love building software. I love building and solving the problems that come to this. I'm not super great at building things that are rope mechanical. And so maybe I'd be good at building one chair, but I'm not going to build eight chairs, go around a diagram table.
1:16:50That would just drive me crazy. The discovery and the learning is what... You can build the machine that builds the chairs. There you go. That sounds great. Spend 10 years building a thing. Three weeks, crank it out. Yeah. Any call to actions? Are you hiring? Yeah, we're hiring a small number of elite nerds. And so if you care about GPU programming, you care about AI models, you care about inference, you care about Kubernetes and cloud scale stuff, please check us out. We expect to grow a lot more later this year. The other thing is that we have a ton of open source code. And so if you've heard about Mojo, but you looked at it a year ago, guess what?
1:17:29Everything's completely different now. And so if you are interested in a lot of these things, if you're interested in learning about GPUs, we have a ton of content that will teach you about GPU programming, GPU puzzles and things like this. People are now picking up Mojo and putting in a lot of like Leet GPU. It came out today and there's a whole bunch of other people that are taking the stuff and putting it out there. And I think it's just such an exciting time because I think that lots more people should be programming GPUs. I think this is a huge opportunity for the industry. And of course, if you're enterprise and you're having trouble scaling your AI, let us know.
1:17:58We can help. Awesome. Thank you so much for coming on here. Inspiration as always. Yeah. Well, thank you for having me.
1:18:10Thank you.
From the publisher
Chris Lattner of Modular (https://modular.com) joined us (again!) to talk about how they are breaking the CUDA monopoly, what it took to match NVIDIA performance with AMD, and how they are building a company of "elite nerds".
X: https://x.com/latentspacepod
Substack: https://latent.space
00:00:00 Introductions
00:00:12 Overview of Modular and the Shape of Compute
00:02:27 Modular’s R&D Phase
00:06:55 From CPU Optimization to GPU Support
00:11:14 MAX: Modular’s Inference Framework
00:12:52 Mojo Programming Language
00:18:25 MAX Architecture: From Mojo to Cluster-Scale Inference
00:29:16 Open Source Contributions and Community Involvement
00:32:25 Modular's Differentiation from VLLM and SGLang
00:41:37 Modular’s Business Model and Monetization Strategy
00:53:17 DeepSeek’s Impact and Low-Level GPU Programming
01:00:00 Inference Time Compute and Reasoning Models
01:02:31 Personal Reflections on Leading Modular
01:08:27 Daily Routine and Time Management as a Founder
01:13:24 Using AI Coding Tools and Staying Current with Research
01:14:47 Personal Projects and Work-Life Balance
01:17:05 Hiring, Open Source, and Community Engagement




