In short
Notes on The TWIML AI Podcast Episode #634: Mojo: A Supercharged Python for AI with Chris Lattner
Podcast Overview Title: The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) Host: Sam Charrington Guest: Chris Lattner, Co-Founder and CEO of Modular Episode Description: Discussion on Mojo, a new programming language designed to enhance AI development by making it more accessible and high-performance, particularly for Python programmers.
Key Concepts
- Mojo: A new programming language aimed at AI developers.
- Designed to simplify the programming stack, making it easier for researchers and developers who are not compiler engineers.
- Functions as a superset of Python, allowing for high-performance execution and compatibility with existing Python libraries.
- Modular Engine: Powers the Mojo programming language and provides support for both TensorFlow and PyTorch.
- Aims to streamline the deployment of AI models and enhance performance across different hardware.
Guest Background
- Chris Lattner's Experience:
- Known for his work on LLVM and the Swift programming language.
- Transitioned into AI around 2016, contributing to Google TPUs and TensorFlow.
- Experience spans multiple levels of the AI technology stack.
Discussions
Mojo
The New Language for AI
- Motivation for Mojo:
- Simplifying the traditional three-world problem of AI programming, which involves Python at the high level, C++ for performance, and CUDA for hardware acceleration.
- Aims to unify programming across various hardware platforms including GPUs and TPUs.
- Features of Mojo:
- Compiled language (as opposed to interpreted), leading to performance improvements (up to 35,000 times faster in specific scenarios).
- Offers type annotations that enhance performance without sacrificing the dynamic nature of Python.
- Utilizes compiler technologies like MLIR, enabling it to communicate efficiently with various hardware.
Performance Improvements
- Mojo provides significant speed enhancements over native Python.
- Users can expect 10x performance increases for code run with Mojo versus standard Python.
- Performance gains can further increase with type annotations and additional optimizations.
Compatibility and Ecosystem
- Mojo maintains compatibility with existing Python libraries such as NumPy and Pandas without requiring rewrites.
- Users can seamlessly integrate Mojo into their current workflows.
- Community Collaboration:
- Chris emphasizes the importance of community involvement in developing Mojo, inviting programmers to contribute to its growth.
Challenges and Future Directions
- Complexity in AI Stack:
- The podcast discusses the challenges posed by the complexity of AI tools, which often leads to long deployment times for models.
- Modular aims to simplify this complexity, allowing more researchers and developers to participate in AI advancements.
- Roadmap for Mojo and Modular Engine:
- Ongoing development and improvements to the Mojo language.
- Building a robust community around Mojo for collaborative growth and shared knowledge.
Key Takeaways
- Mojo is set to revolutionize the AI programming landscape by making it more accessible and performant for Python developers.
- The Modular Engine provides significant performance boosts for deploying AI models, solving long-standing issues in the AI development lifecycle.
- Community-driven development is critical for the success of Mojo, promoting collaborative innovation in AI.
Conclusion Chris Lattner and Modular are working towards a future where AI development is simplified and more efficient, addressing head-on the complexities that have historically hindered productivity in the field. Mojo's unique approach as a supercharged Python offers promising solutions for AI researchers and developers alike.
For more information and detailed show notes, visit: [TWIML AI Podcast Episode #634](https://twimlai.com/go/634)
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:07All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington, and today I'm joined by Chris Lattner. Chris is the CEO and co-founder of Modular AI. Before we get into today's conversation, be sure to take a moment to head over to Spotify, Apple Podcasts, or your listening platform of choice. And if you enjoy the show, please leave us a five-star rating and review. Chris, welcome to the podcast. Hey, Sam, it's great to be here. It is great to have you on the show. The last time we got a chance to speak was, I think, back in 2020 around this time for the big great ML language undebate.
0:46That was a fun time. I think you've switched teams from a language perspective since we last spoke. It's pretty funny, the connection, right? There's a new contender. There's a new entrant on the field. How about that? There's a new contender in town, yes. And we will get deep into that conversation. but before we dive into Mojo, the new contender that we're speaking of and will be speaking of and all the work that you're doing on it, I'd love to have you share a little bit about your background, refresh our audience with you and some of the things that you've been up to. Yeah, sounds great. So I've been kicking around the software industry for a number of years now and have built and worked on a lot of different kind of low-level languages and compilers and other technologies in the developer tool space.
1:32I have a lot of fun with that and have been learning a lot. And so I'm most well known for open source things like the LLVM compiler, the Swift programming language, things like this. But I got interested in AI in 2016. And 2016, it feels like forever ago now, but at the time I felt like all the best work had been done. And it was just such a outrageous new approach to solving old problems. And so I just got into it deeper and deeper and deeper. And good news, not everything in AI is done yet. So I didn't quite miss the boat. But from there, I went through many different parts of the journey, worked on Google TPUs, TensorFlow, and a bunch of other things like that, built more production systems, worked on hardware, and have touched many different parts of this elephant.
2:15I bring a lot of experience with a lot of different parts of the stack, and we're trying to help lift AI to the next level. And at least a part of that is in developing and promoting a new language for AI, and that is Mojo. Can you talk a little bit about Mojo and its significance? Yeah, absolutely. I mean, I think that if you zoom out to understand what Mojo is, you have to understand where it came from. And so when we started Modular, our quest is to make it much easier to build, deploy, and evolve AI research. And so taking research, lifting it to new levels, and then getting that research into production.
2:52This is a quest that many people have been on for a really long time, but it's really about making this whole technology stack more accessible. and make it so more people can play in it so that experts at many different levels of the stack don't get stuck in one level. And one of the things, if you zoom into something like TensorFlow or zoom into something like PyTorch, you'll find is that many people work at the Python level, which is fantastic, and they know how to build models and things like this. But researchers who want to push the boundaries end up having to work at the C++ level. That's one of the dark truths of Python is that deep down underneath it, when you get down to things that care about performance or care about hardware, you quickly end up in C and C++ land.
3:29But AI is even worse and more challenged than most Python systems code because now you bring in GPUs and TPUs and accelerators and all this kind of stuff. And so now you end up in this actually three-world problem where you have Python, the high level, you have C++ in the guts, and then you have things like CUDA and other accelerator languages underneath. And so Mojo is a solution to this equation, right? Where at Modular, we're building and solving and tackling a lot of these old problems in terms of how do you get models to be expressed in a natural way? How do you map it onto accelerators and different kinds of heterogeneous, fancy hardware and the people you're coming out with?
4:07And how do you make it hackable for researchers? And to do that, you have to get rid of this three-world problem. And the stack we built is really novel and the way it works underneath the covers is quite unique. And so we needed a way to program that whole stack top to bottom. And so we needed one language that could scale. And so Mojo is kind of that, right? It starts from this requirement of let's pull together this three-word problem into something that is consistent. But then we needed a syntax. And so when we decided, okay, well, we have a really interesting and cool to compile a set of compiler technologies under the covers to enable all these accelerators and all this fancy low-level heterogeneous, all the technology stuff.
4:47We needed a user interface. And so as part of doing this, we said, well, you know, Python is the obvious thing, right? Python powers so much of AI, so much of data science in general. And so what we decided to do is build Mojo into a superset of Python. So that first of all, it feels like Python and it's accessible and Python programmers already know Mojo. But then we can also give Python superpowers where now Python can scale down and can be high performance and can run accelerators and can do these things that it hasn't been able to do before. Awesome. To what degree does the work you're doing with Mojo build on top of or depend on some of the things that you've done in your past lives around LLVM?
5:26Is LLVM an enabler for this new tech? Yeah, absolutely. So there's a number of different things that Modular and Mojo build on top of. And so you can say Modular is a fairly young company. We're about 18 months old at this point, but it's built on many years of experience building a lot of technologies in a lot of different places. And so a lot of the research has been done in other contexts. One of the pieces of that is this compiler framework called MLIR. MLIR is, you can kind of think of it as an evolution of LLVM that has enabled a new generation of compiler technologies. MLIR is now widely utilized across the entire industry for AI accelerators and has been very rapidly adopted.
6:06It's something that I and the team built at Google and then we open sourced and it's now part of this LLVM umbrella of technologies. LLVM, as you say, is also a really important part of the component stack. So LLVM is an umbrella project that includes things like MLIR, and it includes the Clang compiler for C and C++ that many people know about. But it also includes fundamental building blocks, like code generation for an x86 processor and things like this. And so we build directly on top of a lot of that technology as well. And so that's all kind of integrated into the stack, and that's one of the you-make-the-hardware-go-brr kind of things.
6:41And so that's all super important. And so when you think about, you kind of painted the picture of this three world problem. Every time you say that, I think of three body problem. It's a science fiction book and trilogy. Think of this three world problem and how as an AI developer who is trying to actually get work into production, you have to think about, kind of think really deeply in the stack. is the idea with Mojo that you want to make it easier to go deep in the stack, or do you want to make it more transparent to the user so that they don't have to go down in the stack and everything is just kind of working underneath without their kind of needing to switch boundaries?
7:24Yeah, so at Modular, we have a couple of different goals, right? So one goal is meet people where they are, solve today problems, build a faster horse, right? And so in that department, nobody wants to rewrite their models. They want the code to just work. And so they want new capabilities, but they want to fit within their existing ecosystem. Now, when you deploy a model, this is something that I think many AI practitioners don't talk about quite as much, or maybe the practitioners and the researchers don't have coffee enough, because where it's pretty well understood how to train a model, deploying a model is another completely different set of problems.
8:01And so you can take this in many different ways. One example of that is that Python is great for research. It's maybe not the best for production deployment at scale. And so many teams will end up rewriting their entire model in C++ just to get it to go. If it's a dynamic model, for example, language model. Now that, and there's a bunch of interesting work and there's a really smart people that do that kind of stuff. But why is it that we have to rewrite our production model, our research models to get them into production? That's really unfortunate, right? Yeah. And so we like things to just scale.
8:32And so one of the things that Mojo does is it's way faster. And also, if you use it the right way, you can also make it so it deploys without, you know, into a single A.O. executable and things like this. And so it has new capabilities that Python natively doesn't provide, which enables it to go much further. And so it can be useful that way. Now, another piece of it is we're building high tech, what we call the engine that powers AI. And we have the fastest inference engine that's unified across TensorFlow and PyTorch. now, right? And that engine is built entirely on top of Mojo. And so it's not just about building a faster horse and like enabling the existing use cases.
9:10It's about like unlocking this potential of this next generation hardware. And to us, like that's equally important, even though many people see Mojo as being, it helps out Python and that's what you can look at it as moving Python forward, but really where Mojo came from is working backwards from the speed of light of hardware. and so you know we talk about mojo can be up to 35 000 times faster than python because at the limit of what the hardware can achieve and mojo some people will see it as it looks like a faster python or a python that has no gill or a python that types enable performance or you know things like this but really about what can the hardware do how do we unlock the full potential and how do we do that in a way that python programmers have direct access to and when you look at it from that perspective you said python that has no gill that's like uh the interpreter lock or something like that.
9:58And it is one of many limitations that inhibits the performance of native Python. Yeah. I mean, I think that if you zoom into Python, I don't know how deep you are in the internals of Python. A lot of folks use Python, but they don't dig into it like I do. And so - I don't dig into it like you do, no. Yes. I think you're in the majority. And so folks that use Python know that it's maybe slow. It doesn't scale super well. It can't use all the processors on your machine without a lot of workaround and things like this. There's many aspects of the technology within the Python implementation that make that so.
10:35And so it has an interpreter. Interpreters are slower than compilers, generally. It has what's called the GIL. The GIL prevents effective use of multiple cores. The implementation within Python puts all of the objects on the heap in a very specific way. And there's a bunch of implementation details is going to how it works. Mojo is, I mean, interesting in different ways. First of all, it's compiled. Second of all, it gets rid of the global interpreter law. Third, it changes its representation. Fourth, it adds types. Like you can keep layering and all the differences here. But the consequence is that it really is a different animal.
11:10It has different characteristics in what the Python implementation provides. And so because it's a first principles programming language, it really has addressed a lot of the problems that Python users have felt as symptoms, but have not dug into, you know, why is Python this way? Yeah. You mentioned that it adds types. You know, one of the biggest things that's happened on the JavaScript side of things is the emergence of TypeScript as being kind of this JavaScript compatible language, but that is strongly typed. Is Mojo have that same kind of relationship to Python? Yeah, there's a bunch of very good analogies there.
11:49So TypeScript is super popular. A lot of people use it and it fits right into the JavaScript ecosystem. And so Mojo has a similar relationship to Python where it's a superset. It works with the existing ecosystem. All the packages in Python just work in Mojo, which is really important to us. And so we don't want to break the Python community. Many folks went through the Python 2 to Python 3 transition. It was really quite difficult in various ways. And so we don't want to relive that. And so you can look at Mojo as a Python superset. And so by doing so, you can pull forward all of the existing code and all that ecosystem into a Mojo world.
12:25There's a big difference, though. And so actually, if you zoom into Python 3 as it is today, Python allows you to add types. And those types, if you add them to your code, are there for some linter tools or checker tools that can identify bugs and can identify obvious mistakes in your code sometimes. But those types in Python aren't used and can't be used by the implementation for runtime. And so because of that, you can detect certain errors, but you don't get good performance out of that. And so what Mojo does is it kind of takes that next step. And so you can use the existing lowercase i to say it's an int and declare it as an integer that way.
13:00Or you can use capital I. And if you say it's capital I, that's a Mojo strongly typed integer and it's checked and required. And then also is used for performance. And you get 10x, 20x faster performance if you just add a few type annotations. And we have a couple of demos of that. Yeah, just kind of carrying forward that TypeScript analogy, what I've appreciated about it is like, well, a couple of things. One, you can add types without like fully buying into all of TypeScript and needing to know all that, but still get like a little bit of benefit without going all the way into kind of this new paradigm.
13:35And also when you are looking at code that you're not familiar with, that is kind of fully adopting the new paradigm, it's still familiar. Like you can kind of make your way through it without knowing that there's things that you don't know. If Mojo enables kind of that same level of flexibility, I would think that's a good thing. Yeah, well, so you come back to this two-world problem or the three-world problem, right? Where you have Python and Python lives on top of C++. So being a superset means everything you do in Python works in Mojo, right? So obviously types cannot be required because Python doesn't require types, right?
14:13And so that's all true. But in the traditional world of Python, if you run into performance problems or you need access to system software or low-level thingies, you have to go build a hybrid package where it's half C or C++, half Python. And so the value prop that Mojo provides is you can continue writing dynamically typed code. That's all good. But instead of switching to a different language to do high-performance, lower-level things, just as you say, you add a few type annotations, right? Or you use some lower level syntax within your existing code, and then you can put more effort in to get more performance instead of having to switch to a completely different language where the debugger no longer works on both sides.
14:53Got it. You mentioned that Mojo gives Python superpowers. That made me think of it. I'm probably not alone in this, that the first place I learned about the dunder functions in Python was from Jeremy Howard in the Fast AI course. There's probably a lot of folks listening who came across it in the same way. Are you accessing these superpowers through Python native structures like that? Or are they annotations? Well, first of all, what are beyond the ability to tap into lower level structures? structures, like what are some of the kind of superpowers or enhancements that Mojo adds, and then how are they accessed?
15:36Yeah. So, I mean, you mentioned Jeremy. Jeremy's been a huge influence on me personally. I mean, you can say, you can go back to saying like, why does Mojo exist? And a lot of that's Jeremy's fault just between us. He's been pushing for years specifically for hackability, researchability. Jeremy's got the unique kind of brain where like the whole problem fits in his head. And so he can understand all the different parts of the problem. So yes, so Mojo has all the Dunder methods. And so if you want to add, want to make the plus operator work, you can implement the underbar underbar add method and things like that.
16:06But then it goes a little bit further. And so if you look in the space of system programming languages, you enter, you enter the realm of things like Rust and C++ and like these kinds of languages, right? And the systems programming world for a long time has been pushing towards bringing safety into this world. So C, C++, you have a pointer, pointer dangles, bad things happen, your app crashes, you have security problems, all these kinds of things. Rust and Swift and other languages like that have gone further into making it possible to get good performance without sacrificing safety. And so we've brought a lot of those ideas directly into Mojo.
16:41And so in Rust, there's a notion of lifetimes and ownership and these kinds of things that enable safe pointer usage and things like that. So Mojo brings that in. Now, these are features that, obviously you don't have to use unless you're writing low-level code and you care about getting high performance in certain use cases. But having that available gives you a very accessible whole stack solution that allows you to go all the way down and get Rust-style performance out of a CPU. And similarly, we talk about this hardware stuff. Well, at the bottom, even on a CPU, you have many cores, you have these crazy vector units and matrix extensions.
17:16It's really interesting to see the evolution of hardware because if you go back 10 years ago, It used to be that there was a CPU thing and a GPU thing, and these were points in the space that were very different, and they were completely unrelated from a hardware perspective. But today, that whole line has gotten blurrier because GPUs have gotten more programmable. CPUs are getting more AI stuff in them. CPUs these days have Bflow 16 and all these other AI things that are being built right in. And so we're getting a spectrum of programmability. And so a lot of what Mojo is about is unlocking that for people and making it accessible and making it so that, again, you don't have to switch languages just to get access to this stuff.
17:53And you're rightly focusing on CPUs and GPUs, but there's a wide variety of other options and perspectives, TPUs and other newer and more specific, more exotic. That's a great word. Yeah, exactly. Approaches to this. Are you building Mojo such that it is anticipating all of these options? Or is, you know, when you're focusing on making Mojo better use acceleration, are you really talking about, you know, GPUs or maybe GPUs and TPUs? So I spent a couple of years working on Google TPUs. And Google TPUs are, I mean, they're an impressive set of technology and machines because they scale up to exaflops of compute.
18:33They're highly specialized for AI workloads. They're also internally really weird. And so to plus one, exactly what you're saying, right? AI isn't just about a GPU, right? I mean, so much thinking around AI technologies, okay, I just need to get the GPUs lit up and then go. But particularly if you start deploying, well, if you're running on a smart camera or something, the AI chip is going to be completely specific to that camera, right? If you're doing Google scale training on crazy distributed machines, like that hardware is quite different. And so this is where one of the things that's, I think, very exciting to me as a technologist about Mojo is that it's built on this MLIR compiler.
19:11So MLIR is, again, this thing that we built, started back at Google. Now it's being used by basically the who's who of all the hardware industry. And MLIR talks to all of these things. And so if you're familiar with LLVM, LLVM is now a 20-year-old technology. It's widely adopted and talks to all the CPUs and some of the GPUs. But LLVM has never been successful at targeting AI accelerators and video optimization engines and like all the other weird hardware that exists in the world. And that's the role that MLIR provides. And so Mojo, one of the ways that it's implemented is it fully exposes that power and brings all the nerdery that goes into the compilers and exposes up to library developers.
19:51And so it's actually quite important that you can talk to, for example, TPUs or other things like that in their native language, which in the case of a TPU is this 120 by 128 tile. And being able to expose that out in the language is really quite important. So anyways, that's a long way of saying, yes, it is more than CPUs and GPUs. Though CPUs and GPUs are the starting point, obviously, for lots of really good reasons, but we've built this thing to have really long legs that can bring us into the future. And do you see it extending to things that are even more exotic, like your graph cores and SambaNovas and like the things that take a very different approach to the underlying compute?
20:28So let me bring you back to where Modular is coming at this because Mojo is one of the components of the AI stack as we look at it. So Modular is building what we called a unified AI engine. And so this unified AI engine, what the heck is that? Well, it's an engine. It's an engine. It's not a framework. And so people are familiar with PyTorch and TensorFlow and these machine learning frameworks provide APIs. And so you get NN module and the APIs that we're all familiar with. Underneath the covers, there's a whole bunch of deep technology for getting things onto a GPU, getting things onto a CPU.
21:02And so PyTorch 2 just came out with this Torch Dynamo stuff and all these exotic level technologies that make the hardware work. On GPUs, CUDA is a major component of the technology stack that everybody builds on top of. And so our engine fits at that level of the stack. And the cool thing about it, particularly when you're deploying, is that it talks to lots of hardware. It also talks to both frameworks. And so when you're taking a model from research, for example, you have a nice PyTorch model, you get off hugging face. Lots of people do this, of course. You want to deploy this thing. Well, you don't actually want all of PyTorch in a production Docker container.
21:39You want a low dependency, efficient way to serve the model. And so that process of getting from PyTorch and into a deployment thing is what the modular technology stack can help with. Now, as you say, coming back to answer your question, Graphcore, Salmonova, all these hardwares can't talk about any relationships. But from a technology perspective, they're all slightly different in high-level ways. So Salmanova's chip is, from my understanding, what's called a CGRA, which is a super parallel, really crazy thing that has almost nothing to do with CPUs. Graph cores are apparently lots of things that look like CPUs, but their memories are all really weird and different.
22:23And the way they communicate is very structured. And we all know CPUs and GPUs. And so what our technology stack enables is if you're the Sominova or Cerebrus is another example of a really crazy system, those people need to implement a compiler for their chip. And so they're the experts on their chip. They understand how this works. And what Modular can do is provide a thing for them to plug into so that they get all of TensorFlow and PyTorch. And one of the major problems we have today with hardware accelerators, particularly ones that are not the dominant player in the space, is that their tools don't actually just work.
22:59I'll pick on Apple, for example, right? So Apple has a deployment technology called CoreML. CoreML talks to the neural accelerators, and they have all this amazing hardware on a Mac or an iPhone. But CoreML is not actually compatible with all the models. And so getting something onto an Apple device means fighting with this translator and trying to get it to not crash, you know, doing all these things that the production world struggles with. And I talk with many people, many leaders at software companies that are building AI into their products. And a lot of software leaders, they see the symptoms.
23:31They see, OK, it takes three months to get a model into production. They see symptoms like I need a team of 40 people to be able to deploy things. And they're very expensive, very specialized people. Why is it this hard? Right. And the answer to those questions are that the tools, the technologies are not anywhere near the tools and technologies used for training. And so there's so much suffering, so much, so many problems in these things. And the root cause is the technology I've been working on for years, which is for any one of these chips, people have had to build an entire technology stack from the bottom up.
24:04And there's very little code reuse across hardware and hardware vendors. Again, I'll pick on Apple, but I love Apple. So it's not out of anger. It's very difficult to track the speed of AI. PyTorch moves super fast, right? This is stuff that you need a very dedicated team. You need to be super responsive. You need to be on top of this stuff. And also the compiler problems and the technology problems to make the hardware work are really difficult. And so there have been a lot of really smart people working on this, but if you're always focused on getting the next ship out the door and you can't take a step back and look at this whole technology stack, then you can't make the leap that modular is driving forward.
24:40Interesting. So you said something earlier, kind of describing the engine and its place. And it made me think of for ages now, right? We've kind of decried the kind of stranglehold, if you will, put a negative spin on the CUDA has on like the low level programming interface, which basically kind of ensures that NVIDIA has long-lasting position and makes it very difficult for, you know, say an Intel to come out with a CPU with some numeric capabilities and displace it because there's all this, A, there's all this code that's been written in these three worlds that you've mentioned, and like it's not as easy as just swapping out the hardware, right?
Read the full transcript
25:20Absolutely. Are you envisioning that this modular engine is this kind of replacement for CUDA that is multi-hardware capable? Is that the core idea? Yes, I mean, that's one of the value props we provide. So if I zoom out and look at the steps the industry has been going through, so we as an AI industry owe a huge debt of gratitude to CUDA. If you go back to the AlexNet moment, for example, a lot of people talk about it was a confluence of ImageNet and the datasets and things like this. It was a confluence of hardware. And the fact that GPUs enabled an amount of compute that could cause AlexNet to happen.
25:59But a lot of folks forget that Kuda was what enabled some researchers to go write convolution kernels and actually get a machine learning model running on a GPU, which the hardware is definitely not designed for back in the day, right? Today, AI has taken over and it's a little bit different. But back in the day, that initial breakthrough was really in large part thanks to Kuda. And so one of the things that's happened is that as AI has taken over, right, a lot of technology has been built on top of Kuda. It's a very good thing and it's very powerful and flexible and hackable and it's great. But as you say, it's kind of put us into a mode where one vendor has this dominant position and if you're a hardware vendor at even an AMD or some other widely known company that has really impressive hardware to be able to play in this ecosystem.
26:42Now, what's happened and one of the things that led into the thinking that went to modular existing is that there have been a lot of compiler technologies that have been built. For example, there's this XLA compiler that I worked on at Google. There are new compilers every day being announced by different companies where they're saying, I will build a compiler that will make ML go fast, for example, on GPUs. And so several years of work, lots of cool technology, lots of examples of these systems exist, and the names keep changing, but the technology is very powerful. The problem with that is that they have lost one of the things that made CUDA really powerful, which is the programmability.
27:19And so what has happened is the compiler nerds, which I'm a member, so I love the compiler nerds, but those compiler nerds have went and turned AI code generation and things like this into a compiler problem. But that has excluded all the non-compiler people, right? And so if you look at TPUs, for example, TPUs can express everything you can do in this XLA compiler. And so it can do matrix multiplications, convolutions, element-wise ads, et cetera, et cetera, et cetera. But it can't do sparse operations, can't do data operations, can't do pre-processing. And so AI, you're an expert, you know this, AI is not just about matrix multiplication.
27:58It's about data loading, pre-processing, this full parallel compute problem that is part of AI. And so what has been lost over the several years of trying to solve the CUDA lock-in problem is that people have tried to make this a compiler problem. and now you've turned into a different lock-in. But instead of locking into hardware, you're locking most smart people out of the ecosystem. And these compilers haven't been super successful of being compatible with code and things like this, right? And so what Modular is doing is we're saying, okay, again, I love all these people. I've been working on this stuff for a long time myself.
28:33But what we're doing is saying, start from a different perspective. What is our assumption? Our assumption is people don't want to rewrite their code. What that means is you have to have all the operators, all the systems that go into something like TensorFlow or PyTorch need to work. Thousands of operators each, and it's a really messy job, but we handle that job for the world, right? The other thing we say is, okay, PyTorch is really popular in research. TensorFlow is still quite popular in production. What we see out in the industry, again, every shop is a little bit different, but a lot of people have both TensorFlow and PyTorch.
29:03And so they don't want to have this bifurcated stack built on top of these things. They want to actually have one system that they can scale out. And so we make our problem even more complicated by building a unified solution. And so now it's not about 2 ,000 on the TensorFlow side, 2 ,000 on the PyTorch side. It's about 4 ,000, right? And it's actually even worse than that when you bring in some of the other technologies. But now you talk about hardware, right? It's not just about Intel CPUs and NVIDIA GPUs. It's this other access that then does a multiplication to this whole problem and says, okay, well now I have many different, there's probably a hundred or a thousand different kinds of hardware.
29:41And so where traditional teams have built a point solution saying, okay, I'm going to build a fancy compilery thing for one hardware, for one framework and, you know, in one direction along this. And they built one of these often very good tools, but they're very purpose built in one case. You know, we're having sympathy for all the software people that have to deploy because software people, they don't have one piece of hardware. They don't have one model. They don't have one framework. They don't have one product, right? Their products evolve over the course of decades sometimes. And software lives a long time.
30:13And so they need to be able to talk to lots of different generations of this stuff. And so at Modular, what we've done is we've said, okay, well, this is suddenly a very different problem from a technology perspective than building a point solution. And this problem, this I need to solve this massively complicated space where you have hardware on one side, you have the sheer scope of AI on the other space is what drove Mojo to exist because we need a way to make this entire stack accessible, hackable, understandable to people that are not themselves compiler engineers. We need people that know really fancy numerics and sparse algorithms and convolutions and, or people that know their hardware.
30:52We need to know like all these people that are involved in all of this massive technology stack that we've been building to be able to collaborate and work together and build cool stuff at a high velocity, right? And that's where we think that Mojo is really interesting, because as far as I know, nobody's done that. I mean, it's like a completely unique creation in the space, and we hope that will really simplify the world. One of the things we kind of joke about it, that our biggest enemy, you know, the mortal enemy that we struggle with at Modular is actually just complexity, right? In the AI space, there are so many systems, so many technologies, so many layers of stuff that has been built up.
31:29And if you zoom out, coming back to 2016, I thought I was too late to do anything important in AI. What you realize is that AI is still not done, right? The stack you're building on is adolescent. It's in its teenage years. And so what we need is we need to get to that next level where everything actually works way more predictable. It's actually hackable. When you try and experiment as a researcher, the tools don't break out from underneath you. And when you achieve that, we think that the impact of AI can go much further and that many more people can participate. When you talk about the complexity and diversity of underlying components, and then you talk about kind of how the lifespan of software kind of extends over generations of underlying infrastructure, it makes me think of like dependencies and dependency management and packaging and all these things as like huge problems that need to be solved is, does that play into what you're doing at all?
32:25Not directly, but your pattern matching, your neural net there is doing a very good job of pattern matching and seeing what we're talking about here. The packaging problem is often because you have all these incompatible systems that are lashed together. And so if you zoom into Python packaging, I mean, there's a lot of things going on there. I'm not an expert in Python packaging. People I talk to that are, a big part of that is because of the C parts of these Python packages. So you pick our old friend NumPy, for example, right? NumPy has a ton of C code inside of it, as well as the Python API.
32:56Well, packaging that means you're not actually packaging Python, you're packaging C code. C's never had a package manager that's any good, right? And so it's funny, you look at these old problems that we've been struggling with. Well, you get rid of the C code, and suddenly packaging is way simpler, right? And so this is one of the things that Mojo provides, is providing unified language. And more generally, every time you see one of these fissures, like you're talking about the hardware divide. Here we're talking about Python C++. Talk about CUDA versus SICKL versus HIP versus like all these other crazy things that exist in the world.
33:28Like each one of these things is at the bottom of our stack driving complexity up. And so at the end of the day, you'll have a researcher who very reasonably says, hey, I just want to run this model on an AMD GPU. No big deal, right? You should flip a switch, right? But the problem is that at the very bottom, all this stuff is very different and all the cracks go up. And if you take reliability and it's 90 % reliable, and then the next step is 90 % reliable, the next step is 90 % reliable. You start multiplying together all the 0.9s and you get something that's 10 % reliable, right? This is the AI stack that we all depend on.
34:02And you've got this easy problem, which is, well, okay, let me be careful here. You've got this one class of problems that is very challenging, but it's easy to deal with. And that is when you're trying to use all this stuff together and it just doesn't work. Like it doesn't compile or it doesn't run or whatever. But then you have this other problem where it works, but you don't know that it's actually not working. Yeah, exactly. Because of like semantic differences or what have you, it's either not performing well or you're not converging, your results are out of whack. And like you're digging deep into underlying libraries, trying to figure out like, why are your answers like crazy?
34:43Yeah, I give you one example, right? I mean, just go through the lifecycle of deploying a model. So to just make up a scenario, but to just double click on what you're saying, okay, I want to deploy a model. Well, now I need to get it to go through Coromel or one of the many things for deploying to some piece of hardware. Results don't work. Well, now to just plus one you a hundred times. Now you need to know not just PyTorch, not just your model, not just Coromel, but also the translator, also all these things. And you dig in and dig in, dig, dig, dig, dig, dig, dig. And you find out it's handling the edge padding on a convolution slightly differently.
35:15Right. and so now wait a second so like all of these tools were supposed to be making it easy but because they're not all reliable like it's this leaky abstraction now you have to understand all of this complexity right this is what causes it to take three months to deploy a model fundamentally this is something where i think that many folks that are building ai products and they're managing they're the vp of software some technology company right they just see the symptom of why is it take so long to get this model in production? But they don't realize that the tool set, this fundamental technology that all this stuff is built on top of, it's not up to the standards of a software tool set.
35:52No C programmer would tolerate AI tools and their quality. It's just crazy. But again, this is just the maturity of the AI technology space. And by solving that problem, what we want to see is like way more people, way more technology, way more inclusion in the kinds of companies that are able to work with AI and do things. And we think that'll be a really big impact on the world. We've talked about Mojo. We've talked about this inference engine or the engine that we've referred to. In the context of Mojo, you've talked about like 35 ,000 X performance improvements over a standard Python. Do you need the engine to get that level of performance improvement?
36:31Switching, using Mojo, like lock you into using this engine? Like what's the business model there? Do you have licensing issues? I have a bunch of questions kind of coming out here. And they span kind of technical and like business licensing kinds of questions. How does all that work? Great question. So you've identified the right players. There's Mojo, which is a programming language. It's a programming language that's a member of the Python family. It's really useful on, for example, just CPUs, which is the only place that Python plays. And so many people see Mojo as just being a better Python.
37:05Now, we have the engine. The engine itself can stand alone and you can use the engine as a drop-in replacement. It works with TensorFlow PyTorch. It'll make your BERT models go 3X. And you're using it as a drop-in replacement for what exactly? For a traditional TensorFlow implementation. Okay. Before I answer your bigger question, let me dive into that. Yeah. So what the modular engine does is you replace the TensorFlow with our TensorFlow or your PyTorch with our PyTorch or if you're using TorchScript or things like this, and so you just put a new thing in your Docker container. Got it. And what you get from that is massively better performance.
37:42And so, you know, TensorFlow is quite good at production, but we're showing 3 to 5x better performance on, for example, an Intel CPU or an AMD CPU or an ARM-based Graviton server in AWS. And so you think about that and you see 3 to 5x better performance. Well, that's a massive cost savings. 3x less GPUs that you need to... Exactly. That is a massive cost savings. Well, and it's also a massive latency improvement. And so many of our customers love that because then they can turn around and make their models bigger. And so now you can have a better product for your customers. And so you get direct impact on your costs, direct impact on your product.
38:22And this is a huge deal for people. I'm a technology nerd sometimes, right? And I love how it's built, but the impact on products is phenomenal. And the engine is a really big deal for getting production AI to scale. So just kind of continuing down on that line before we click back out, then I would imagine one of the commitments that you need to be making to folks that are thinking about using this thing is how close you're going to stay to the development of that stack, right? Yep, absolutely. Well, so I mean, one of the things also that customers love is that Google and Meta don't actually support TensorFlow or PyTorch.
38:58right these people forget but these are not products right these are open source projects they are hobbies maybe for the megacorps and so you're essentially offering like supported performance optimized version of tensorflow and pytorch absolutely and right but then if i'm going to think about using this i need to know that i'm not going to get left behind like you're gonna you know i'm going to wake up one day and i'm three versions behind the latest thing in tensorflow and it has something that I need in order to make my 500 trillion parameter LLM work. Yep. So, I mean, we're committed to doing that.
39:32So I don't know if this is like a binary question, but yes, we do that. But the thing that the enterprises we talk to that care about their costs, right? Often they want somebody that they can call, right? And if you think about it, right, it's analogous to who wants to run a mail server themselves, right? You can run SendMail or something, right? But nobody in their right mind does that, right? Why do we do this with AI infrastructure, it's because there's no choice. There's been nobody to reach out to, nobody that actually can do this. And the thing that I think many folks forget is that Meta and Google, their technology platform has diverged a lot from what the rest of the industry uses.
40:05So they both have their own chips they build, right, for example, right? And they have their own specific use cases. And so they're not actually focused on making the traditional CPUs, GPUs, and public cloud use case actually really good. That's one of the reasons why we have such high value we can deliver. So yes, this is a product for us. That means we actually support it. That means we invest a huge amount of energy into it. This is one of the reasons why we have such phenomenal results as well. And also to your other question, one of the great things about being a drop-in replacement, from a customer perspective at least, is that it means you can un-drop in.
40:37You can use our technology and if you want to switch back, you can always switch back at any time. And at some point we'll make it back to that broader question, but we've talked about Mojo as being this better Python, but what makes Python usable in AI is not just kind of the core Python. It's all these other things, NumPy and pandas and many other packages. You mentioned, you know, we know they have C at the heart of them. So at some point, there's a significant number of packages that you also have to kind of rewrite that need to be Mojo native, I would think, in order to get the full performance?
41:14Let's dive into compatibility. So Mojo is still a young language. We haven't talked about that, but it's still not, it's not done. And I think it will take another year or so of development before it gets to be like solving all the world's problems that we want to solve, things like this, right? But even today you can import and use arbitrary packages like NumPy, Pandas, TensorFlow, PyTorch, whatever directly into Mojo. And so a really important part of how our stack works is you don't have to rewrite all of your port or touch all of your Python packages. I mean, many people have their own Python code.
41:43It's not just big packages like NumPy, right? And so Mojo talks directly to all those packages. You don't have to write wrappers. It all just works. This is a really big piece of that. Now, if you choose to move your code into the Mojo universe, then you can get the benefits that Mojo provides. And so if you're just talking to an existing package, well, it'll still run at Python speed. It will be fully compatible, but it will also run with the same implementation, this default Python implementation. And so moving your code to Mojo can then unlock these new capabilities, but then you can choose to do that a package at a time or however you'd like to do that.
42:17I guess I'm curious, like how much of the like surface area of AI related packaging have you built? Or am I thinking about this the right way? Like in order to fully provide the performance benefits that you're talking about, did you need to port NumPy over to kind of a Mojo native or to run on MLIR or whatever, at whatever level that makes sense? Did you pandas all these other like, how much did you need to do and how much of that is done like percentage wise relative to what you expect will need to be done to be? Yep, absolutely. Well, so the answer is zero. So our solution is like our solution enables to talk to the entire Python ecosystem out of the box.
43:02So matplotlibscipy, numpy, all that stuff just works. And that, again, come back to being pragmatic and productive. Like, we can't, I'll make fun of you, I'll make fun of me from our last call on the great language debate, right? The problem with any new programming language is that a new programming language has no community, has no package ecosystem, right? And so, again, like myself on that previous call and all the other lovely people there, right? You want to get ML out of Python for whatever reasons is very exciting, but it's not very pragmatic because the entire data science ecosystem is all wrapped around Python.
43:36And Python's also pretty great, right? I mean, I think that that's something that people in other communities like to make fun of Python because of indentation or whatever it is. But Python's beautiful. Subjectively, I'll say it's my opinion. And so what Mojo does is enables you to use literally everything in the Python ecosystem. And then if you want to invest more effort to get more performance, then you can do that, but you don't have to. This is the major value problem. Now, in the case of Modular and why we built Mojo, like business objective of his go make ML really awesome. We care about the matrix multiplications and the convolutions and the core operations that people spend all their time on in AI.
44:14And so we rewrote all of that stuff in Mojo. And so this isn't rewriting Matplotlib, this is like rewriting Intel MKL equivalent, right? Or rewriting the CUDA implementation of these CUDA kernels equivalent, right? And so that's where we've put our energy into because that's what enables unlocking of the hardware, enables unlocking of performance, enables unlocking of usability. We have really exotic, fancy compilery features that enable kernel fusion, automatic kernel fusion and things like this that no normal ML researcher should ever have to know about. They just see, okay, it runs 10x faster in this use case.
44:50Well, that's pretty cool, right? And another thing that I think that folks are struggling with is that take transformers, for example. I mean, you know transformers, I know transformers. We all love transformers. They're eating the world. But one of the problems with this is that because it became so important to so many different use cases, we got all these very hyper-specialized software stacks for transformers. And so these exist at the low levels. So NVIDIA, for example, has a set of kernels called Faster Transformer. These are at the high levels. And so there's all these distribution frameworks for transformers and things like this.
45:22And so you get this very transformer-specialized stack, which again, forces you into this very narrow view of what a transformer is. And it works for the benchmark. But if you're a researcher, You want to go push the boundaries and try slightly different transformers. Or, you know, maybe there's a thing beyond transformers. Like I hear that RNNs are coming back in and maybe FFTs will have their day, right? I mean, there's like all these different theories. And if we can't enable people to do that research, right, we may be missing out on that next big step. And so the specialization that's inherent in things becoming important really cuts against generality.
45:58That's one of the things that we've seen and that we really want to like, again, like if you dramatically reduce the complexity in these stacks, you can make it way more hackable. And that we believe will enable people to invent new things. Yeah. I want to push on this one more time just to make sure, see if I can figure out what I'm, expose any kind of fissures in my understanding here. Is what you're saying that, or is it the case that, you know, in thinking about the relationship between like, NumPy or Pandas and Python, that those libraries that we all use as part of, that are kind of ubiquitous from a machine learning perspective, I can imagine a couple of things.
46:40One, they're sitting on top of the underlying, they delegate enough of what they're doing to the underlying Python that you kind of replacing, fixing that underlying Python gives you some percent of the performance benefit such that you don't need to deal with the upper piece. Who is you here? So are you asking how it works internally? Are you asking how a user uses it? Are you asking when somebody should do something? Which piece of this elephant are you touching? I'm primarily trying to make sure that I understand how it is that you're able to offer the performance improvements that you're boasting without needing to touch any of the libraries that people depend on.
47:32And so I'm kind of asking about internals, but also like how they're used. So I'm imagining like several scenarios. A, for whatever reason, A is like your 35 ,000 number. That's kind of a made up number that doesn't actually rely on any external dependencies. And it's kind of a useless performance boasting metric. That's one possibility. Another possibility is NumPy delegate, you know, these libraries delegate enough of their operations to the underlying Python that you can get the significant performance gains even without touching those things. And hey, if somebody did touch those things, maybe it would be 70 ,000 or whatever.
48:13I can break it down for you if you want. Got it. Okay. Let me break it into a couple of categories. So one is you have unmodified Python that is just imported. So you take map plotlib, just to pick on something that's not performance sensitive. There's no reason to rewrite map plot lib. It's fine, right? And so you just import it. If you import it, the way that runs is that runs with the existing CPython interpreter and Mojo talks to the CPython interpreter. And so that code runs 100 % compatibility. Everything just works great. And this is why the entire ecosystem works, but it's no faster. And so really what you're getting as you're getting at, what are my trade-offs?
48:49What are the levers I'm pulling here? And so full compatibility, but no performance benefit. Those things go together. Another thing you can do, if you go to modular.com, you can see our video and you can see Jeremy giving a demo, Jeremy Howard giving a demo. And there we can see is you say, okay, I just take some Python code I put into Mojo and now it runs, you know, depends on the code, but you know, roughly 10X faster out of the box, maybe 15X, 16X. I mean, there's more we can do to push it further. We just haven't focused on that. And that's running the same code, but in Mojo. And the reason that you get performance is it's compiled instead of interpreted.
49:21It has a new fancy compiler stack, all this stuff under the covers, but it's still running fully dynamic typed code. It's just running dynamically typed code in a better way. And so you can get 10X out of the box. It's pretty good. That's quite nice. Then you start layering in and saying, hey, I want to add types. Okay, well now you're talking about like changing the in-memory representation of this thing to be way more efficient. Well, that's 10X. Now you say, give me threads. Okay, well that's 10X. Okay, now I want to use vectors and new hardware. That's another 10X. And so if you stack all these things up, This is where you get into 35 ,000 times.
49:54And I will agree with you, by the way, that the 35 ,000 number is a cherry-picked number. This is an extreme result on Manelbrot, right, which is a simple algorithm we can explain and people can play with in a notebook and stuff like this. But we have lots of people, just random people on the internet using Mojo that are getting hundreds and thousands of times speedups. And so the 35 ,000 may be cherry-picked, but reasonably expecting getting over 100x is... 100x is pretty big. I consider that to be a pretty big deal. Yeah. Right. And you can look at that as 100x over Python, or you can look at that as saying Python's now 100x more relevant for keeping me out of C.
50:31Both of those sides of that is really cool. Now, anyway, so coming back to your categories. So you tie together. Let me add that third category, or maybe it's more a scenario than a category. It is also, I think, largely the case, but maybe you can validate this for me. You're probably using a lot more of these libraries when you're doing EDA and the early stages of building a model. But then you finally have your model in the form of a graph in TensorFlow or PyTorch. And, you know, at that point, the things that you're relying on are kind of much lower level as opposed to like your pandas and your scipy, all this kind of stuff.
51:11And so your exposure or your need to pull in all these libraries at the kind of the point where you're in the kind of core training loop or inference is less. Yeah, you want to like deploy the model. And so if you zoom into the third part, let's just call, because the things we just talked about are actually completely generic software engineering things, right? We talked about using arbitrary Python package off the shelf. We talked about take Python code, arbitrary Python code in an arbitrary domain and just make it go fast, right? Which is fun, right? But it has nothing to do with AI. Now let's talk about AI.
51:46So AI, third category, super important. And it turns out many of your readers or your watchers want to think about AI, right? And so AI is this really fascinating technology stack that, yeah, you talk to it in Python, but underneath the covers, you have kernel fusing graph compilers and all this accelerators and like all this other cool stuff, right? Yeah. And so this is where the modular engine comes in, right? And Mojo is an implementation detail in the modular engine. And Mojo makes it all super extensible and hackable. But this technology space actually really has nothing, very little to do with syntax or with a programming language.
52:20It's a completely different technology stack that's much more similar to these XLA compilers and the internals of CUDA or the internals of Intel MKL or these kinds of things. So these are all different. But come back to your basic question, several layers up in the stack, which is what is the relationship between Mojo and the modular engine? Because that's also really important. So the modular engine is really focused around high performance, production deployment, go solve problems in AI. And so it's an AI thing. Mojo as a language is actually a new member of the Python family. Yeah. Right. And so for modular, we see the engine as being a product and we see Mojo as a technology.
53:00And both of these things stand alone. So you can use Mojo as just a better Python if that's what you want to do. Or you can use the modular engine as it drops into TensorFlow and PyTorch, and then you just have a better TensorFlow and a better way to deploy your models. But there's a much bigger and hopefully, I believe, much more important in the long term, better together story here, right? Because putting a custom op into TensorFlow or PyTorch is very difficult. We talk about the three-layer problem, right? The Python C++ CUDA. Well, if you want to put a custom CUDA op into PyTorch, you have to write C++.
53:37CUDA is like a C++ thing, right? It's a C++ thing that doesn't have a debugger. A C++ thing with a whole bunch of weird constraints where you might wedge your GPU. And that complexity makes it so people don't do the kind of research that they might otherwise do. And obviously, if you have to hack C++ or even if you just have to rebuild TensorFlow, who in their right mind knows how to do that? I know these people and I love these people, right? But this is monstrous, right? And so the better together here is that if you're an AI person, you're building, deploying models, you're training, you're doing research.
54:08Well, what Mojo inside the modular engine allows you to do is make this whole thing hackable. So you can define custom ops, so you can get kernel fusion, so you can get all this stuff for free. And then when you want to go push boundaries, you can go crack open the box and say, okay, I'm going to write a custom sparse thingy for my domain or a custom summary function that does some fancy domain specific reduction before I send all the data across the wire. and making that possible is, I think, really cool. Interesting. Yeah, I feel like we've covered a lot and there's still a lot to cover, particularly in this dimension of hackability, but we don't have time to cover all that.
54:43To kind of wrap things up, I'd love to have you maybe riff a little bit on future directions, roadmap. What are the big things that you need to attack next to kind of build out this vision? Absolutely. So Modular just came out of stealth. And so we have a nice video on our website at modular.com if you haven't seen it. There's a whole bunch of new drops that we'll be adding to the product over the coming months. And so you can sign for a newsletter on that. The thing I'll say is that Mojo is still quite early. And so it's not ready for production use as a general drop-in Python replacement. But we have an amazing community of people already coming together and we're developing it in the open.
55:22This is, I think, a pretty big deal for something that I hope will be important to a wide range of different use cases. I mean, Python goes everywhere, right? And so I think it's really important that we as a community build and do this together. And Modular is obviously driving this because it's really important to us, but we don't have all the smart people in the world. And so I'd really love for people to join us on our Discord forum and other places where we can interact and build this together. Awesome. You mentioned that Jeremy was a big inspiration. I'm glad he wasn't able to inspire you to ride it around Pearl.
55:54Yeah. Again, he's an incredible person. So he gives a killer demo in our launch showing how to take matrix multiplication. And he doesn't get it 35 ,000 times, but it's 20 ,000 times or something in a notebook, which is pretty cool. That's awesome. He's an incredible person. Don't underestimate Jeremy. Absolutely. Well, Chris, thanks so much for getting us up to speed on what you're working on with Mojo and Modular. Yeah, it's really great to talk to you, Sam. Same. Thank you. all right everyone that's our show for today to learn more about today's guest or the topics mentioned in this interview visit twimla.com of course if you like what you hear on the podcast please subscribe rate and review the show on your favorite podcatcher thanks so much for listening and catch you next time
From the publisher
Today we’re joined by Chris Lattner, Co-Founder and CEO of Modular. In our conversation with Chris, we discuss Mojo, a new programming language for AI developers. Mojo is unique in this space and simplifies things by making the entire stack accessible and understandable to people who are not compiler engineers. It also offers Python programmers the ability to make it high-performance and capable of running accelerators, making it more accessible to more people and researchers. We discuss the relationship between the Modular Engine and Mojo, the challenge of packaging Python, particularly when incorporating C code, and how Mojo aims to solve these problems to make the AI stack more dependable.
The complete show notes for this episode can be found at twimlai.com/go/634




