Modal and Scaling AI Inference with Erik Bernhardsson

31 Jul 2025 · 40 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Summary: Modal and Scaling AI Inference with Erik Bernhardsson

Episode Overview

  • Podcast Title: Software Engineering Daily
  • Episode Title: Modal and Scaling AI Inference with Erik Bernhardsson
  • Description: Erik Bernhardsson discusses his company, Modal, a serverless compute platform for AI workloads that allows rapid iteration and autoscaling of GPU-enabled containers.

Key Participants

  • Host: Sean Falconer
  • Guest: Erik Bernhardsson, Founder and CEO of Modal, former CTO at Better.com and a key figure at Spotify.

Main Topics Discussed

  1. Background of Erik Bernhardsson
  2. Seven years at Spotify focusing on data and AI, including building the music recommendation system and the Luigi workflow scheduler.
  3. Realization of tooling gaps in AI and machine learning led to the founding of Modal.
  1. Motivation and Genesis of Modal
  2. Identified a market gap in ML and AI tooling during his tenure at Spotify.
  3. Aimed to create a tool that simplifies the deployment and scaling of machine learning applications, addressing inefficiencies in developer productivity due to slow feedback loops.
  1. Developer Productivity and Feedback Loops
  2. Highlighted the lag of AI/ML engineering tools compared to traditional application development.
  3. Emphasized the importance of fast feedback loops in enhancing developer productivity, contrasting the dynamic nature of front-end development.
  1. Modal's Core Features
  2. Serverless Platform: Focused on running data and AI workloads, enabling fast deployment of GPU-enabled containers.
  3. Multi-tenant Model: Shares resources among users, allowing for efficient scaling and lower costs through pooled demand.
  4. Usage-based Pricing: Charges only for the compute time used, eliminating the need for capacity planning.
  1. Technical Infrastructure
  2. Developed a custom file system and container runtime to achieve quick container cold starts, enabling fast deployment and scaling.
  3. Functions run in isolated containers, allowing for diverse configurations (e.g., different Python versions).
  1. Use Cases for Modal
  2. Predominantly driven by generative AI applications including:
  3. AI-generated music
  4. Computational biotech (e.g., protein folding)
  5. Geospatial analysis
  6. Noted that Modal also supports smaller projects like web scrapers.
  1. Challenges and Future Directions
  2. Plans to improve low-latency processing for real-time applications and to decentralize the control plane for faster decision-making.
  3. Future work includes enhancements in distributed training capabilities for more efficient model training.
  1. Perspectives on Security and Compliance
  2. Discussed how multi-tenant architectures require robust security measures and the evolving perceptions of cloud security.
  1. The Evolution of AI Infrastructure
  2. Emphasized the ongoing development in AI tooling and the growing demand for better applications, with a focus on high-code tools that enhance productivity.

Key Takeaways

  • Developer Experience: Central to Modal’s philosophy is enhancing the developer experience through speed and efficiency.
  • Market Gap: There remains a significant gap in AI tooling that Modal aims to fill with its serverless platform.
  • Future Growth: Modal is positioned to evolve with the AI landscape, focusing on both advanced use cases and broadening accessibility for developers.

Closing Thoughts Erik Bernhardsson's insights reflect a strong vision for the future of AI workloads and the importance of efficient tooling for developers. Modal aims to bridge existing gaps in the market, fostering innovation and productivity in AI and machine learning applications.

For further details, listen to the full episode [here](https://softwareengineeringdaily.com/2025/07/31/modal-and-scaling-ai-inference-with-erik-bernhardsson/).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Modal is a serverless compute platform that's specifically focused on AI workloads. The company's goal is to enable AI teams to quickly spin up GPU-enabled containers and rapidly iterate and autoscale. It was founded by Eric Bernhardsen, who was previously at Spotify for seven years, where he built the music recommendation system and the popular Luigi workflow scheduler. In this episode, Eric joins Sean Falconer to talk about the motivation for founding his company, the market gap in ML and AI tooling, optimizing container cold start, Modal's interface design, and more. This episode is hosted by Sean Falconer.

0:38Check the show notes for more information on Sean's work and where to find him.

0:54Eric, welcome to the show. Thank you. It's great to be here. Yeah, thanks so much for being here. So I was diving a little bit into your background preparing for this. And so it seems like you spent a lot of your time working in data throughout your career, which also kind of matches my own experience. You know, you previously were at Spotify for a number of years, you were the CTO of better.com. Now you're the founder and CEO of Moldal. You know, were there certain things in these prior roles that led to identifying some sort of need for Moldal? Like, what's the story behind essentially going off and deciding to start this company?

1:25Yeah, for sure. The answer is yes. And the long story is I was at Spotify for seven years, built the music recommendation system. But as a part of building that, I also realized there's kind of a general gap in the tooling. I ended up building a vector database called Inoin that we use today, and also a workflow scheduler called Luigi that very few people use today. But generally realized like as a part of building all of that stuff at Spotify and also did a lot of other stuff, that there's very little tooling in data, AI, machine learning. There's more today, but I still never really felt like much later, like in 2020, 2021, when I started thinking about building a company, I realized there's kind of a gap in the market for like a tool I always wanted to have myself.

2:05So that was kind of the genesis of Moto. It's almost like building selfishly for what I always wanted to have, you know, throughout my years at Spotify, building a music recommendation system. And also, just to make sense, I did better. I was a CTO, so a little bit more like general role. I spent a lot of time thinking about platforms and data and stuff like that, too. Yeah, I mean, sometimes people talk about how, you know, the discipline of essentially, or like the tooling and the things that are available for data engineers lag somewhere, you know, five years maybe behind, you know, more traditional application development.

2:35Where would you say that something like ML engineering from an applied sense of being able to actually run these things in a company that may be user facing lags behind, you know, what we think about from traditional application engineering? It's definitely behind. I don't know what the number is in terms of number of years. I think like to me, it all comes down to developer productivity and how much time is like wasted on like tooling stuff versus actually delivering like business value. And the only way I found that's like, I think it's like somewhat correlated with developer productivity is like fast feedback loops.

3:07Like how fast are your feedback loops, right? And it's not like necessarily a perfect metric, but I think it sort of speaks a lot about like developer productivity. Like when you ask developers if they feel productive, a lot of it, I think it comes down to like having fast feedback loops. And when I look at other types of software engineering disciplines, like if you look at like front engineers, like I don't know if you do a lot of front of it. I kind of enjoy it because like you have like editor in one window and you have to browse in the other windows. And like today when you're writing front, it's like so crazy fast.

3:33Like you save the code and it just like automatically just like reloads the other screen and you see it right away. So you have this like, you know, sub second level feedback loops. And I think there's something similar about backend. Maybe it's a little bit different. You run unit tests or whatever. But with data AI machine learning, all of that stuff, I would almost argue we're taking a step backwards in that. As much as I love the cloud and I'm a massive proponent of cloud and it gives us tremendous power to do stuff, for whatever reason, using the cloud and suddenly you have to build Docker containers, you have to trigger things, you have to mess around.

4:07infrastructure ends up like just being this like massive friction in that loop when you're working with AI and machine learning on the engineering side. And that was like my sort of, you know, frustration. A lot of like what model came from was just, I just wanted to be fast. And I've like, maybe I have like lack of patience, but like, I want to write code and then like run it in the cloud in less than a second. So that was like my kind of starting point is like, how do I solve that problem? Yeah. To your point about front end, there is something, I don't know, like there's just immediate gratification.

4:36Like I can go in into string and I see it, you know, and I think that's probably why, you know, you see a lot of people who are starting to get into engineering probably start with front end because you have sort of that immediate feedback loop. And also even from languages, you know, adoption of languages, there's sort of a more immediacy of getting to that like aha moment and those feedback loops, perhaps using something like Python versus using like a lower level language where you might struggle a bit more to just kind and get started. And even older languages, you know, now even things like Java and C Sharp have tried to reduce sort of that barrier to entry of doing that simple like hello world, essentially equivalent application because of how easy it is in some of these other languages that have like severely simplified it and then led to a high adoption curve.

5:21For sure. And kind of as a side topic, like one of my beefs was Rust. I love Rust. It's a great language. I think one of the things that got wrong is like, it's just, you know, like they made a language that makes developers very productive in certain ways. but then you have the compiler getting in the way. It's like, if the compiler was just fast, I would actually really, really enjoy writing Rust. Now it's like, I enjoy it, but it's like, to your point about the older language, I used to write a lot of C++. And having that compilations thing take minutes every time you want to just run, it just takes the joy out of coding.

5:53So, I mean, can you break down a little bit more about, what are you essentially doing with modal? You want these super fast, essentially deployment times for ML workloads, but what's that end up looking like? Yeah, totally. Yeah, and maybe we should take a step back. We talked a little bit about like making developers productive and happy, but like what is modal, right? Like I tend to think, you know, a lot of like running data, AI, machine learning stuff today is about building code and shipping it into the cloud and scaling it up, running it on GPUs, mapping over large data sets, et cetera, right?

6:23So that's like originally what I set out to build. And as a big part of that, like, as I mentioned, I wanted the feedback loops to be very fast. we spent a couple of years focusing on that platform as kind of a core concept. Like we felt there's such a huge opportunity. This was also during the Zerp era. So we had like no pressure to make money. But two years in, you know, where we started seeing a lot of traction was with Gen AI applications. And so a lot of our use cases today that we see that, you know, as driving a lot of the growth of modal is various types of GPU inference use cases in Gen AI.

6:55So large scale, particularly like video, audio, music, like people who deploy models, often proprietary models, that need to scale up to sometimes thousands of GPUs and they don't want to handle the infrastructure or they're running like very large spiky batch jobs. So Modal is both like sort of an infrastructure provider, like in the sense that we have like a big pool of compute in the cloud, like a lot of GPUs, a lot of CPU. And then kind of the separate side of Modal is we have an easy to use Python SDK that lets you iterate very quickly. And in a way where you don't have to install anything, You don't have to configure anything.

7:29You just have to write a little bit of code. And then when you run that code, it runs in modal. And that makes it very easy to take a particular machine learning code and scale it up. And in terms of like this big cluster of GPUs, essentially your customers are sharing. Is that like a shared model? It's a multi-tenant model, right? Which is, you know, different from, I think, traditional applications. I tend to think it's the future of the cloud. There's so many benefits of us having a shared multi-tenant pool. Just to give you one example, if someone needs 100 GPUs, we can often get you that in a few seconds.

8:03Because we can pull all this variable demand, means we can do capacity planning at a very large scale. And so there's been many benefits of having this shared tenancy model. Not to mention also the fact that there's zero installation. Well, you have to do a pip install modal. But then once you do that, you can immediately start running code because we manage the infrastructure in the cloud. for people who are, you know, managing their own GPUs, do people end up typically like under utilizing the GPUs they have available? Totally. Which is another thing that modal solves beautifully is that because of the multi-tenancy, we have a model where everything is pay as you go.

8:40So it's all usage based. So we only charge for the time the GPUs are actually running, right? So there's no capacity planning with modal. You just start running modal and then we charge it per, you know, essentially GPU second or CPU second. Actually, I should say, There's many use cases where people don't use GPUs with Modala. And again, kind of going back to the multi-tenancy model, part of the reason why we can offer this is we can pool a lot of people's very bursty workloads and run an underlying shared compute pool. The other thing we also had to spend a lot of time on is doing things like very fast container cold starts.

9:09Because in order to have a usage-based pricing model, you need to be able to start containers and stop containers very quickly. Interestingly, that is a problem we have to solve in a separate context, which is like we talked about this previously, right? I wanted the ability for users to write code and immediately run it on the cloud. And that's actually the same problem, like fast container cold start. So that is another thing we've had to spend a lot of time on. We built our own file system. We built our own container runtime. We built our own scheduler. And all these things like let's just have, A, very fast feedback loops, B, a fully usage-based pricing, and C, I guess, fully managed infrastructure running in the cloud so people don't have to think about it.

9:48So can you break that down? like, you know, what's happening essentially behind the scenes. So I write, I've done pip install modal. I write some Python code, presumably like I tag it in some fashion. So I know, you know, what I want to run within the cloud. I run it, but what's essentially the magic that's happening between my machine and being able to run this on modal. Yeah. So modal is an SDK, which means like you import modal and basically like the way to think about, I think the like easiest mental model to think about modal is function as a service. It's similar to like AWS Lambda, if you're familiar.

10:17Like the idea is that you can take any function in Python and turn that into a function that runs in the cloud on invocation. So you can have, you know, all these functions in Python. And you can say this one should run on H100. This one should run on whatever, T4. Different container environments, like different, you know, even drivers, right? And you have to specify that in the code using decorators that you apply to these functions. And then you can have these functions call each other. And so you have this like function as a service programming model. And under the hood, the way that works is like we take the code, we're able to build a container image and launch that container image in the cloud in about a second, right?

10:54And more for like very large images. Like if you have an image that's like 100 gigabytes, it might take a few seconds. But we spend a lot of time on like, how do we take the code on the local computer, stick it in a container in the cloud with arbiter dependencies and launch that on a worker in the cloud that might have, you know, GPUs or many workers, maybe you need 100 GPUs that we need to spin up that container 100 times over on different workers in the cloud. Does each function that I'm specifying, let's say I have function one, function two, but I want to run them on essentially different GPUs, do those end up being separate containers that have to get deployed?

11:28Yeah. So every function ends up being a separate container. In fact, every function is auto-scaling. So if you start issuing many requests to the same function, we will just auto-scale it up to as much as is needed in order to serve all those requests. In some ways, it sounds like as a developer of these, you're sort of, it's almost like you're coding a monolith, but then the deployment is able to automatically sort of scale this out as more of like a market service architecture. Yeah, sort of, yeah. And, you know, I think that's one way to look at it. Like every function ends up being its own kind of container.

11:59I mean, you can have like a lot of code running inside a single container, but you can have like very, you know, isolated pieces of code. You know, you can have one container running, you know, Python 3.9 and another container running, Python 3.12, calling each other, just like a normal Python function call. You have this convenience of just, it just feels like local code, just functions calling functions, including handling tracebacks and exceptions and all these things. It just feels like Python, even though it's running in a distributed way in the cloud. Right. How does that function calling happen behind the scenes?

12:31Essentially, are you using some sort of GRPC call or what is essentially allowing you to call functions across these containers? Yeah. So under the hood, we use gRPC just like for internal communication and also the client library talks to the server using gRPC. There's like a few different layers, but we also like within Python, we use Cloud Pickle, just like serialize all the payload and send it between containers. But again, like that's not something that developer necessarily has to think about because it just feels like you're just calling a function in Python. So we handle the serialization, all the exception management, all of that stuff.

13:03But yeah, there's a lot of work, obviously, under the hood. Basically, it's all complex queuing theory and scheduling and doing that fast at a large scale. I mean, I think we're serving something like 100 ,000 requests a second right now across the scale of modal. So managing that, the state management and the scheduling and constantly scaling up and down, obviously, there's a lot of work to build all that infrastructure. Yeah. How's the state management and persistence in these distributed applications work? I would say modal right now focuses on compute. So there is some level of state management, You can create simple distributed.

13:37We have a primitive to create a distributed key value store and also Q. Most of what people use modal for today is things like Gen.AI. Gen.AI is actually kind of like a weird use case in many ways because it's extremely compute intensive, but it's actually kind of low IO. So if you think about stable diffusion, for instance, you send this tiny piece of text to a GPU, and then that GPU does like a trillion operations, and then it sends your little JPEG back that's like 100K or whatever. and that's like the type of applications that we do really well it's like very compute heavy stuff but not necessarily like super io intensive we're not necessarily like trying to compete with like spark in that sense which has like these operations you can wrangle like very large data sets and stuff like that so there is some state management in modal we also have a distributed file system you can set up in modal and you can attach to all the containers basically just like a local postix compliant file system which means like you can basically interact with it just like using normal file system operations.

14:33But anything you read or write is like globally distributed across all the other containers. So that's like another way you can also work with large data sets in modal. But we don't necessarily handle the sharding and partitioning and shuffling. That's not something we've built yet. When I think about something like an iPaaS for traditional applications, a lot of times people run into these challenges eventually where essentially the abstractions start to get to a place where they don't work. They want to get in there and sort of adjust things, tune it to their specification, we'll say we use a certain scale, and it relieves to this graduation problem.

15:05Is that a challenge that people building on Moldo in this world face? I think so. I mean, I think that's something I've spent a lot of time thinking about. It's like, the right abstraction sacrifices a little bit of power, but makes the remaining stuff so much easier to do. So I think any abstraction layer you add, of course, you're going to sacrifice a tiny bit of like you know capabilities but you know by doing that it turns out you can actually the remaining stuff like make it like so much easier to build i think the bad abstractions they basically like sacrifice 80 of the capabilities you know and then like turns out they're like very limiting so like how do you make modal like fairly general purpose so that we can do almost anything in modal i think that's very hard and that being said i mean i think like so far this is something I'm like obsessed with, by the way, thinking about, you know, the right abstractions and thinking about the developer experience and making sure, you know, modal is fairly general purpose.

16:02That modal is not necessarily like super frameworky. I think that's sort of like, I think I don't like about a lot of modal modern framework where like either they're like very config based or they kind of lock in a certain way. Modal is opinionated, but like I want modal to be like a bunch of Lego blocks and then you can like take those Lego blocks and build whatever you want. And I think to a large extent, it lets people do that. We've always thought about programmability. We always thought about making it possible to run any code you want. We're thinking a lot about doing non-Python stuff, for instance, as one example.

16:35So I don't know. I mean, I'm obviously biased, but I tend to think that modal is a platform that doesn't sacrifice that much capabilities. You can do almost anything in modal that you could do just using lower level cloud primitives with like 5 % of the effort. Yeah. This episode is sponsored by MailTrap, an email platform developers love. Go for fast email delivery, high inboxing rates, and live 24-7 expert support. Get 20 % off for all plans with our promo code SEDaily. What are the typical like Gen.ai applications workloads that people are running on this? Is this primarily people running their own model and they want to be able to run inference across these GPUs or are there other types of workloads that are also there?

17:23Yeah, so one user that I love, one of our customers is a company called Suno. They do AI-generated music. So they run a lot of their inference on modal, very large scale. So basically, what they run on modal is GPU-based. They have their own proprietary model and it generates music. Super cool application. I think a lot of our customers fall into that pocket. It's like there is a proprietary model. It does some very cool magic stuff on a GPU, typically in audio, video, image stuff, music. It's been kind of the domain we've seen most traction. There's other use cases too. We've recently seen a lot of traction coming from computational biotech, which I think is super exciting.

18:03So like protein folding, multiple sequence alignment, like there's all kinds of medical imaging, processing, like very large data sets of, you know, using applying computer vision to like assays. I don't know too much about the field, but like I find it incredibly exciting. And kind of in that vein, like we've seen people use model for like geospatial analysis, physics simulations, like turbulence stuff. And then there's a lot of people who just like the developer productivity, and they don't necessarily run big things, but they have a little web scraper running in modal or a little web server.

18:36There's a lot of that stuff. So modal is, the goal was always to build a fairly general purpose platform. We found initial core product market fit and the Gen AI. That's been the main use case for us. But there's just so many other things that people always surprise me. I was just talking to someone the other day who was running a chess engine on modal. I don't know why, but it was really cool. I mean, some of those examples you gave are pretty high compute examples, which I think make a ton of sense. Are there certain workloads that don't make sense? I guess you mentioned things that maybe require a high I.O.

19:06are probably maybe not the right fit. Yeah, high I.O. I think we're like pushing into that. I think it's going to be a big focus for 2025 to also handle. Similarly, I would say like very low latency things is something we're like very excited about. Right now, there's like an overhead in the system of like 100, 200 milliseconds of every function call. a little bit less if you run it in the same region. But that's enough for a lot of use cases, especially in Gen AI, like stable diffusion. No one really cares if it takes 200 milliseconds because the inference in itself takes a couple of seconds.

19:34So that's pretty negligible overhead. But it's not enough for something like real-time streaming of audio or real-time video. So that's another thing we definitely want to push into over the next year is how do we get the overhead of the system down to 10 milliseconds, 5 milliseconds, whatever. In terms of getting to a place where you're able to run these for things like real-time, you know, video, audio streaming? Like, what do you see as the main sort of technical hurdles that you have to solve in order to reach that kind of performance? Today, it's mostly about geolocation or basically like we run a distance.

20:05Exactly. Like the speed of light is pretty high, you know, and as you may know, like, you know, it doesn't take that many 100 milliseconds to like send something to Australia and back to the US, like, you know, 200 milliseconds or whatever. But, you know, it is a challenge when you're, you know, doing real-time stuff. So a big challenge for us is like, while we have a distributed data plane, like all our workers run in many different regions, many different cloud providers, our control plane right now is not distributed. So one of the things we want to do in 2025 is decentralizing the control plane so that we can use, you know, smarter ways, basically like kind of route to, you know, multiple edge.

20:39I feel like the word edge is overused, but you're running a model in many different regions across the world, like also the control plane so that we can route things and execute it faster. That's a big re-architecture. So in this multi-tenant architecture that you have today, you have your control plane, you have your data planes, your data planes are distributed across these regions. Is primarily the sort of reuse of resources happening in the control plane? Yeah, the control plane makes all the decisions, right? And it sort of has the global state of the world, which is another thing to some extent.

21:10When you have a very large worker fleet and many different scheduler running in different regions, I think you also kind of need to change the truth of the system to be owned by the workers themselves, because they always know the latest. Because, again, speed of light is not always as fast as we wish. So there's a lot of that state management. Where do you make the decisions? Who has the authoritative view of which worker is running which containers? And how do you propagate that information? I mean, these are hard technical challenges, but luckily we like those at Modal. Was the plan from the very beginning always to build this out as this like multi-tenant structure?

21:46It is challenging sometimes because I think people are just not entirely comfortable with that model yet. But like, I don't know, I've been coding for 30 years. And like, I remember when the cloud came in 2007, it was something like that. Like, you know, AWS launched or EC2, I think something like that, 2006 maybe. And my first thought was like, that's insane. Why would I put my code on someone else's computer? And then like, you know, a few years later I was doing it. I was like, this is kind of nice. I like this. Yeah. So like, you know, it took a few years. I mean, it's still taking a few years.

22:16Like many people still run on Prime, right? But like, I think there's been like, obviously, you know, people are seeing that the cloud makes sense for a lot of stuff, even though it's like arguably like a shared resource. If you look at, you know, a company like Snowflake, they were like, hey, we're going to build a cloud native database. And at that time, which maybe doesn't seem that crazy now, because basically everybody's doing that. But at that time, it was like, what do you mean you're going to take this thing that we run in, I don't know, on-prem Oracle today and move it to the cloud? Like, we're not going to do that.

Read the full transcript

22:47Yeah, it's funny. You took words out of my mouth because I was just going to talk about Snowflake next. It's funny. I actually interviewed with Snowflake 2012, and they told me the idea. I'm like, I don't think this is going to work. And then I turned out the job offer. It was obviously a very terrible decision. But I think what Snowflake showed is like, you know, beyond just the cloud vendors, also Snowflake showed that you can be infrastructure as a service and, you know, host people's data. And eventually, like people will be comfortable with that. And so, I don't know, I think security compliance is like shifting.

23:17I think people, you know, a lot of customers today, like they don't necessarily worry about the fact that something is multi-tenant. They worry about like best security practices. And we take that extremely seriously. We think a lot about, you know, how do we encrypt all the storage? How do we encrypt all the data in transit? How do we run the containers in a way where like it's impossible to break out of containers? Those things are very, very important. Like exactly where things are running in terms of networking or in terms of VPCs. I don't necessarily think those are like the prime concerns.

23:46And I think over time, there'll be even less of a concern. So multi-tenant is definitely like a change in mindset. It's a change in like how security and how compliance operates. But I think, you know, the trend is our friend here. Yeah. I also think from like a security perspective, when you look at things like breaches or other security vulnerabilities, I don't see a lot of reports that are a result of some sort of compromised multi-tenant, you know, cloud offering. It has a lot more to do with just like a lot of times simply human error of like, hey, we accidentally committed, you know, our API credentials to GitHub and that's, you know, available online.

24:22Yeah, that's always what happens. Have an unencrypted log file that has the social security numbers of all our customers in it or, you know, things like that. Totally. And even network segmentation, I think, is kind of, to me, like an outdated model. There was like a hack a few years ago where like Target got hacked because it turns out their HVAC system was running on the same network. And like there was like a vulnerability and they're like, whatever, the HVAC system and someone was able to get into it. So I think network segmentation is like to me is like never like a strong security model.

24:48And I'm kind of glad like people are not, you know, there's like the beyond car corp and zero trust. That to me is like very clearly the future of cloud security. You know, in terms of things like capacity planning for inference, like even outside of modal, is capacity planning for inference and Gen.AI applications like a fundamentally difficult task for businesses today because it's hard to know in any given time, like how many tokens you're going to be generating and, you know, what the workloads actually look like. Yeah, I think so. Because there's like so many sources of noise, right? I mean, first of all, you have sort of a daily variation, like kind of a sine curve over the day.

25:23But then often you have like kind of a noise. I mean, it's like a Poisson distribution at any point in time. I have always like a little bit like stochastic noise. But then there's so many other things, like people running batch jobs, like suddenly you launch something and it goes viral and hacker news. Like there's all these different sources of noise. And, you know, the distribution is extremely fat tailed. And so it's fundamentally hard to plan, you know, do capacity management for inference. I think for training, it's a little bit easier. Like when you're training a model, you can just like buy, you know, 1 ,024 GPUs and just like kind of make sure that GPUs stay warm.

25:56But for inference, it's much harder because you can't fundamentally plan. And I think that the thing that also exacerbates that is sort of the only way in the past, at least up to very recently, to get, you know, high-end GPUs was to make big reservations, long-term reservations like three year or whatever right and and that's fundamentally kind of a hard matching problem like you you have this like unpredictable demand where you have to like go out and buy like fixed capacity which means that a lot of people are running things very poorly utilized however i actually tend to think resource pooling is kind of a free lunch in many ways if you take a lot of people's noisy workloads and you aggregate them you can run the aggregate at much much higher levels of utilization.

26:38So that's one way we can save a lot of costs for people. Even though, frankly, our prices is sometimes higher, we can still save money because they run things that effectively 100 % utilization from the point of view of paying modal. How do you see some of the landscape around AI infrastructure evolving over the next couple of years? Maybe we're still very much in early days. How do you think that's going to continue to change and evolve? Yeah, we're in the super early days. This is so incredibly hard to predict. I think it's also, you know, like there's so many different layers, there's so many different boxes.

27:10I don't know. Like I'm very bullish on high code tools. I think at the end of the day, you know, looking back at my 30 years of, you know, coding and, you know, going even further back, I think the like story of software engineering has always been, you know, better tools drive more demand. And then, you know, making engineers more productive is always fundamentally like what, where the value is created. So I'm very bullish on building better tools. I'm very bullish on the infrastructure layer, making it easier for people to build these applications because clearly there's a lot of demand. So that's what modal focuses on.

27:44In terms of the types of things that people are doing today in the industry to build AI applications, using things like vector databases for building RAG applications. And there's been all kinds of takes now on RAG. I see a new three-letter acronym with, on a weekly basis. But what are your thoughts on vector database as a dedicated storage for AI? Just given that you've spent a lot of time thinking and working in this space, do we need that? Is that sort of the right form factor for building some of these applications? I don't know. I mean, I think it's needed on some level, right? Like we definitely, you know, and I started using vectors and built my own vector database at Spotify back in 2012, I think, something like that, and open sourced it.

28:24And for a while, like, you know, it was called Annoy, still called Annoy. A lot of people actually used it. But funny, actually, when Twitter open sourced their recommendation algorithm a year or two ago, I was looking through the source code and apparently they were using annoy for some of the stuff. Anyway, so talking about vector databases, like, I think there's definitely like a need for a vector databases. I think that being said, like, when a space is new, I feel like no one really knows what's the actual, like the ultimate abstraction and the right, like interface boundaries. So with vector based databases, I think a valid criticism is like, maybe that should just be a part of Postgres.

28:58Maybe Postgres should just do vector. And I think that's valid. But I also wonder if like, in the long run, I don't even know if we know where the boundaries are going to be. I think right now we're kind of, it's like easy to look at Postgres and say, yeah, we should just put vectors in it. But like in the long run, things kind of end up redrawing themselves. I could see a world where, for instance, you know, in one direction, you can say, maybe people shouldn't even think about vectors. Maybe people should think about a database. You can insert text, you can insert images, and then you can search for which images and texts are close to each other, whatever.

29:27Vectors is kind of a low-level primitive. So that's like one way you can think about it. It's like maybe the interface shouldn't even be vectors. Because right now it's kind of tedious to it. Like you first have to embed it, you know. So like maybe the embedding should like sit in the database. I don't know. That would be like one direction. Another thing I've thought about a lot is LLMs are kind of in a way like vector databases. They store like very large matrices. And the way they store these matrices, this would be like the other direction. And they do the state lookup through like very expensive matrix multiplications.

29:55Maybe there should be like a differentiable vector database inside every LLM. My point is just like, I don't know, we're still kind of early with vectors. And I don't know, like in the long run, abstractions, interface boundaries, like all of these things may change. And categories never look the same, you know, when you look at it. I think it's too early. It's too early to say. Yeah, I mean, it is a little bit strange that currently, and I think this is just a sign of the times and things being early, that you have to think about actually like generating a vector or even think about what, you know, what a vector is.

30:23in order to search a database. It's a little bit like having to really understand, and there is some value to this, like if you're like a DBA or something, but really understanding like the underlying tree structure that's used in indices to think about like, you know, optimizing my lookup and so forth. Like for most people doing simple application development, they don't necessarily need to be like, you know, digging into that level of detail. Yeah, I don't think that's the interface people necessarily want in the long run. And for that reason, And I don't know if the shoe warning vectors into Postgres is going to be the right boundary.

30:56But we'll see. We'll see. What are your thoughts on some of the role of things like, you know, lakes and warehouses when it comes to building, you know, Gen AI applications? I think in like traditional machine learning, you know, a lot of times we would, you know, we're aggregating specific data down into a particular location, going through a process like feature engineering to build like a bespoke model that we need to deploy. And maybe we don't update it that often. But in sort of Gen AI applications, we have like a very general model, and then we're sort of massaging it for application-specific behavior by, you know, adjusting the prompt during prompt assembly, which is a lot less about sort of the old world of batch updating these models, but really sort of real-time updating the prompt in order to generate some sort of behavior that's going to be relevant for the user.

31:43So I'm curious about what are your thoughts on that? Is this sort of a shift in the way that we need to think about the role of, you know, lakes and warehouses? I'm not sure. Like, maybe this is like my boomer perspective. I feel like, in a way, maybe we're like throwing out the baby with the bathwater. When I look, you know, five, 10 years ago, if you look at like a search application, right? Like they would have like a multi-stage use three-volt process and then like a re-ranker to use an ensemble. They would have all these like feature stores, you know, generate a lot of features. And in the end, they would run some sort of XGBoost to like aggregate all those features and then rank based on that.

32:16My feeling is like, and recommendation systems kind of similar, like searching, ranking, like all these things had this, you know, kind of complex setup. And it was kind of hard to build. And then, you know, my feeling is like, there's so much demand for these applications, but they're kind of hard to build. So then when like LMs came, people are like, actually, let's just turn this into like a bunch of prompts instead. And doing that, I kind of feel like we, it's almost like low code, like retrieval. My feeling is like doing that, we kind of threw out the baby with the bathwater. So like we're like in the long run, I wonder if a lot of those prompt engineering things will just become features into like a multi-step retrieval process that will also combine other features.

32:58So like the pendulum will swing a little bit back towards those more like traditional models. LLM ends up being like one feature, a very powerful feature, but just like, you know, you might have different prompts, like generating different features. But in the end, you sort of go back to sort of more traditional view of like a multi-step retrieval. I'm talking specifically about those types of applications. There's many other types. But like in general, like I wonder if LLMs, you know, may just become like one feature out of many other features. And that's like, will be a step backwards towards the traditional sort of feature stores and all that stuff that has, you know, been a very powerful paradigm for many years.

33:36In some sense, I mean, I think you could think about like agents to like, you know, this sort of multi-step process where you can weave in, you know, traditional ML models or even other types of workflows. And a core component of that, of course, is the LLM or the foundation model, but it's not necessarily, it doesn't have to be responsible essentially for all behaviors in there. You can use a combination of things to get essentially the behavior, the output that you want. Yeah, for sure. I'm very bullish on the high code approach. I think LLM is a little bit the low code. But like over time, clearly there's a lot of demand for people wanting to build AI applications and machine learning applications.

34:11So like I think there's going to be 10x more machine learning researchers in the long run. They're going to use LLM. So they're going to be very happy with it. But they're also going to use the like underlying like core models as well. What are your thoughts on the energy consumption required to do model training? I think the latest open AI model consumes more electricity than the city of Pittsburgh. I don't know. I think it's kind of overstated. Part of why I think it's somewhat overstated is actually the cost of a GPU over its lifetime versus the energy consumption of its lifetime. It's actually still, the main cost of training a model is actually still on the GPU side.

34:46Building the GPUs is like a lot more expensive than the energy required to run the GPUs. Even if you run a GPU at full capacity for its entire lifetime, the cost of the GPU, which to some extent, I think like speaks to just the cost of building the GPUs. So like, I don't know if running, it's actually the other way around for CPUs. Like CPUs, if you look at the cost of operating CPUs, the energy cost is higher than the cost of buying the CPUs if you've run at full utilization for a few years. So I don't know. I tend to think we focus a lot of the energy consumption. I don't know. Humanity always needs more energy and they always find uses for more energy.

35:22I'm an optimist. Otherwise, I wouldn't have started a company. So I'm very optimistic when it comes to, you know, things like climate change and, you know, energy consumption. I think we're always going to find more energy sources and the cost of, you know, and the GPUs, like the energy consumption is going to come down and the GPU costs will come down. And all these things will be easier and, you know, do, you know, more better, you know, in the end, they're going to create a lot of value for human beings. humanity. Speaking of sort of having this like optimistic view of the world and that being sort of a core part of, you know, surviving the hardship of, you know, building a company.

35:53When you were thinking about modal originally, and you had this sort of vision of being able to do, you know, these deployments at, you know, less than a hundred milliseconds for running ML workflows, but were there parts of that that you weren't sure you'd be able to actually solve in order to realize that vision? Like, were there certain things, you know, technical challenges that you were really scared about whether you'd actually be able to be successful with? For sure. And I think a lot about like one model I have of like a startup is like, you have to pick a problem that's like hard enough that, you know, you create a lot of compelling value, but it's not hard enough to, you can't solve it.

36:28Right. I look at a lot of AI startups and like, I feel like, you know, there's a lot of companies in all three buckets. Like there's companies that I think are doing too, solving too easy problems. And for that reason, they're just kind of rappers and it's they don't have like a lot of pricing power and it's like you can sort of you know i sort of doubt that they're gonna have the competitive advantage long term then on the flip side there's companies that i think have like way too ambitious goals and they're like we're gonna do agi or like whatever like we're gonna solve you know we're gonna do this agent thing that can act autonomous for everyone and like that sounds very hard i think the trick of a startup is like kind of picking a problem that like could conceivably be solved in about three years, you know, and like, I think three years, four years, maybe five years, it's like kind of a good timeframe.

37:09And so yeah, you know, I've spent most of my career in infrastructure. And for me, like looking at this problem, I knew all the components would be possible to solve. Like I knew, you know, looking at containers, I was like, yeah, it's possible to start containers quickly. Because, you know, Docker works this way, and Kubernetes works this way, and they're doing a lot of unnecessary stuff. So like, that being said, I mean, a lot of my early VC conversations, like people are like, why aren't you just using some existing system? But I was like, kind of adamant. I'm like, no, we're going to build our own thing.

37:40And I'm very happy I did that. And it took about three years. But I think doing that means now we have like a pretty strong competitive advantage. And we have a very unique set of infrastructure primitives that lets us build things and deliver much better developer experience than anyone else. Are there particular like performance optimizations that you did that you're particularly proud of? So much, but I mean, I think, you know, at a high level, like a big part of it was building your own file system for serving container data. As it turns out, like Docker, I mean, as much as like, I think, you know, I respect like Docker for like, you know, introducing a very new paradigm.

38:11It's quite inefficient in how it stores images and pulls and pushes images. So what we realized, like looking at like what happens when you start a container is that most of the data is never read, you know, that is in a container image. and the data that's read is highly redundant between images. So what if we can switch to using a content address system and then we built a Fuse-based file system that caches all the data under the hood. And then we then spent like two years like figuring out how to optimize the page cache and how do we like all these things, right? But like that is like a big part of why we can start containers very quickly is just optimizing the hell out of just the file system side of it.

38:49How do we send the data very quickly? Got it. So as we start to wrap up here, is there anything else you'd like to share? What else? I mean, like we're working on distributed training. I think that's going to be really exciting. We're hoping to launch that pretty soon. And the focus is not like super crazy large, like running thousands of GPUs. But for companies who are training models, and maybe it takes too long to train on a single GPU, we make it pretty easy to scale up to eight GPUs right now. But we want to go beyond a single box. So pretty soon we'll make it really easy to scale out to 16, 32, 64, maybe, maybe even 128 GPUs and get this like super fast feedback loops, just like, you know, kind of was always like the core value of like modal.

39:28Now you can also hopefully get that up to, you know, 100 GPUs. So I think training is going to be super exciting. What else? I mean, I think throughout the year, like, you know, we're spending a lot of time on security compliance, we're kind of moving up market and focusing a lot on like enterprise customers and serving their needs. you know all the range you know the whole range from like sso to sock to building custom telemetry integrations and stuff like that so that's another area i'm also like super excited about awesome well eric thanks so much for being here yeah it's great thank you so much for hosting me yeah cheers thanks

40:13Thank you.

From the publisher

Modal is a serverless compute platform that’s specifically focused on AI workloads. The company’s goal is to enable AI teams to quickly spin up GPU-enabled containers, and rapidly iterate and autoscale. It was founded by Erik Bernhardsson who was previously at Spotify for 7 years where he built the music recommendation system and the popular

The post Modal and Scaling AI Inference with Erik Bernhardsson appeared first on Software Engineering Daily.

More from Software Engineering Daily

All 195 episodes
Modal and Scaling AI Inference with Erik BernhardssonSoftware Engineering Daily · 40 min
Listen in VO