In short
Eye On A.I. Podcast Episode Notes
Episode Title
#229 Mitesh Agrawal: Why Lambda Labs’ AI Cloud Is a Game-Changer for Developers
Episode Summary In this episode, Mitesh Agrawal, Head of Cloud/COO at Lambda Labs, discusses the company's evolution from a style transfer app to a leader in AI compute infrastructure. Mitesh highlights Lambda's innovative approach to providing scalable GPU solutions and accessible cloud platforms designed for developers and enterprises, tackling the global GPU shortage and optimizing AI workloads.
---
Key Topics Discussed
- Lambda Labs’ Origins
- Began as a style transfer app called DreamScope, converting photos into artistic styles.
- Shifted focus to AI compute infrastructure upon recognizing the profitability of building hardware and software for deep learning.
- Lambda Stack: A collection of libraries (CUDA, PyTorch, TensorFlow) designed to simplify machine learning implementation.
- AI Cloud Infrastructure
- Lambda Cloud: Focused on deep learning applications, allowing remote access to GPU clusters similar to AWS or GCP but optimized for AI workloads.
- Emphasis on transparent pricing compared to hyperscale providers, aiming for affordability and accessibility.
- Pricing Philosophy
- Transparent pricing strategy helps reduce market confusion and builds trust among developers.
- Emphasis on growing access to compute resources supports the AI community.
- Managing GPU Supply and Demand
- Discusses innovative supply chain strategies to cope with global GPU shortages.
- Orders GPUs months in advance based on anticipated demand, maintaining high utilization rates.
- Evolution of Workloads
- Acknowledges the shift from training workloads to inference as foundational models become more complex and integrated into applications.
- Lambda's infrastructure is designed for both training and inference, leveraging cutting-edge NVIDIA GPUs.
- Future of AI Compute
- Explores the potential rise of localized data centers to reduce latency for applications.
- Predicts an increasing need for compute power as AI applications evolve, particularly in areas like video generation and reasoning models.
- Challenges and Constraints
- Regulatory issues impacting access to GPUs in countries like China, limiting AI development.
- Observes that while sanctions may pose challenges, innovation persists in various forms within China.
- Lambda's Growth Trajectory
- Currently operates multiple data centers with aspirations to scale significantly in the near future.
- Targeting sustained growth rates over 100% as demand for compute resources rises.
- Long-term Vision
- Aims to democratize access to AI compute while fostering innovation across industries.
- Encourages developers to explore Lambda Cloud for their deep learning needs.
---
Key Takeaways
- Lambda Labs has successfully transitioned from a consumer app to a powerhouse in AI infrastructure.
- The company prioritizes affordability, transparency, and accessibility in its pricing and services.
- A significant shift in AI workloads from training to inference necessitates a robust and adaptable infrastructure.
- Localized data centers and innovative supply chain strategies are crucial for meeting growing demand.
- Global regulatory challenges, particularly around GPU access, can significantly impact AI development in various regions.
---
Contact Information
- Lambda Labs Website: [lambdalabs.com](https://lambdalabs.com)
- Custom URL for GPUs: [gpus.com](https://gpus.com)
---
Episode Structure
- Introduction (00:00)
- Origins (01:37)
- Pivoting to Deep Learning Infrastructure (04:10)
- Building Lambda Cloud (06:23)
- Pricing Strategies (09:16)
- GPU Supply Management (12:52)
- AI Workloads Evolution (16:34)
- GPU Preferences (20:02)
- Future of AI Compute (24:21)
- Global Challenges (28:30)
- China’s AI Landscape (32:13)
- Scaling Growth (39:50)
- Advancing AI Models (45:22)
- Optimism for AI's Future (50:24)
- Accessing Lambda Cloud (53:48)
---
Conclusion This episode provides insightful perspectives on the future of AI infrastructure, the growing demand for computing resources, and Lambda Labs’ role in shaping the landscape for developers and researchers in the AI field.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00I think it's important that people stay very optimistic about this technology. Of course, we have to have, you know, controls around a little bit controls around how people it doesn't go into the wrong side of things. I want to encourage people to to be in the know about this. Underlying it all, like, of course, we want to build a big, successful, long term business. We really care about that. It's kind of our baby to build that. But underlying of it all, we really want to be one of the ones that provide a lot of compute and ease of compute to people that are building the next step of applications for this world.
0:32And we really focus on the deep learning aspect of it. We spend a lot of time building for it. And for a lot of listeners, again, I would encourage to try Lambda Cloud if you're building any deep learning application. What does the future hold for business? Ask nine experts and get 10 answers. Bull market, bear market, rates rising or falling, inflation going up or down. Can somebody please invent a crystal ball? Until then, over 40 ,000 enterprises have future-proofed their business with NetSuite by Oracle, the number one cloud ERP, bringing accounting, financial management, inventory, HR into one fluid platform.
1:18With one unified business management suite, there's one source of truth, giving you the visibility and control you need to make quick decisions. With real-time insights and forecasting, you're peering into the future with actionable data. If I were a larger organization, this is the product I'd use. Whether your company is earning millions or even hundreds of millions, NetSuite helps you respond to immediate challenges and seize your biggest opportunities. Speaking of opportunities, download the CFO's Guide to AI and Machine Learning at netsuite.com slash ionai. That's netsuite, N-E-T-S-U-I-T-E dot com slash ionai, E-Y-E-O-N-A-I, all run together to get the CFO's Guide to AI and Machine Learning.
2:21The guide is free to you at netsuite.com slash ionai. netsuite.com slash ionai. Why don't you start by introducing yourself and how you got to Lambda into this business? I'm very happy to, and thanks for having me here today, Craig. Like, look, I think the best place to start is always with the beginning of a company and for Lambda. You know, I think the best thing was we all were engineers and in machine learning fields. And, you know, when Lambda was initially started off as a face recognition API company. So basically, we had an iPhone app that used to take your photos and convert it into a Van Gogh-style painting or a Picasso-style painting.
3:17It was called transfer-style learning back then. And we used to have our processing done on the cloud. And it used to cost a lot of money. We were never profitable doing that business. What we realized at that point of time was that actually we could purchase the hardware that we were running it on, the GPUs, and we could deploy it and we could install what we now call Lambda Stack. But at that point of time was just what is called a Debian repository, just a collection of all the libraries like CUDA, CUDNN, PyTorch, TensorFlow, languages on it and make it super easy to run. and it immediately like took out all of our kind of spend on the cloud.
3:59So we saw that there was a gap in the market where, you know, you could buy a Dell, HP, whatever hardware, but really like to make it work for machine learning, you have to install this software stack and really support it. And that's kind of how Lambda got started. And for me, I think, you know, kind of an interesting journey, you know, being in Bay Area and Stephen, who is the founder of the company and machine learning researcher himself, um you know we we were just introduced and we're talking and and he was like hey there's this opportunity we we're pivoting into infrastructure um i i was from a background of chemical engineering and and have been working in quantitative finance but was trying like starting to see early inklings of kind of at that point of time you know it was just like advanced statistics basically like just prediction prediction base and things like those was trying to understand these models and i was hey, this is a good way for me to foray into this world.
4:55I really understand the computational layer. And I joined in as a very, very early employee, number fifth at Lambda. And pretty much as the saying goes, did everything that the company needed it to be. Was it an engineer? Was it a support person? Was it an HR person? Was it an office manager? Was it a salesperson? Everyone did everything. and that's how Lambda kind of grew really from both my perspective and this external perspective is like we really branded ourselves into this infrastructure world but only very much focused on what we call deep learning applications right really like making sure that at that point of time the state-of-the-art models were convolutional neural nets the res nets all of those models really making it work for that and then over time you know look company's mission was really focused on bringing compute as quickly as possible to the machine learning community.
5:53And we realized we had an opportunity in 2020, 2021 with cloud starting to grow. But what we noticed was the hyperscalers were still, and everyone was in early innings, actually. And no one really foresaw ChatGPT, maybe other than OpenAI. And even they would probably say that they didn't foresee the enormous success of it. But we realized that the compute will be a very critical category and compute design for AI will be a very critical category. And we kind of stepped into that where it's like instead of just selling infrastructure to be deployed in an end-person environment, whether that's through a desktop or laptop at your office or through a cluster in a data center, we said, why not build?
6:37We are experts at building clusters. We're experts at the software stack about it. let's build it in our data centers and give people remote access, aka like an AWS or GCP-like service, but for deep learning. And really optimize on kind of the networking stack, the software stack, the speed at which you could do the model. And that's kind of how Lambda got the first, like, really, like, foot down into the space. And then, you know, we saw the growth curve. We got attracted to the growth curve and really then doubled down on building this kind of AI cloud and cloud for developers and researchers and enterprises, again, keeping to a theme of only focusing on deep learning applications, but really building around that.
7:19So that's been the last journey of eight to nine years for Lambda. I'm just curious, this style transfer app that you guys have, what was its name? Yeah, it was called DreamScope. So, you know, it was like, look, I mean, of course, we look at it with very like loving eyes and it was like the mid-journey before the mid-journey right you know i think mid-journey is a brilliant brilliant application that has been developed right now and and and we we we did this back in 2016 2017 uh and really it got a niche following in the arts community because you can imagine right it just it wasn't like a gender develop right you couldn't type something and it would generate it but what it would do is take an existing piece of whether it's an image or a series of images and then convert it into any piece of art that you want.
8:10And we really focused on kind of the artistic, like the noir style, Picasso, all the artistic styles that people do. And it got over a million downloads within a few months of launch. We had a package called, a paid package called$9.99 and we had a lot of paid users at one point touching well over 100 ,000 ARR, which at that point in time, I mean, now it's very small compared to what the AI companies do, but at that point in time was a new and radical kind of way of approaching the consumer marketplace. And, you know, it's a baby, but like, you know, we had to, first of all, a lot better models came out, people who developed their own models.
8:51And then second, because we focus so much on the AI learning infrastructure, we realized that focus is key to growing this market. And so we, as a company and as an expertise area, we really pivoted into building out infrastructure for deep learning engineers because we ourselves were users of those. Yeah. The reason I ask is I was fascinated by style transfer back in 2016, 2017. And I did use a few apps. I used Picasso with a K, P-I-K-A-S-O. And there was another one I can't remember. Yeah, even I'm blanking on the name. It was a very popular Russian app. Damn, I'm blanking on it. But yeah, I know too.
9:34I've used, yeah, exactly. Stryle Transfer were the first really consumer applications of deep learning. I'm actually surprised that it didn't take off or become more popular. And then there's another company that's doing something kind of similar to Lambda, PaperSpace. I got to know those guys years ago, and I can't remember what they were doing, but it's the same kind of story. they were doing something in that you guys were working on style transfer, and they saw this opportunity for GPUs, and they moved into that space. Why are you cheaper than a hyperscaler? Yeah, I think that's a fair question that you can imagine we get asked across the board, whether it's investors or the media or just users.
10:38Look, I think there is a few things that go into it. First and foremost, let me be honest. There's not like a huge magic wand here that is somehow making like, you know, Lambda has like a way of producing the cost structure or because we get GPUs that are cheap or nothing. Nothing like that. So there is a few things that go into it. One is just Lambda's philosophy. I'll be very frank with you and with everyone listening in and we have been with our investors too. We really want to grow the access to compute for all of both, not just like incoming developers and students and things like those, but across even enterprises.
11:14Before the advent of this kind of AI cloud companies, including Lambda, the GPU marketplace was a little hazy in terms of pricing. You had your website, even today, if you go to AWS and you look at AWS prices on the website, for Hopper, it's like$5 or$6 or$8. Almost no one is paying for that. And that's been the reality for a long time. But you had these different pricing strategies that the hyperscalers did. Lambda went in with a simple task of transparent pricing, straight pricing to the market that we believed allowed us to both grow and capture enough margins to continue to run the business and be successful.
11:57So that was one, it's just philosophy around pricing. Two is this, and this is one of the technical reasons. Again, not a magic wand, but if you think about hyperscurses, they have to support many, many hundreds or thousands of use cases, right, as they go. So they have microservices and supporting infrastructure for those services that they have to build. So when they deploy a GPU cluster, they don't just think about the GPUs, the CPUs around it, the memory, the storage, the networking. They have to think about a hundred more myriad different things, how to get interconnected to all the data center regions, how to have all the services.
12:35So what that results into is a slightly an additional cost structure that they want to recoup, but also a little bit time where if Lambda were to get the GPUs today and let's say Hyperscale were to get the GPUs today, we would get that deployed way faster. Not again, not because we have more people or anything like that. It's just because we don't have to think and optimize for that. We have to optimize for deep learning. So the focus there helps. And as such, if you have less time to market, so you'll spend less on operational costs, you have to don't think about the microservices and things like those.
13:08So it really allows you to price way more competitively. And the last thing I will say is this. Look, almost all the hyperscalers are public companies. They have certain expectations to meet on what their gross margins are, what their operating margin structures are, and things like those. For Lambda, I think being a private company here and really optimizing for growth allows us to take lesser margins than hyperspace. Again, as long as the business is sustainable, we go through that. So there's just multiple things that contribute to it, including, as I said, the first thing that really contributed to it was us, ourselves, as a philosophy, we wanted to provide as much compute as possible.
13:50And we were advised by fellow deep learning researchers and engineers who help us at the company and who are advisors to us saying that, look, the best way you could grow is by making sure that everyone understands and learns about deep learning and that there is an accessible way to use that compute. So it kind of makes sense from that perspective. Yeah. And then growth, as you grow, do you, as demand grows, you add GPUs or do you add GPUs and then sell the demand, sell the capacity? How does that work? And do you have any trouble getting GPUs since there's famously a shortage of GPUs globally?
14:33Yeah, two very kind of circular situations or catch-22 situations, as you mentioned, right? I think, I'll be frank, I think the last two years have been, we've been blessed in the sense that on the demand curve, the demand curve has grown so fast that no matter how much supply you're providing the market, it was above the supply curve, right? The demand curve was above supply curve. So to your point, whether, you know, how many of our GPUs we kind of ordered and deployed is a function of less of what the demand curve of the market is. So to your point, whether we saw demand and then deployed, no, it was more a function of how much capital do we have to purchase the GPUs, how much data center do we have space and power do we have to deploy these GPUs, and how much can we, to your point, how much can we get allocation of the GPUs.
15:27So it was more a function of those, you can think of like commercial factors rather than really the demand curve. The demand curve has always been there. So from Lambda's perspective, we pretty much always, if someone is doing like a reserve, what is called a reserve contract. So like committing for a year, two year, three year for compute, we're kind of always sold ahead of the curve. Like so people have come to us and say, hey, I want a thousand GPUs. I want to 5 ,000 GPUs. and you say, okay, great, give us three months and we'll get it deployed for you or things like those. So that's on the reserve side.
15:59Now where Lambda has kind of differentiated ourselves is a little bit more on the on-demand compute where to your point, we deploy compute without having an associated reserve or committed contract to it because we want to. We want to create a marketplace where their people can come to Lambda, see a GPU available and spin it up and use it the moment they want it for how much of their time they want it and then spin it down. We want to create that. So we encourage that marketplace. So we basically assign a certain percentage of the GPUs that we get for that market. Now, we have been, again, as I said, both lucky and as a business, really, you know, I guess having that kind of visibility into the pipeline that GPUs as a demand curve has always stayed of supply curve.
16:45So it has allowed us to, even when we have booked it for on demand, And those have pretty much immediately sold out. So we reach high levels of utilization within a few days to weeks of launching those GPUs. So far, it's been always ahead of the curve. In terms of securing GPUs, I think for Lambda, I think NVIDIA has created a partner-friendly ecosystem. So if you think about it, how much GPUs Microsoft would like or want, probably nvidia could park all its gpus at microsoft you know maybe i think uh but no nvidia has been very i think uh much more i guess in a way democratic and i think in the sense of like how they have uh provided gpus and and part of it is just that they follow the usual traditional supply chain principle of first in first out so in the sense that you mentioned how do we get allocations you know again not a huge no matter kind of what you see in the media it's not like a like a, oh, like, you know, somehow we go to NVIDIA and say like, hey, we are your favorite people and can you please give us GPUs?
17:52It's more around the fact that like, look, we have this demand curve. We are placing a purchase order six months, nine months ahead of time. And we are committing to this purchase. And as such, you get into the queue of where the GPUs are going and you get those allocations. And so you have to do the supply chain planning very much so to get the allocation. You know, so I think over time, you know, as Hopper, when they first came out, it was hard to get for everyone in the amount that you wanted to get, because, again, demand was about supply. But, you know, everyone got whatever they placed their queue into the, how much ever they placed their order in and queue in.
18:30And accordingly, everyone got the ratios as NVIDIA also ramped up its production system. Now it's like very, very much accessible to get a Hopper GPU in the Edge generation, Edge 100 or Edge 200. but now the same story is unfolding for the Blackwell generation you see it in the media report and they're constantly saying they're sold out and they truly are because all of us have placed you can imagine hyperscalers and us all of us have placed orders for Blackwell GPUs now almost probably like 12 months to 6 months previous to today's date for Blackwells that are going to get delivered later this year, early next year so you are in the queue to get those but let's say if you were to start a cloud company today or go out and place an order today, I mean, you're potentially waiting for like eight, nine months to get those Blackwell GPUs, such as the demand for those NVIDIA Blackwell GPUs, right?
19:24So yeah, that's kind of how the allocation curve goes. So you have to take certain risks to get the allocation and you have to constantly assess what your sales pipeline looks like, what the market demand curve looks like. But so far, it's always you kind of plan your GP deployments well ahead of time by both having visibility on the demand curve and also assuming that demand curve continues to stay above supply curve, which we really believe in that way. Yeah. A couple of questions. The workloads that you're running, are they primarily training workloads? So that's been an evolution. I think if you had asked me that question, like, let's say, 12 months ago, I'd have very strongly said yes.
20:09I think a lot of the new spend dollars in infrastructure, especially in the deep learning space, is going towards primarily training. And I include fine-tuning as part of the training bucket in this, especially since open source has really kind of allowed people to fine-tune a lot more now with the models being of very high quality. But even until very recently, I think training was the new major spend because for training, you need a huge chunk of compute all at once, all interconnected in a way and so on. And inference is an application-based usage, right? You know, if you are a consumer like company that is providing some application, if it takes off, great, you need a lot of GPUs.
20:52If you're just starting now, you can start off with a few GPUs. So until very recently, training was the majority spend, and that reflected in Lambda's own deployments and in Lambda's own usage as well. I do want to very specifically answer generically. The way Lambda Cloud is built, it can be used for both training and inference. That's a result of both what NVIDIA has done with its high-end GPUs, where they have made it extremely good both for training and inference. So if you look at Hoppers, if you look at Blackwells, they are extremely effective training and inference GPs. You really can't say that, oh, it's really great for inference, but it sucks for training.
21:31It's really not true. They're absolutely top of the line for both those applications. So first of all, that architecture and infrastructure itself helps both training and inference. And then the way Lambda designed its data center too, it's like, of course, we design more for training in the sense that we try to get as many GPs all interconnected together. but the same setup can be used for inference from that perspective. And so far, inference is not bottlenecked by some of the factors that you think they'll be bottlenecked by, such as latency, right? If you're a user in New York and your GPU is in San Francisco and you're using chat GPT or an image generation, you're okay because the networking time lag, the latency lag is hundreds of milliseconds.
22:17but the actual computation time is in thousands of milliseconds or even seconds. So that networking lag is not effectively your bottleneck. It's still the model processing time. So right now inference actually works very well. It doesn't matter whether you run it in a training type cluster. And as models are becoming bigger and bigger, they actually need that kind of interconnected GPUs. So in terms of the market use case, Yeah, training was the primary driver of revenue or GPU use cases for Lambda. I think if, you know, we constantly do service. It's very hard to always know, like have posts on the market exactly how much usage.
22:55But when we right now speak to our users, I think inference probably in terms of the new dollar and new GPU spin up, I think inference has overtaken training in terms of how many GPUs are spin up. Like we still have training customers. We still have fine tuning customers. but in terms of when we're deploying new GPUs and how they're getting allocated, it's more towards inference. When Blackwells do go live, though, I expect that ratio to, again, flip a little because the first foray into Blackwell will all be training customers, right? Because that's what they're going to try and get advantage of the Blackwell.
23:27So it kind of switches like that. But overall market, I think inference is starting to take over in terms of the dollar spent as more applications are coming. Yeah. And the reason I ask that, yeah, I think the same thing, too, because, you know, foundational models are the foundation, but applications on top of them are all about inference. And I've been talking to Andrew Feldman at Cerebris and Rodrigo Leung at SambaNova. Are you starting to buy those chips for the inference business? Or if not, why are you sticking with GPUs when these are so much better for inference or Grok for that matter?
24:20Yeah. So the first answer I can very explicitly clearly say, we currently only deploy NVIDIA GPUs, both for training and inference applications. And we have not yet forayed into deploying ASICs or alternative accelerators. inference demand is definitely starting to take off i think one of the key points there just quickly on that note is if you think about foundational models you know when the whole kind of ai we've especially around transformer started there were way more than three four dozen companies all trying to do foundational model training i think now it's getting concentrated the models are becoming massive like i mean you hear about 100k gpu clusters and so on but it's getting concentrated on how many people do it too right so more people are going towards fine tuning and then figuring out applications for it, which is also driving inference use case.
25:11Now to your point where companies like Cerebrus, Grok, Sambanova, even a few others, they're focusing on inference as a service and inference as an application, right? And all of them have their own kind of ways of niches. For example, Grok really focuses on the speed, tokens per second, throughput. So Rebrus can fit a big model, but is also competing on tokens per second throughput. So the speed, they compete on more speed than a GPU, Nvidia GPU could. For us, there's two things. One, these are still kind of up and coming accelerators and use cases that yet don't have that massive market economic drive for us to be the first adopters.
26:03Look, Lambda, I'll be frank, we are a growing company. I would say we have a limited amount of money in our balance sheet, and we want to basically put that money towards guaranteed revenue generating ability. And NVIDIA GPU today is that, right? And it's a universal card. Every single application runs on it. People optimize for NVIDIA GPUs. People build for NVIDIA GPUs. You know for a fact that every single application will work on that. I think it will be very interesting and it's really, I think even NVIDIA will accept that it's great to have competition because it keeps everyone on their toes and from that perspective.
26:44But from our perspective, really right now, it's more around both capital allocation and what we see as all-purpose applications, like where we want to be. We can't yet determine the niche of where do we want to really support that. So that's kind of what has led us to really focus on the NVIDIA GPUs in general. I think as a software stack on these companies become more robust such that we can deploy infrastructure with them, I think there will be considerations, of course, for this. But to be frank, from our viewpoint, I don't yet know things can change, but I think we'll be looking for guidance at other hyperscalers, at other players to first introduce this to market, see the market capture, and then maybe consider it.
27:30But for us, very clearly right now, it's, you know, NVIDIA GPUs are, I mean, if you look at the market capture of NVIDIA GPUs, it's well over 90 % of the total compute, right? So we will stay within that lane for now. Yeah. The other question is you were talking about inference using your cloud and the latency is all in the computation, not in the transmission. How far away are your biggest customers? or rather how... The distance between data centers. Yeah, I mean, are people globally using you guys to train or is it really a U.S.-based clientele? Yeah, I mean, look, I think it doesn't matter which cloud you're speaking to, the majority of revenue is coming from U.S.-based clientele.
28:35We do have, as it's always the case, we do have global customers. And the part of it is for training, truly, we have had customers from Europe and Asia because it doesn't matter. And vice versa, we have deployments in certain other countries that U.S. customers have used for training. For inference, though, I think majority of the user base for us is U.S.-based, first and foremost. and a lot of our inference deployments are in the United States, and it's the usual regions, whether it's kind of West Coast, East Coast near Virginia, Chicago area, Texas areas, kind of thing, and the new up-and-coming regions like Salt Lake City area and some other kind of regions where there's more power, like Atlanta and things like those.
29:27um but uh yeah generally it's it's it's united states based and and by the way for the latency ones for the smaller models it is starting to become more important right and again this is you know not kind of trying to harp on it but like going back to your point around grok and samanoas of the world they kind of for the the throughput that they they have the speeds that they have latency would matter right like especially for smaller models like an 8 billion parameter or now a 3 billion parameter model. I think that's where latency does start to matter. And I think if you probably, I think Andrew will agree with this or things like those, as AI starts to feed input into other AI applications, latency will also matter.
30:12Right now, even if you are the top reader in the world, the throughput speeds are high enough that throughput can be faster than yours or at least my reading speed. Let me just talk with myself. So it kind of doesn't matter to me when that output is like, whether it's faster by 10x or 2x, as long as it's faster than my reading speed, great. But as those outputs get inputted into other AI models, and there's a chain of AI models or agentic behaviors or actionable behaviors from me, I think it would really matter the speeds and latency. And that's, I think that's why these guys are also correct in saying that like latest, like throughput speed is important.
Read the full transcript
30:52So it will matter in the near future as more agentic AI applications come in. You know, people will try where, you know, we do do load balancing where like a user in West Coast is hopefully hitting our West Coast GPU and a user in East Coast is hitting. It's just not as important today for us. Like right now, we basically hit them wherever there is capacity. Like, you know, if you're sold out in our data center in West Coast, like it doesn't matter if a San Francisco user is hitting a room, just divert it to Atlanta. But as you go forward on inference applications, you will see this optimization.
31:25I guarantee you hyperscalers already do it by regions. You will see that across the board, but it's still not as important, but things move fast. And especially smaller models, as they're getting better, you will absolutely see latency matter on being closer to the metropolitan areas. But right now, first of all, most of the demand in the world is US-based because almost all startups, and again, I really don't mean to say that they're not amazing European and Asian companies, but if you look at the percentage spread out of the demand curve for compute, a lot of it's driven. Yeah.
32:03As I said, I don't know if I said it to you, but I spent a lot of my life in China, and I still travel quite a bit. I was in Armenia recently, and the AI community there is struggling because they don't have GPUs or adequate GPUs. China, of course, is under these export controls. They can't buy the GPUs. But all of these places theoretically could access GPUs in the cloud. So is that, do you see much of that? I mean, and is it practical for a Chinese company? I mean, I don't know what the regulations are in China, but to access a GPU cluster in the cloud to train a model, does that work? China is a, for China, no, because it's on the restriction list for United States.
33:09So, for example, we have to have a block on all IPs that originate from China. Oh, is that right? I wasn't sure. Yeah, yeah. And that's true because, you know, even though the block states are like, okay, you know, NVIDIA can't sell their GPUs, but even then, like, it's even up on the cloud access. Now, like, look, it is a tough one to, because, like, before Chinese companies did used to have access, you know, You have the biggest clouds in the world like Alibaba and other clouds as well that have these GPUs. And China has, if you read reports, China has figured out a way to get access in other countries and things like those.
33:49But for Lambda, very clearly, we have to follow the U.S. regulatory guidelines. And as such, with China itself, but if you're speaking about an European country or an Asian or an African country that are not on the sanction list, they absolutely can spin up a GPU on Lambda Cloud, doesn't matter where they're accessing it from and do it. We do see a decent amount of demand coming from, you can imagine, Europe. I think that's another hub in terms of the machine learning community and starting to see from other nations as well. This is part of where when you see NVIDIA's report Or do they talk about sovereign clouds?
34:31Or more than sovereign clouds, I mean, they use that word, but I think it's more like internal nation cloud. So for Indian users, like GPUs in India, for French users, GPUs in France. I think there is a lot of these things happening, both hyperscalers going into different countries and deploying this, but also local companies spinning up. But yeah, China, Russia, these are the sanctioned nations, which they definitely... There's a lot more regulatory kind of thoughts and regulatory laws that you have to understand to make sure that you're not getting used there. But for other nations, yeah. So it depends on where you're from.
35:14China, definitely. I mean, like both. They want to have access. I mean, I didn't touch upon China. I mean, other than the United States, the biggest user or if you think about demand computationally will be from China. It's just we don't see it or we don't get to talk about it because we don't see it. Because if it was a free trade, then we would probably be talking about China as probably competitive to U.S. demand usage in terms of the total usage in the ad. But we don't see it, so it's all internal there from a perspective. Yeah, and this is completely off topic, but do you think those sanctions and restrictions and shortage of getting their hands on GPUs is going to cause China to fall behind in its AI development vis-a-vis the US?
36:05I can only give you my personal opinion on this. I am nowhere close to being even a novice in geopolitical realms of things like those. But look, what I can tell you is from technical side, that's my foray or way from viewpoint. Just today, there was a model launched by Tencent, which was an open source model. I don't know if you saw it. It's a mixture of experts. based on the performance evaluations it beats llama 4 of 5b which is the current state-of-the-art open source model at least in the united states right and and so what that tells me is the answer is no uh we have had the sanctions for multiple years now um if you look at and okay so that's one one thing like if you look at what are the best video generation models there and and just now the first open source like really good quality first open source model like by genmo you know Mochi one got launched and it's an amazing open source model.
37:03That's a US based company, but we have had video generation close source models, both in the United States by companies like Runway and Pika, but then by Chinese companies like Kling and then others, right. Those models, if you look at the results, they're really, really good. Right. So I think overall it's, you know, you know, people say that, Oh, like, you know, they are maybe behind in LLMs and things like those, But really, the evidence is starting to go flip side. Now, they don't have an O1 reasoning model yet, at least not that we know of in terms of both closed source and open source. But even other than OpenAI in the United States, no one else has that either.
37:41And I think it's just about development time for even other companies to catch up to that because it's a new research stream. So I do not think from a technical perspective that they are behind. Maybe it has slowed them. I don't know what the hypothetical looks like if they had free access to the GPUs, what that would look like. But from a technical perspective, at least what I can say is the answer is no to me. Yeah. And with this Tencent model, I mean, they haven't said what chips they trained on, did they? No, no, no. I mean, certainly they were stockpiling GPUs, but at some point that supply has to run out.
38:24and I don't know. It'll be interesting to see what happens. I mean, this is the part of that, and this is like, I know that this is a little bit outside of our thing, but like outside of discussion point, but like, I think that's the interesting part. I mean, people do speak about how companies within China get access to those GPUs and none of us can comment on it in a sense definitely, whether they do or not. I think that's not in our kind of domain area, but what I can say is that if they don't have access to it And they're getting models like what the Tencent model is, that in itself is an amazing feat of technological progress from them.
39:03If they do have access, well, then in a way, then that's just the sanctions are not working, right? So I think that's kind of how you look at it. But again, and then the interesting thing is the Chinese companies have open sourced a lot of their models, which is quite an interesting tactic that they've taken, right? And I think people do use Gwen, and now I'm sure they'll use the Tencent model and then so on and so forth. Yeah, yeah. Okay, so where are you guys going? I mean, first of all, scale. You're not a hyperscaler, but how many locations, data centers do you have around the world? How many GPUs do you have deployed?
39:46And what's the growth trajectory? Yeah, I think the interesting thing is that hyperscalers are also growing so fast that I think what our current deployment would have been considered like a hyperscaler deployment just like literally three years ago. But now, like, obviously, like I think the amount of GPUs that are getting both manufactured and deployed is really unseen of, which makes sense because the demand curve. For Lambda, right, look, we have data centers in now double digits, right? Now, I want to also be transparent. It's not like hyperscale where we have like, oh, each data center is 50 megawatts, so people can multiply 50 times 12 and say you have 600 megawatts.
40:28No, because I think when we started, we had smaller data centers, you know, 3, 5, 10 megawatts. Now, as we're growing, now we have data centers 30, 50, maybe 100 megawatts in the future, right? So it's a much more of like an uneven mix of the sizes of data centers, but we have now double-digit data centers that we are managing both in the United States and what we are publicly discussing in terms of international, we have a place in South Korea that we are deploying GPUs in as well. In terms of number of GPUs, I think, you know, the scale keeps on growing. I think, you know, for Lambda, we are, I mean, there have been public reports and I think they're pretty close to the mark in terms of we are in multiple tens of thousands of GPUs.
41:14and interestingly enough, it's not as interesting a number because when you hear singular companies like even Meta and X who are not cloud companies having GPUs in hundreds of thousands. So that's kind of where we want to go towards scale if you think about the growth rate, Craig. So it's about how do we get more and more computation under our management and then build the software stack on top of it. So, yeah, Lambda is in, I would say, somewhere in the middle tens of thousands today. And then, you know, ideally really like gunning for that kind of six-figure mark in terms of number of GPUs deployed, hopefully in the next year or two as our data centers, the bigger data centers come up and as we get more and more capacity under us.
42:08So that's kind of what the scale is. in terms of growth rate, look, I think for us, starting from a lower denominator, the growth rate looks amazing, right? You know, I think, you know, kind of last year, you know, you look at almost a magnitude growth. Now we're talking about, you know, over 100 % in three figures percentage growth rates. We want to continue that. We want to be near that, like, you know, over 100%, 150 % growth rate mark. It'll be interesting to see what the, you know, this is the point where we constantly think about, like, how do we keep on innovating both on software side as well as really offering what people want?
42:48Because to continue that, you think about, like, what hyperscalers are deploying in terms of GPUs as well. You know, it's going to have competition. But underlying thesis still remains. I think, you know, everyone still thinks that, hey, there's going to be some plateauing of the compute curve. personally look I have roasted their glasses on and so you know of course people should take you know what I say with a grain of salt but like I do not believe so that there is a plateau I think that with the rise of inference demand with the rise of new types of model structures whether that's reasoning curve whether that's multimodal whether that's video generation models I think you know where we'll end up is just a continued rise in demand and as long as the underlying market force stays strong growth rates will continue.
43:35I mean, think about the video models, right? You are in podcast and video generation industry. Someone like me, I don't take out time to generate videos of my own because I know that there's millions of creators in the world, but that's still a barrier to entry among the 8 billion population worldwide. But imagine if you could just type something and say, hey, make Mitesh sound very smart about, I don't know, investing. And I upload my photo and it generates a nice video about me being as smart as Warren Buffett in investing in value stocks. And imagine how many people will generate that. Or I want to see myself as Luke Skywalker in a Star Wars movie.
44:17And keeping aside the legal content generation, all those things aside, you think about what that unlocks, whether it's content generation, whether it's reasoning model, you think about how AI today is used in a lot of, you think about kind of activities that are not doers, right? You think about chatbots and stuff. Yes, you have interactions with it, you get answers, you use it in legal and support services. But when we go into doing what people, they use the buzzword of either agent applications, so like writing code and taking your appointments and things like those, but all the way from that to literal physical activities, whether in robotics or whether in any kind of machining or whether in civil engineering jobs, jobs, chemical engineering jobs, jobs that require complex problem solving that comes from reasoning.
45:06I think that the amount of computational required, like the curve will in fact, keep on getting steeper as we figure it out. Now, whether it takes some time to figure out, that's sure that that might be the, that might be the air pocket that will cause it to people to think that it might plateau. But personally, I mean, as I said, I truly fall on the extreme side of that belief of where we are going with the demand. And on the agentic side of that, as you were saying, that latency will become much more important when you have networks of agents working together. And I wonder then if that means that, you know, eventually every community will have its own AI data center, you know, GPU data center serving its community.
46:07Where are your data centers? You don't own the data centers, right? you're placing machines in data centers. Is there a logic to building your own and operating your own data centers? Or is that just another business? I think the way we look at Lambda for us, we are much more focused on the layers about data centers, so putting an infrastructure, making it run very well, putting software stack on top of that infrastructure layer for people to use it as a cloud service, both for training and inference and then keeping on going up the stack. But I think if you look at the marketplace, I mean, hyperscalers have certainly gone towards building their own data systems as they've gotten bigger in size.
46:51So never say never. I think this at a certain scale, operationally starts to make sense to get into either a building yourself or getting into a partnership to build one. But for today, Lambda, yeah, absolutely. We work with the big data center infrastructure companies. They're excellent at it. They're efficient at it. they're fast at it, why do you think about all the digital infrastructure players that are out there? And I think, so we do like a long-term leasing partnerships with them to deploy this. In terms of the local data center that you talked about, I think, look, I think the extreme version of it is, you know, obviously people will move towards like, you know, local kind of inference at the edge.
47:34There's going to be absolute part of that. I always think of it as like, you know, not anything, but like kind of like, you know, like different sizes of brain, like small brain problems, medium brain problems, big brain problems, right? Small brain problems solved on your edge device, whether it's laptop, phones, glasses, whichever edge devices that humanity creates in the future, right? I think then you think about kind of, like you talked about either local data centers in areas or even kind of like generators, right? You might have computing devices per house to run your local household robots and decision-making.
48:11or maybe they'll be all on right now all of the humanoid robots it's all on either local the robots itself like the computation is there or they interact with the cloud and things like those so it'll depend on kind of how that latency requirements progresses but I won't be shocked where people have like kind of we have a local water heater and local generator like you kind of have that on a house level as well but I don't I think that's too far away still I think but like you were absolutely spot on that I think what we'll have is, I do believe, I think as more application, like, you know, you think about whether it's, you know, robotics application, agentic applications, I think, and AI kind of daisy chaining into other AI, I think you will see more and more local and local creations of data centers that go there.
49:01And then there will be absolutely the massive ones. I mean, right now, all the talk in the industry, and it is rightly so on this, you know, 400 megawatt data center, half a gigawatt, a gigawatt data center, nuclear power to power a gigawatt, solar renewable power to be a backup, figuring out ways to get it to multi-gigawatt. I mean, if you think about the level of the data center powers where we are going, you think about three, four, five gigawatts, that's like a New York City-style power consumption by a data center that is built in a hundred acre plot of land. I mean, the density of power that is required there.
49:37And And that's where a lot of focus today is because the big architectures are getting built here, you know, 100 ,000, 400 ,000 GPUs kind of structures that are going to come in the market in the next year or two years. So that's where the focus is, I think. And because the underlying thesis is scaling is critical for making better AI models. I think once that moves on to how do we assign more computational spread to multiple areas, more reasoning, smaller models, but like doing a lot more of it a lot faster, then the conversations will grow into like, hey, probably a conversion of office real estate into data centers, like, you know, all that thing.
50:18So there will be waves of this kind of talk points around how we build data centers and power areas. Yeah. Wow. Fascinating stuff. Is there anything I haven't talked about that you want to talk about? I think from my perspective, what is interesting, like, and again, because for me, the most interesting thing are kind of the technical progresses that are happening, both from the model architecture and the applications perspective, right? Like we today see what we see, like whether it's chatbots, whether it's support agents, whether it's code writing assistant tools, you know, people question sometimes the quality of the outputs and then so on.
51:02But just for me, the essential thing is that that's just like any technology, right? It takes time. Like if you look at an iPhone in 2009 or 2007, you know, you'll think it's a shitty iPhone. It's like, wow, this is so slow. You know, you can barely do call. You can do calls plus maybe a little bit of mail and things like those. Right. I think, I think every cycle, the technology pace has become faster. You know, what it took for us to, to get from, you know, like getting really internet being invented to people using internet at a wide scale. It took like 10, 20 years of life cycle. Phones or smartphones was 10 years.
51:40Web services and cloud was like eight or seven years. I think AI will be faster in terms of how it gets really percolated in everyone's life. And I think the most interesting thing for me is to really keep track of some of these kind of interesting points. Chatbot was one. I think now for me is the video models, how they get integrated for us. I think that will be the next kind of milestone point in terms of how easy it is to do multimodal and video models in our day-to-day life, how easy it is to use it. then it'll be this like what 01 did was a reasoning models around you know kind of complex problems and then you'll be like the low latency on the edge devices kind of these are kind of the step function points that i'm looking forward to and it's very interesting how quickly we are moving through this right and i think even when we go from like you know what was gpd4 was state of the art to now cloud 3.5 now back to 01 these are happening in months so so from my perspective I think it's important that people stay very optimistic about this technology.
52:43Of course, we have to have controls around, a little bit of controls around how people, it doesn't go on the wrong side of things. But I just want to make sure there are a lot of listeners to your podcast. I want to encourage people to be in the know about this. I think it's a very fascinating field. And then lastly, of course, on Lambda side, I think we are here to really, underlying it all, of course, we want to build a big, successful, long-term business. And we really care about that. It's kind of our baby to build that. But underlying of it all, we really want to be the ones that provide a lot of, or one of the ones that provide a lot of compute and ease of compute to people that are building the next step of applications for this world.
53:30And so, you know, to that point, I think, you know, we really focus on the deep learning aspect of it. We spend a lot of time building for it and for a lot of listeners. Again, I would encourage you to try Lambda Cloud if you're building any deep learning application out there. Okay. And how do they reach Lambda Cloud? What's the URL? Well, our classic URL is lambdalapse.com. but our fun one is gpus.com. We own gpus.com, so we've won for a long time. That was a coup, yeah. Exactly, a cool grab. That was done six years ago. So at least on that front, we were ahead of the curve for grabbing that domain.
54:16What does the future hold for business? Ask nine experts and get 10 answers. Bull market, bear market, rates rising or falling, inflation going up or down. Can somebody please invent a crystal ball? Until then, over 40 ,000 enterprises have future-proofed their business with NetSuite by Oracle, the number one cloud ERP, bringing accounting, financial management, inventory, HR into one fluid platform. With one unified business management suite, there's one source of truth, giving you the visibility and control you need to make quick decisions. With real-time insights and forecasting, you're peering into the future with actionable data.
55:11If I were a larger organization, this is the product I'd use. Whether your company is earning millions or even hundreds of millions, NetSuite helps you respond to immediate challenges and seize your biggest opportunities. Speaking of opportunities, download the CFO's Guide to AI and Machine Learning at netsuite.com slash ionai. That's netsuite.com slash ionai, E-Y-E-O-N-A-I, all run together to get the CFO's Guide to AI and Machine Learning. The guide is free to you at netsuite.com slash ionai, netsuite.com slash ionai.
From the publisher
This episode is sponsored by Netsuite by Oracle, the number one cloud financial system, streamlining accounting, financial management, inventory, HR, and more.
NetSuite is offering a one-of-a-kind flexible financing program. Head to https://netsuite.com/EYEONAI to know more.
In this episode of the Eye on AI podcast, we dive into the transformative world of AI compute infrastructure with Mitesh Agrawal, Head of Cloud/COO at Lambda
Mitesh takes us on a journey from Lambda Labs' early days as a style transfer app to its rise as a leader in providing scalable, deep learning infrastructure. Learn how Lambda Labs is reshaping AI compute by delivering cutting-edge GPU solutions and accessible cloud platforms tailored for developers, researchers, and enterprises alike.
Throughout the episode, Mitesh unpacks Lambda Labs’ unique approach to optimizing AI infrastructure—from reducing costs with transparent pricing to tackling the global GPU shortage through innovative supply chain strategies. He explains how the company supports deep learning workloads, including training and inference, and why their AI cloud is a game-changer for scaling next-gen applications.
We also explore the broader landscape of AI, touching on the future of AI compute, the role of reasoning and video models, and the potential for localized data centers to meet the growing demand for low-latency solutions. Mitesh shares his vision for a world where AI applications, powered by Lambda Labs, drive innovation across industries.
Tune in to discover how Lambda Labs is democratizing access to deep learning compute and paving the way for the future of AI infrastructure.
Don’t forget to like, subscribe, and hit the notification bell to stay updated on the latest in AI, deep learning, and transformative tech!
Stay Updated:
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
(00:00) Introduction and Lambda Labs' Mission
(01:37) Origins: From DreamScope to AI Compute Infrastructure
(04:10) Pivoting to Deep Learning Infrastructure
(06:23) Building Lambda Cloud: An AI-Focused Cloud Platform
(09:16) Transparent Pricing vs. Hyperscalers
(12:52) Managing GPU Supply and Demand
(16:34) Evolution of AI Workloads: Training vs. Inference
(20:02) Why Lambda Labs Sticks with NVIDIA GPUs
(24:21) The Future of AI Compute: Localized Data Centers
(28:30) Global Accessibility and Regulatory Challenges
(32:13) China’s AI Development and GPU Restrictions
(39:50) Scaling Lambda Labs: Data Centers and Growth
(45:22) Advancing AI Models and Video Generation
(50:24) Optimism for AI's Future
(53:48) How to Access Lambda Cloud




