The GPU Myth: State of AI Compute 2026 | Stephen Balaban

18 Jun 2026 · 1 h 14 min · 37 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

“Neocloud boom” and the 2026 state of AI compute, focused on why GPU cloud is not a commodity, how demand and scaling laws drive underbuilding, and what it takes to build/finance gigawatt-scale AI data centers.

Guest background

Stephen Balaban is co-founder and CTO of Lambda, a neocloud powering AI workloads. He founded Lambda in 2012 after working on facial recognition deep learning; Lambda later built consumer AI hardware (a camera baseball cap) and an image generator (DreamScope). Lambda grew from workstation sales to a cloud business (now under a billion $ run-rate) and has fully exited hardware.

Key claims

  • GPU compute rental indices can mislead because on-demand and long-term rates move differently.
  • Cloud compute remains vertically integrated (land, construction, HPC design, virtualization, orchestration), so it won’t commoditize like “GPU rental.”
  • The industry is “generally under building” due to ongoing scaling laws and expanding model use cases.
  • Main bottleneck is land with megawatt utility commitment plus data center MEP equipment.
  • If models become 10x more efficient, demand expands (more tokens), so compute demand doesn’t collapse.
  • GPU usable life exceeds accounting depreciation; “five-year lifetime” naysayers are wrong.

Notable examples

  • Lambda’s “one-click cluster” can partition 16 to 4,000 GPUs via orchestration software.
  • Frontier distributed inference/training uses spine-leaf networking (InfiniBand/Ethernet) and sharding (e.g., mixture-of-experts).
  • Cooling/water claims: modern direct-to-chip liquid cooling with dry coolers; misinformation about water use.
  • Financing approach: asset-based loans for GPU/offtake and credit underwriting of on-demand cloud.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

Debunking NeoCloud Commoditization

1:15 to 1:50

Discussion about misconceptions surrounding NeoClouds and GPU compute.

“Please enjoy this amazing and very educational conversation with Stephen.”

The Complexity of Cloud Services

1:50 to 2:15

Exploration of why cloud compute is not merely a commodity service.

“It is a very complicated, highly vertically integrated type of service that spans everything from land entitlement, construction, HPC, high performance computing design, software virtualization, cloud services on top.”

The AI Cloud Service Model

2:15 to 2:40

Overview of the AI cloud service model and its unique characteristics.

Rental Rates and Market Dynamics

2:40 to 3:24

Analysis of rental rates for GPUs and market trends affecting them.

“But really what it was, was it's a cloud service designed for the age of AI.”

Lambda's Unique Offerings in NeoCloud

3:24 to 5:30

Discussion on Lambda's software and innovations in the NeoCloud space.

“if not increasing long-term rental rate and very consistent and increasing on-demand rental rates.”

Future of the NeoCloud Ecosystem

5:30 to 6:16

Predictions on the competitive landscape of the NeoCloud ecosystem.

“And then there's, as you mentioned, innovation on the finance side of things where, you know, we're coming up with new and unique ways to finance, underwrite, package these large scale capital projects, really.”

Demand for AI Compute and Scaling Laws

6:16 to 8:18

Insights on future demand for compute driven by scaling laws and model capabilities.

“And I think that the fundamental reason for that, kind of going back to what drives, I guess, market structure.”

Bottlenecks in Building Lambda Labs

8:18 to 10:41

Discussion on current bottlenecks within Lambda Labs and the industry.

“And as long as that continues to hold, I think that we still have in store for us.”

Community Response to Data Centers

10:41 to 14:01

Exploration of community concerns regarding data centers and their impact.

“Where's the main bottleneck these days that you're experiencing building Lambda Labs?”

Understanding Data Center Misconceptions

14:01 to 15:00

Explore common misconceptions about data center water consumption and cooling methods.

“You'll see people talking about how data centers consume a lot of water.”
Show all 37 chapters

Compute Units Explained

15:00 to 17:12

Learn about compute units, energy production, and their measurement in data centers.

“People talk about things like flops and GPU hours and tokens and MFU.”

Maximizing Chip Utilization

17:12 to 19:20

Discover strategies for maximizing the value extracted from chips in data centers.

“If two companies have the same chip fundamentally, how do they extract more value from it?”

Networking in High-Performance Computing

19:20 to 21:43

Understand how high-performance computing clusters operate and their networking strategies.

“Actually, a lot of NeoClouds are in that position where they don't even have the infrastructure to be able to run a real cloud service.”

Cost Factors in Cloud Services

21:43 to 23:53

Examine the major cost components associated with cloud services and compute costs.

“generally speaking when you're doing a training run you might think there might be some sort of split between the backwards pass and the forward pass on the model.”

NVIDIA's Dominance in Chip Manufacturing

23:53 to 26:11

Discuss the competitive landscape of chip manufacturers and NVIDIA’s market position.

“And within that is basically some sort of bill of materials for the servers that are in the data center, which is by far and away the biggest portion of the cost.”

Optimizing Performance with CUDA-NN

26:11 to 28:00

Learn about the advantages of NVIDIA's CUDA-NN library for matrix multiplication.

“Do you think that today or in the near future, we're going to be in a multi-silicon kind of world?”

Optimizing AI Compute with Networking and Storage

28:00 to 29:19

Learn how advanced networking and storage solutions optimize AI computing environments.

“and QDNN means that you don't have to go and do the optimization yourself.”

Understanding Cloud Infrastructure for AI

29:20 to 31:18

Explore the components and coordination required for effective AI cloud infrastructure.

“whether it's the data that you're using to train with or whether it's the data that's coming in and streaming in from your end customers.”

Challenges of Building a Neocloud

31:19 to 33:39

Discover the complexities and investments needed to establish a functional neocloud.

“Let's say you've got a cluster of 10 ,000 GPUs.”

Vertical Integration in Data Center Construction

33:40 to 35:38

Learn about the benefits and strategy behind vertical integration in data center operations.

“And so like that is, and then to have it all work with the storage.”

Market Focus and International Operations

35:39 to 37:18

Understand the geographic focus and operational strategy of Lambda's data centers.

“And it's been great because we've been able to kind of, again, bring that engineering mindset to this problem, which was historically mostly run by people in real estate.”

Latency and Governance in AI Compute

37:19 to 40:04

Examine how latency considerations and governance affect AI compute services.

“I get this question a lot and people, they're like, well, does latency matter?”

Financing AI Compute and Market Dynamics

40:05 to 42:00

Explore how financing structures and market demand influence the AI compute industry.

“And there's a vibrant private and there's a vibrant credit market for that.”

The Neocloud's Financial Market Maturation

42:00 to 43:38

Explore how the financial market for computing units is evolving.

“We have GPUs that we've commissioned, and we're one of the earliest neoclouds.”

Lambda's Origin Story and Early Days

43:38 to 45:40

Learn about Lambda's beginnings in facial recognition and deep learning.

“But I think for right now, people are starting to realize that it's a great credit investment.”

Transitioning to Hardware and Cloud Services

45:40 to 49:12

Discover how Lambda evolved from software to hardware and cloud computing.

“We launched this face recognition API, got a couple thousand users, but it wasn't really generating a ton of cash.”

Navigating Challenges During COVID

49:12 to 52:20

Understand how Lambda faced difficulties during the COVID-19 pandemic.

“We were so, so scared that doing this CapEx was going to put us out of business.”

Leadership Transition and Future Vision

52:20 to 56:00

Examine the decision to bring in a new CEO and its implications for Lambda.

“Hardware companies, the docks were closed.”

The Experience of CEO Roles

56:00 to 57:29

Explore the secret feelings of founder CEOs about their roles and responsibilities.

“And I think there's plenty of founder CEOs who absolutely love every aspect of their CEO job.”

Rapid Data Center Deployment

57:30 to 59:28

Learn about the strategies for rapid deployment of data centers and the unique challenges involved.

“What was XAI's record like when they launched?”

Neural Software: The Future of Computing

59:29 to 1:02:35

Understand the concept of neural software and its implications for how we interact with computers.

“You had a quote where you said that AI won't write software, it will become the software.”

Agent-Driven Development Paradigms

1:02:36 to 1:07:29

Delve into the future of software development with self-assembling software and agent interactions.

“Vibe coding and a neural operating system or neural software.”

AI Factories and the One Person One GPU Vision

1:07:30 to 1:10:03

Discuss the concept of AI factories and the vision of one person utilizing one GPU.

“The other side of that, eventually, once the models get smarter, I think that they're quite not quite there yet.”

The Evolution of Personal Computing

1:10:03 to 1:11:33

Explore the timeline of personal computing from 1984 to future predictions.

“The Macintosh came out in 1985 or 1984, excuse me.”

The Future of GPU Accessibility

1:11:33 to 1:12:05

Discussing the need for computational power in daily life and its implications.

“more to just do their daily work, enjoy life, whether it's getting access, whether it's getting entertained, whether it's being productive, whether it's being creative.”

AI Insights: Overhyped vs. Underrated Concepts

1:12:05 to 1:13:16

Analyzing AI ideas, distinguishing between overhyped and underrated concepts.

“What is one idea in AI that is overhyped?”

Agent-Based Workflows and Software Development

1:13:16 to 1:14:00

Delving into the practical applications of agent-based workflows in software development.

“Inflation would be, it wouldn't be inflation.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00It's pretty clear that we have an amazing system that can take in money and output software. The people who are the naysayers, you're going to throw these GPUs out in five years, are completely wrong. They're completely wrong and they've been wrong the entire time. We continue to be generally under building. Most people that are sort of in leadership positions at neoclouds or within the market have been recognizing this insatiable amount of demand for large language models to do everything from being an assistant to code generation. We continue to see no end to the scaling laws. Hi, I'm Matt Turk from FirstMark.

0:37Welcome to the Matt Podcast. My guest today is Stephen Balaban, co-founder and CTO of Lambda, one of the top neoclouds powering the AI boom. This episode goes deep on the physical layer that everything else in AI runs on. We get into why GPU compute was never actually a commodity, how you finance billions of dollars of data centers and chips, why a 2023 H100 can be more expensive to list today than when it was bought, and what it actually takes to stand up a gigawatt scale AI factory. We also cover Lambda's wild origin story, from a facial recognition startup to a baseball cap with a camera in it to a near billion dollar cloud business today.

1:15Please enjoy this amazing and very educational conversation with Stephen. There was a moment in time in Silicon Valley a few years ago, if you had asked most people, they would have said that NeoClouds were going to be a commodity, in particular because GPU compute was going to get commoditized. And if you fast forward to today, it seems to be exactly the opposite, both Lambda, but several of your competitors seems to be absolutely a ripping. So what is it that naysayers got wrong then and continue to get wrong today? The big thing is that cloud compute is not a commodity service. It is a very complicated, highly vertically integrated type of service that spans everything from land entitlement, construction, HPC, high performance computing design, software virtualization, cloud services on top.

2:15and there's a reason why the biggest companies in the world these multi-trillion dollar market cap businesses whether it's amazon microsoft google oracle are all in the cloud computing business is because it's a great business and so i think that's like probably the fundamental thing that was misunderstood is that oh this is somehow a little bit different than a normal cloud service But really what it was, was it's a cloud service designed for the age of AI. But there is some element of commoditization, right? The price of rental of a GPU is going down. But what you're saying is that to some extent it doesn't matter because it's only one layer of the cake.

2:57Yeah, so when you look at, for example, I think it's actually worth doing is to try to like kind of dig into some of the methodology on, for example, an index. like there's the index that's on Bloomberg for H100 rental prices. And what we're actually seeing in the market is that, first of all, there's two different rates. There's a public cloud on-demand rate, and then there's a long-term rental rate. And I think that some of these indices don't properly take that into account because what we're actually seeing is a very consistent, if not increasing long-term rental rate and very consistent and increasing on-demand rental rates.

3:43And so what happens is if the index mix, for example, if the methodology and the index biases towards long-term contracts being a bigger part of the volume, that will look like a decline in the index when the reality is it's just a decline in the mix that the index is covering. Fascinating. So I'm curious about your thoughts as a key leading player in the NeoCloud ecosystem about how you see the market evolve. How much of the competitive advantage that you guys are building and other players are building is based on technology versus financing race? There's a few different layers on it, which is there's a lot of differentiation and work that's being put into, for example, the cloud software orchestration layer, which allows us to, for example, take a very large scale GPU cluster and partition it up for our customers.

4:41So we've got, for example, our one-click cluster product that allows us to do that. And that's something that's quite unique in the NeoCloud space. Most of the other NeoClouds either don't have the ability to launch a cluster from their website or max it out, let's say, 32 GPUs. Whereas Lambda's designed a piece of software that allows us to give you anywhere from 16 up to 4 ,000 GPUs in a web interface. And then there's innovation on the data center construction and design side of things, which is also really important, right? Because that's like the physical layer underneath the high performance computing equipment.

5:20And, you know, we're working on a lot of different ways to dramatically reduce the time it takes to construct and stand up new megawatts. And then there's, as you mentioned, innovation on the finance side of things where, you know, we're coming up with new and unique ways to finance, underwrite, package these large scale capital projects, really. And so I think it's like innovations happening on every layer of the stack. And it's a very complex coordination style business. Yeah. And do you think that ultimately the neoclid ecosystem becomes a winner take all or is there this room for multiple very large players?

6:05No, I think it's absolutely room for multiple very large players, just like the traditional cloud business has shown that there's room for multiple large winners and multiple large players. And I think that the fundamental reason for that, kind of going back to what drives, I guess, market structure. And I'd say, generally speaking, when you have an industry that has technology moats and capital formation moats and economic moats, that tends to be oligopolistic in its market structure. When you have markets that have more sort of network effect moats, those tend to be a little bit more single winner take all.

6:48What are the various scenarios in your head as you think about the future, about how it all play out? Are we overbuilding? Are we underbuilding? Nobody knows. How do you think about it? Well, I think that we continue to be generally under building. And most people that are sort of in leadership positions at neoclods or within the market have been recognizing this sort of insatiable amount of demand for large language models to do everything from being an assistant to code generation. you know you you can kind of look back to some of the talks that i've that i've given in the past around i kind of called hey in a couple of you know months to years we're going to be a point in time where you can put money in and get software out the other end and now at that at that point in time when you're predicting when i was predicting that it was maybe not why not as widely held of a belief.

7:54But now with, you know, let's say with the release of Opus 4 or 5, I think it's pretty clear that we have an amazing system that can take in money and output software. And I think the part which makes me feel so confident that there's going to continue to be demand is that we continue to see no end to the scaling laws, which are like the underlying line idea that you put more compute in and you get better intelligence levels out of your models, you know, as you increase the capacity of the model and train it with more compute, train it with more data, you get more intelligence out. And as long as that continues to hold, I think that we still have in store for us.

8:38It's hard to predict exactly when scaling laws might start to reach sort of a diminishing marginal return type of part of the curve. But right now, it's very clear that we're going to continue to see more and more and more capable models that is kind of expanding the cone of the addressable market, right? Like originally the cone of the addressable market was, all right, this is going to be helpful for customer support. It's a sort of substitute good for Google search and for other search online. And then now it's like, well, this is a substitute for a lot of software engineering roles or a huge augment to software engineering roles.

9:19And so as that cone expands, the total market and the demand for compute expands. And I think that we're continuing to underestimate it. Do you worry about model training and model inference becoming, I don't know, 10x more compute efficient and what that would mean in terms of the buildup? I think that generally speaking, what you're seeing is if, let's say you do become 10 times more efficient, I think that that just means that everybody is able to process 10 times more tokens and there's still the same fixed amount of compute in the world at any given point in time. And so in the early days, it's funny, we used to talk a lot about this back in, let's say, 2017.

9:55Oh, well, maybe there's going to be some new type of model, let's say, that will look more like a random forest model, which the audience might, some members of the audience might know, you can kind of train a random forest model on a MacBook, right? And there was this concern that was kind of persistently raised around like, well, okay, what happens if you have this sort of like adjacent disruption on the model side of things? And so far, we haven't seen that. And again, everything that we're building towards is sort of based on these scaling laws, which is really about scaling up this architecture.

10:32So I don't really foresee a very likely outcome where we have this huge model disruption that would cause a decline in the demand for compute. Where's the main bottleneck these days that you're experiencing building Lambda Labs? Is that GPU, power, electricity? So I always say that bottlenecks are always like kind of local before they're global in terms of, you know, one development might be bottlenecked on, let's say, generators or on UPS systems. That's a function of like the sort of idiosyncrasies of the site. But broadly in the industry, the thing that is the main bottleneck is basically land-powered shell, which is basically land that is entitled to have a certain amount of megawatt commitment from a utility.

11:26And then, of course, the data center and the mechanical, electrical and plumbing equipment, the MEP equipment that goes into that data center. And so that's the main bottleneck that we're seeing in the industry right now, I'd say, across the board. How real is the movement against data centers from the global community? And how do you think about how to respond to it? Well, it's certainly it's like very popular in the news right now. I'd say that it's definitely very real. I mean, I think that rightfully communities that host any type of large capital project, whether it's a power plant or a solar farm or a data center or a distribution center, right?

12:11Those communities want to have a seat at the table. I'd say in general, though, I spend a lot of time reading through a lot of the comments from communities and people want jobs. They want tax revenue. Any major capital development is going to bring a lot of tax revenue and it's going to bring a lot of jobs and it's going to bring investment into their community. And what they really are voicing, I think, is one is having a seat at the table while this stuff is being developed. I think that's an important thing just to have their voices heard and that the developer is coming in and actually understanding the community.

12:50The other thing to kind of, I think, keep in mind is that there's a lot of misinformation out there. So, for example, every single modern deployment of, let's say, a Blackwell class or Rubin class GPU, you know, the VR, GBN VR GPUs, these are oftentimes in a closed direct-to-chip liquid cooling system that's connected to a dry cooler, which means that there's almost zero evaporation. It's not using evaporative cooling. It's using a dry cooler system that does not consume a lot of water. And on top of that, most of these data center developments are bringing a ton of power to the grid. They're either standing up behind the meter power, they're standing up and bringing battery electric storage systems to the grid, and they're bringing all these like sort of ancillary benefits that strengthen and fortify the grid.

13:45And also, you know, eventually in the long term will maintain the costs that are being experienced by the community. And so I actually think that there's a very clear path towards, you know, maybe spreading more of that. The facts around what does the data center bring, because there's a lot of misinformation. You'll see people talking about how data centers consume a lot of water. Well, an evaporative cooling tower might evaporate a lot of water, but practically no new builds in the United States are using evaporative cooling for doing these closed loop direct to chip liquid cooling systems. Do you think we do a terrible job as an industry explaining this to the broader world?

14:28Because like those things keep coming back and they seem to be accelerating. But then when you have the discussion, from a technical standpoint, a lot of this is just simply based on museum information, as you just said. I think that everybody's trying to get better at that kind of communication. And it just takes some clear thinking, writing down what are the benefits, writing down what are the costs, and presenting that clearly and plainly to a community so they can make a good decision about, you know, what kind of jobs and what kind of development they want in their communities. Let's open the hood for a minute.

15:03People talk about things like flops and GPU hours and tokens and MFU. What is the best way to think about a compute unit? Yeah, it's interesting. You know, you said a few different terms and I always like to kind of break it down from like a physics perspective into like the SI terms. So, OK, on the left hand side is all of the energy production. And then on the, you know, my right hand side is sort of tokens being consumed by somebody. And, you know, maybe you can even have the application layer on the far right of that that's using the token. So on the left hand side, you've got either photons coming in per second or molecules of natural gas coming in per second.

15:49And then that through a power plant or a solar farm gets converted into joules per second, which is a measure of electrical power production. And then the joules per second, obviously in engines, there's a level of an efficiency and that's an engine efficiency. It's interesting because like the MFU percentage is kind of like an efficiency up on the higher end of that chain. The power plant or the solar plant then converts that into joules per second, which is watts, which is consumed by the entire data center. The data center itself needs to cool itself, and that's the PUE, and that's actually the efficiency metric that you can use to measure a data center on.

16:32And then you put the servers and all the different networking and storage gear in, and that's producing floating point operations per second or flops per second. Okay. That is what gets consumed. The flops per second capacity is what gets consumed by, let's say, a model builder when they're training a model or when they're inferencing a model. And that gets turned from flops per second into the tokens per second. Then on top of that tokens per second, you might have some level of efficiency that the end customer is actually, you know, turning those tokens into real actual intelligence. That's like the entire pipeline, I would say, from end to end.

17:11Super helpful. If two companies have the same chip fundamentally, how do they extract more value from it? What needs to happen to maximize the usefulness of that chip? If you look at the cost structure of, let's say, one GPU hour of time, we're talking about H100s, the largest part of that cost structure is the depreciation that is associated with that GPU hour. And basically, you can think of a utilization metric as being kind of a multiplicative factor on that. So one over the utilization. So if you if you use your capital asset 50 percent of the time, you will have on a per hour basis twice one over zero point five, the amount of per hour depreciation expense associated with that.

18:00And so I think that the number one way that companies are, you know, sort of gaining a unique advantage is, well, how can I build a cloud product that is beloved by people that is going to drive a high utilization? And, you know, in addition to that, the market, as we mentioned earlier, for on-demand compute, basically, the retail pricing is obviously much higher than the wholesale pricing. So the retail is like on-demand, spin up a GPU, spin down a GPU, normal cloud service. The wholesale is sort of buying 10 ,000 GPUs for five years, for example. And so one of the things that we do at Lambda is really try to figure out, hey, how can we sort of get the most dollar utilization and percentage utilization out of the capital deployments that we do?

19:00And that's by making great cloud software that makes it easy for somebody to spin it up and down. So, for example, if you don't have that cloud software, you can't rent. You can't extract a retail pricing, right? You cannot rent it out to somebody for an hour because you just simply don't have the means to be able to do that. Actually, a lot of NeoClouds are in that position where they don't even have the infrastructure to be able to run a real cloud service. So you have GPUs, but a big part of how those data centers work is transforming GPUs into networks of GPUs. Do you want to explain at a high level how that works?

19:37The general idea is that you've got a large scale, high performance computing cluster of a bunch of... you know, let's say NVIDIA GB300 and VL72 racks. That's 72 GPUs all networked together via NVLink. And then there's a connection between the racks that's either InfiniBand or, you know, high-speed Ethernet. And that is essentially what's called a spine-leaf topology, which is basically a way to say, hey, this is a completely non-blocking. Every port on every GPU can talk with every other GPU in the network. It's fully connected and it's able to provide maximum bandwidth between every individual GPU.

20:26And that cluster is useful for training large models. It's also useful for inferencing. So frontier inference, as we sometimes refer to it at Lambda, is basically, you know, very much a distributed inferencing problem where they actually will, you know, fragment or shard the model. There'll be some sort of sharding strategy for the model where it can be essentially run on multiple GPUs and it uses that high-speed InfiniBand or ethernet interconnect to to do that communication and so what is a frontier in france is that in france uh for the most advanced reasoning models like the more demanding yeah well you know it's not necessarily associated with reasoning models so much as like just a very large frontier model that is you know kind of the domain of let's say three companies in the world or four companies in the world when they're doing their inference it's a very complicated thing that is is fully utilizing all of the interconnection that's available and what you describe for frontier in france is that conceptually the same thing as what happens for training this concept of just distributing a task massively across a bunch of gpus what happens during a training run from a compute standpoint generally speaking when you're doing a training run you might think there might be some sort of split between the backwards pass and the forward pass on the model.

21:58And the backwards pass might be, let's say, two-thirds or more of the compute and the forward pass, which is basically the same thing as inferencing, is the remainder. And one of the realizations that I think has been made over the last bit of time is that the type of infrastructure that you'd want for doing a large scale training run can be reused to do the inferencing of that model. And what I mean by the sort of frontier inference and the fact that the inferencing is being done in a distributed way, you'll have a mixture of experts model and there'll be different basically starting strategies for how you put those experts onto different servers and to different GPUs.

22:46And the models can be very large. They may not fit on one single rack or they may not fit on one single server. They might need to be distributed across different servers to even just do the forward inference pass. And so that's where distributed frontier inference comes into the picture. Because if you're doing a small model, let's say Llama, that the users might be familiar with, or some of the quantized small models can fit on a single GPU. Okay, well, let's just say that like Opus and ChatGPT 5.5 can't fit on a single GPU. And when we think about compute costs, what costs the most money? Is that model size, is that memory bandwidth?

23:35Is that latency? Does context window and like those very, very large context window, do they change anything to the compute cost? What costs the most money? As I mentioned, the biggest component of the unit cost for a cloud service like this is the depreciation expense. And within that is basically some sort of bill of materials for the servers that are in the data center, which is by far and away the biggest portion of the cost. If you talk about the capital stack, let's say, you can go back down to power generation, two to three million dollars a megawatt, two to three billion dollars a gigawatt for a power plant.

24:17The data center is between 10 and 15 billion dollars a gigawatt for building the data center. And then the compute, the servers can be anywhere from 35 to 45 billion dollars a gigawatt. And within that, so you can see the server portion is obviously by far and away the largest. And that's like a big part of the depreciation expense. And then within that, obviously, you have the sort of server and cluster bill of materials, which is primarily the GPUs. if you were to kind of break down nvidia's uh bill of materials then you know you can kind of get better allocation towards uh where those costs are coming from but certainly in the most recent period of time memory expenses you know memory has gone up a lot in price and uh you know there's there's very few vendors right you know for hbm memory with samsung hynix so you guys are a big Nvidia shop at a precise level.

25:25You mentioned some of the names, but like which chips do you use mostly? What's your kind of chip stack? Yeah, so Lambda really loves Nvidia's products. I mean, they're the only server provider, the only chip provider that is available in every single major cloud platform, which is a huge platform advantage. And we stuck with the Nvidia sort of ecosystem for all the chips we've deployed. And we've got everything from V100s, A100s, H100s, H200s, B200s, GH200s, or GB200, B300s, and VR200s coming soon. And so we, you know, use everything in the ecosystem. Do you think that today or in the near future, we're going to be in a multi-silicon kind of world?

26:23Is there like room for different players beyond NVIDIA? Well, I mean, I think that we're already in a world where there's a huge amount of competition from massive, massive multi-trillion dollar companies and they're all trying to fight for the same thing, which is to be the best chip in the world for running and training neural networks, essentially. NVIDIA has built a great product that has gotten a lot of distribution and has a great platform of developers who love what they do. And you have to take into account not just the cost of the chip, right? The price of the chip is one aspect. But, you know, you have to take into account the entire software ecosystem and what's been developed.

27:03So one of the big people talk about what's NVIDIA's moat. One of the big moats they've got is just the CUDA-NN stack. It's not just CUDA. It's, you know, CUDA is, sure, that's like the water we all swim in, but like CUDA-NN has got so many, you know, matrix multiplication routine optimizations baked into it. What is CUDA-NN for everyone to understand? Okay, so CUDA-NN is the, it's CUDA deep neural network library. And it's basically, NVIDIA's, you can think of it like a highly tuned engine for matrix multiplication. And basically, if you were to just sort of naively implement the matrix multiplication algorithm, you would maybe get a certain level of floating points per second.

27:50But they've gone and tuned every single aspect of it and, you know, come in and do Winograd filtering or, you know, a bunch of different algorithms that you would apply to speed up matrix multiplication. and QDNN means that you don't have to go and do the optimization yourself. And so that's one aspect. The other one is Nickel NCCL, which is their networking optimization library where it will sense the topology and the connected nature of your network, your InfiniBand or your Ethernet network. and it will suggest an optimized sort of routine for doing, you know, reduce all and broadcast the different what are called open MPI primitives, which are used for that sharding that we were talking about for both training and for inference.

28:40And so that's like the kind of software stack that I think really is hard for a lot of the new entrants in the chip space to overcome. I think we're already, like I said, we're already in a world where there are multiple options for silicon. You know, the biggest labs in the world are using multiple different types of chips to do their inferencing and training on. What would be a plain English definition? We talked about the chips, but like the rest of the stack, the networking and the storage, just walk us through how it works. When you're running a cloud service, one of the things, you know, you'll train your model or you'll upload your trained model and you're ready to start doing large scale inferencing.

29:19Well, you're going to need a place to put your data, whether it's the data that you're using to train with or whether it's the data that's coming in and streaming in from your end customers. And so having high speed storage is like a really important part of it. And so Lambda offers the basically AI optimized file system service that is significantly faster than like your standard, let's say, cloud file system, which is maybe more of a traditional NFS type of thing. this is like a highly optimized parallel file system that's designed for high performance read and writes and mostly high performance reads that's like the kind of most of the workload and that's something you build in-house completely we have i mean it was in-house completely right it's like you have to ask the question like what is the definition of in-house completely right you know like we've never spun a pcb at the company we have not authored for example you know we use KVM Kimu for our virtualization, for example.

30:20And so we have both commodity off the shelf hardware that has software installed on top of it for some of our storage. We have some storage partners that we work with as well. But generally speaking, everything that we do on the cloud, I would generally say is something that is like we rolled it ourselves with the help of the broader ecosystem because again there's no such thing as rolling it yourself unless you're like you know mining um you know ultra pure silicon from somewhere you know and then coming up with your own uh asml uh you know it's like it's uh it's funny yeah that's the highly optimized uh storage what else the networking part and what other yeah so so i i was talking about this one click cluster product that we've got.

31:15And the way just for everybody to think about this is like, okay, well, look, you've got a bunch of GPUs. Let's say you've got a cluster of 10 ,000 GPUs. Well, I want to partition that cluster up. And so what it is, is it's a bunch of GPUs, some CPU servers as well, because you need to have an orchestration fleet as well. And then you've got some storage. And all of the CPU servers and the storage servers and the GPU servers are interconnected with the storage so they can quickly read and write from it. and so there's and that communication happens over what's called you know the in-band network and then there's the compute fabric which is where I was talking about where all the sort of weights and feature activations are being shared throughout that compute fabric and then there's an out-of-band monitoring network where you've got access to whether it's BMC or some of your DPUs and when you are trying to create a sub partition of a 10 ,000 GPU cluster, you need to simultaneously partition the in-band, the out-of-band and the compute fabric.

Read the full transcript

32:29Okay. So like that complex coordination between we've got a bunch of bare metal systems to, Hey, we've got a virtualized system that has, you know, what's called RDMA, you know, RDMA remote direct memory access that allows them to read and write quickly, not just from the disks, but from each other's memory, the GPU's sort of HBM memory, and allow them to do that sort of direct memory access, allowing it to go directly from a GPU to another GPU without getting copied to the CPU, for example. Having that all work is an immense, immense software undertaking. And going back to the original question, it's like, well, what are people not getting about neoclouds?

33:19Well, first of all, the answer is that most neoclouds don't have this kind of technology. Most neoclouds have not made the really it's like kind of high tens to hundreds of millions of dollars of software investment that you need to make to build a real cloud system that can partition a high performance computing environment like this. And so like that is, and then to have it all work with the storage. Anyways, I guess that sort of summarizes the steps that you need and kind of, you can think about all the different moving parts of a modern, like how does an AI data center work? People talk about AI data center, but really you have to kind of go down that one next level down, which is, because if you were By the way, if you were to ask an AI data center landlord, a traditional one, what's going on inside of the data center, they'd be like, well, look, we're real estate people.

34:15And we really outsource this to the GC, but the GC doesn't know, of course, anything that's going inside. It's their tenants who know. So this is what's actually happening inside of an AI data center. And then it serves the result. Also, going back to the community stuff. If people knew a lot more about like, well, this AI data center is actually just serving the chat GPT requests that I'm giving it. Right. Like sometimes they don't even realize that that's actually what an AI data center does. So you mentioned tenants. Do you rent them? Do you also own some, building some? And what does that fit in the overall strategy?

34:54Yeah. So, you know, initially we started off as being primarily a renter and we've actually started to get into the business of financing some of them, the construction of them ourselves, as well as, you know, we're going now into full vertical integration where we are identifying land, coming to the table with a basis of design, which is basically all the engineering diagrams to construct the data center. financing and constructing that data center, putting the servers in, and then associating that with a long-term offtake agreement with one of the major compute consumers in the world and financing it all.

35:37So we're getting to full vertical integration at Lambda. And it's been great because we've been able to kind of, again, bring that engineering mindset to this problem, which was historically mostly run by people in real estate. In your own data centers, are you the sole tenant or is part of the idea that you can also rent some to others? In a lot of our data centers, we are the sole tenant. In terms of the data centers that we're planning on constructing, we don't yet have any plans to lease that space to others. So we're not trying to get into the leasing data center business. Maybe that's something that you can imagine down the road.

36:17I wouldn't write it out completely. But for now, we have to focus on providing Lambda with the compute that we need to service the market. How international are you, by the way? I'd say that we're very much focused on North America. And so we have data centers in Canada, United States, and Mexico. We're very much, like I say, primarily focused on North America, but really within the United States, obviously. And we haven't had this desire internally to try to go and expand into Europe or too far into Asia. We've done some partnerships with some of our great investors like SK Telecom. And we have a data center that we've operated in Korea and Seoul.

37:06And so we have some experience with international. But right now, we're just like, look, let's focus on the U.S. market. it's where the opportunity is. Do you need for performance reasons to be close to the customer the way you need to have regions in cloud? You know, it's super interesting. I get this question a lot and people, they're like, well, does latency matter? Does, so I'll tell you what matters and what doesn't matter. You can look at your own utilization of whether it's chat GPT or Claude or Grock or Gemini and you can see, hey, a lot of the things that I'm doing, I kind of shoot it off.

37:45I come back later and there's a research report for me. Maybe it's a long running agent workflow. In those cases, latency doesn't matter at all. The only thing that matters is your cost per token. That's all that matters. And so that's been a really interesting change. I think that, you know, the old school, traditional legacy cloud business was so latency focused because of some of the applications. But this new fleet of AI applications are far less latency sensitive. So that's one. But there is the caveat, which is this governance and data governance is becoming an important thing. And a lot of countries are wanting to have the AI compute that their citizens are using be run out of their own country so that they can, you know, at least have their own perception of control or whatever and the you know that is that is another that is an element to it but i'd say that from the latency there's no technical reasons let's talk about the financing stack so presumably it's a commission of equity and debt how does it all work yeah so um the way that it works is that you you know you can really fragment it into this two parts which is like financing your on-demand cloud versus financing an offtake agreement, which is like a longer term commitment.

39:08And on the on-demand cloud, you're kind of looking at Lambda's credit quality. On the offtake agreement, you're kind of looking at the credit quality of the end customer who's paying the bill. And so what you do is you just take your offtake agreement, you take this chunk of GPUs that you're deploying, you take a lease or the property and you kind of put it into a box and you can go to the private credit markets and you can come up with an asset-based loan. There's a variety of different methodologies for financing it. Most of it is just some sort of special purpose vehicle that's designed to finance this particular deployment with a very known and easy to underwrite, which is basically just a fancy way of saying the finance term for just assessing the risks and the downsides of a particular credit investment.

40:08And there's a vibrant private and there's a vibrant credit market for that. On the on-demand cloud side of things, it's not quite as mature as when there's, for example, an investment grade off-taker agreement. But it's becoming more and more mature. And in general, creditors and lenders are really starting to understand the value of an NVIDIA chip. Because, you know, you actually look at the chips that we deployed in 2023, H100s. We're now leasing those out at a higher rate now than we were originally in 2023. So these creditors are starting to look at these assets and say, wow, this is an asset that is very valuable and also easy for us to underwrite.

41:02And of course, while they are underwriting towards the actual cash flows that are coming out of that agreement, just as an asset class overall, people are realizing that this is a really great opportunity. And so creditors are starting to flock to these deals. You're renting an 800 at a higher rate because, why? Because the demand for compute is so rabid that people will take any or the technical depreciation of the product is slower than people thought? What drives that? Well, what's driving it, I mean, certainly it's the demand being high increases the price that you're able to get in the market.

41:41There's no question about that fundamental law. Again, going back to what people didn't understand about this market. There was people who were saying, oh, well, there's a five-year lifetime or three-year lifetime. I even heard some people say three-year lifetime for these GPUs. This is completely false. We have GPUs that we've commissioned, and we're one of the earliest neoclouds. In fact, we're probably the only neocloud that actually has GPUs in our fleet that are fully depreciated from an accounting perspective, which is most people are adapting around a six-year accounting depreciation schedule.

42:17But that's not the usable life. The usable life is longer than the accounting depreciation schedule. And what really matters is the economic usable life. And so what we're starting to see is that like the people who are the naysayers, oh, this is going to be you're going to throw these GPUs out in five years, are completely wrong. They're completely wrong and they've been wrong the entire time. Do you think there is going to be or do you already see happening some kind of financial market for compute units, you know, with trading and derivatives? Is that happening? I'm starting to see some people start to examine what a maybe vibrant spot market – first, you need to have a spot market for something before then you can establish a derivative like a future or other more exotic things.

43:07I'm starting to see that. But fundamentally, I think that the asset class is just starting to mature and creditors are starting to become very comfortable with investing in the credit side of buying NVIDIA GPUs and deploying them into data centers. And we don't need to get too fancy with it. That's kind of like part of my opinion is that I think that that market is starting to mature, that maybe an eventuality is having more complex securities that surround GPUs. But I think for right now, people are starting to realize that it's a great credit investment. And that's what's changed, I'd say, over the last year, is that people have started to really treat it like a more mature asset class.

43:54Maybe quickly just go back to the very origin, because I think you've been effectively in the AI world the whole time, but are coming from a very different angle with multiple pivots. What did you start with and when? Well, you know, with the complexity of the business, you can now see the complexity, the capital intensivity, just the sort of not fitting into a box. And you can see why we've oftentimes not had a lot of traditional venture investors in Lambda. And, you know, all of our investors have done exceptionally well, but they've kind of come from more often than not outside of traditional, let's say, mainline Silicon Valley VCs.

44:34And so just going back to the origin story, I started Lambda in 2012, and we were a facial recognition software company. So I was training convolutional neural networks to do face and image recognition. And we eventually hosted that on API. I was training those ConfNets on a 4X NVIDIA GTX 580 workstation that I had bought from a friend who had built it, actually. And this was, you know, really pretty avant-garde stuff at the time. Most people didn't really believe in what was called the field called deep learning at the time. And that was inspired by the ImageNet 2012 moment or that was even before that?

45:24The ImageNet moment, you know, I pulled the CUDA ConvNet repo off of Google Code. That's how you know how old Lambda is, is that Google Code was still around. And I pulled the CUDA ConvNet code base and was like playing around with it. I got very lucky that the AlexNet paper had been published the same year that Lambda was founded. It's not a coincidence at all. It's not a coincidence at all. We launched this face recognition API, got a couple thousand users, but it wasn't really generating a ton of cash. And sort of as part of that complex story of startups, you know, in parallel, I was sort of I found these guys who had just graduated from their PhD programs.

46:08programs, the gentleman, Zach and Nico. And they had said, hey, we're going to start a company. I said, hey, let me help you guys out. I'm going to work with you for a year. I'm going to learn a little bit more about neural networks. And I helped them out on this company. I helped them get a company called Perceptio started. And I was the first employee there while I was running Lambda. And we were running these confinets locally on the iPhone. And again, this is 2013. So we were using the GPU image library and just straight open GLES shaders, like the shaders that are used for rendering. We were using those to run the confinets on the iPhone.

46:53And eventually, I kind of left to go continue to work on Lambda full time. and probably about a year or so later, they got acquired by Apple. And so if you know the feature on your iPhone where you swipe up on an image and you can recognize faces and search through your library, that's maybe some of the stuff that eventually got integrated into iOS through that acquisition. And then Lambda, we continued on. We had a variety of different products, everything from Lambda Hat, which was a baseball cap that took a picture every 10 seconds with a camera embedded in the tip of the brim for gathering data sets for image and face recognition.

47:30Which is fascinating because fast forward to today and that's a whole segment, right? Like capturing everyday life to train the AI. It goes to show you have to, you know, it's one, it's important to be able to see the future. It's also important to get your timing right as well. Right. And now it all worked out, right? Despite maybe that Lambda hat product not being great, but it taught me a lot about how to build hardware. where I lived in Shenzhen for a little bit, working on the PCB and spinning the PCB and designing the actual hardware product. And it taught me how to make consumer electronics.

48:08And that was actually a huge, huge skill because it totally opened my mind to new ways of doing business that aren't just making apps, right? And eventually, we had this product called DreamScope, which became really popular in 2015 and 16. And it was basically using the Google Deep Dream methodology of using a ConvNet to generate images. It's like an early version of MidJourney or whatever. And Deep Dream and the Leon Gatte's style transfer algorithm allowed you to turn a photo into a painting, basically. And we got like a million users on that process, tens of millions of images, maybe 15 million images or something like this.

48:54And that caused us to have a huge AWS bill. It was like$40 ,000 a month or something. And so to replace that, we'd ended up building a little cluster out of workstations. And then there was a$60 ,000 CapEx that we were terrified to make, by the way. We were so, so scared that doing this CapEx was going to put us out of business. We made it out of workstations because we thought, oh, well, worst case scenario, we can just sell them. And so lo and behold, we did end up, you know, turned it online and it brought the bill down to zero. So it paid itself back in a month and a half. And we thought, oh, this is like, we're saving more money than we're making.

49:32Maybe we should be in the business of providing compute to other AI researchers. And thus, we started selling workstations and servers and started developing our cloud platform, maybe did$3 million of revenue in 2017, that first year selling workstations, then 10 million in 2018, then 30 million in 2019. we grew the hardware business over the next couple of years to probably about$200 million run rate. And then the cloud business, we really started in 2019. And we started development before then, but we started really marketing it. And it kind of was slow to grow, to be honest, because not a lot of people in 2018 and 19 and 2020 wanted a bunch of AI compute.

50:16There was a pretty niche market for it. But eventually, you know, our cloud business continued to grow. And now it's at, you know, a little bit under a billion dollar revenue run rate. We've fully exited the hardware business. And so, yeah, Lambda's got a absolutely wild founding story to summarize. Are some of the people that were there at the beginning still around? I think you started the company with your brother. Is that right? And your brother is still at the company? Yeah. And so in terms of like the early people, basically, it's not, I mean, not even basically of the four people who are making DreamScope, me, Michael Balaban, my co-founder and fraternal twin brother, Chuang Lee, who's our chief scientific officer, and then Steve Clarkson, who's an engineering leader at the company and, you know, has a bunch of folks reporting into him.

51:07Now, you know, they're all still at the company. the next hire one of those gentlemen named Mitesh Agrawal who was one of the next hires in that team he was with the company for maybe eight years or something like this

51:25and then he eventually left and joined another former Lambda team member Thomas Summers to start Positron which is an accelerator company and they're like now valued at over a billion dollars and so not only has like the original team stuck around but we've already started to kind of see what like a Lambda alumni a Lambda mafia network looks like in in the world Lambda lab member alumni How did you keep the band together during the difficult times? Just when you're running a startup company that's this capital intensive, working capital intensive as well. Like it's just you get a lot of shocks to the system as you're growing.

52:16And then COVID. What COVID, I mean, in April of COVID, software companies were feeling great because they could ship software and there was so much more demand. Hardware companies, the docks were closed. You couldn't ship revenue. in April and March. And so, I mean, I remember all these things really distinctly. I think I remember just getting in front of the team like, hey, look, it's really tough right now. And there's certainly a feeling that we're not sure if we're going to make it through this, but the only thing to do is just to suck it up and enjoy the pain, run through it and come up with the solutions to the problems that you're presented with, all in the service of delighting customers.

53:04Because fundamentally, I mean, the big thing is just aligning people towards the only reason we're all here is to build something that people want and they love so much that they tell their friends about it and they give you money. And then everything else, it just follows from that customer experience of delighting customers with what you do. When we do onboarding, for example, I used to do this thing in what's called Lambda 101. And we would show a picture of like a Linux penguin. And he was like on a Lambda workstation. And he was reading the GPT-2 paper and training at a loss curve, which is like what you see and you look at if you're doing machine learning research.

53:45I was like, look, just put yourselves in the shoes of the penguin who's training using this workstation or cloud service to train a neural network and just think about what's going to delight them. You know, whether it's, you know, people on our shipping team who said, hey, let's put some T-shirts inside of the boxes. And so every workstation came with a lambda T-shirt, you know, or members of the data center operations team said, hey, you know what? we should do a white rack because that'll kind of like set us apart and make everything look good. And we'll be really proud to showcase that. And, you know, those are the types of things that as you kind of imbue your company with the kind of delight the customer first mentality that I think helps you get through the hard times.

54:29Recent evolution in that journey is that you just brought on a new CEO and my fellow French countryman, Michel Combe, to run the business. Walk us through the thinking and what led you to make the decision and how that equips the company for the next chapter. It's a huge honor as a founder to get to the point where the company can afford to bring on amazing talent like Michelle in that seat. Because if you think about it, I'd say most companies, it's not uncommon for somebody to say, Hey, look, a lot of people sometimes maybe there's a component of ego involved where they have to be the founder, CEO.

55:14I've never really personally had that. I care about the technology, as you can tell. I care about building a great generational company. And I think there's so many different seats to do that from. And so getting to the point of maturity where we can afford to bring on a CEO like Michel, who has experience, obviously, previously SoftBank International CEO, Sprint CEO. Alcatel. Alcatel. He's on the board of some really amazing companies, including McLaren, which is a fun one. I always did the fundraising and capital formation and day-to-day business management as a necessity and not because that's what I really love doing, for example.

56:02And I think there's plenty of founder CEOs who absolutely love every aspect of their CEO job. I think that privately, and it'll be very hard for you to get this out of any founder CEO oftentimes, but like secretly when I talk with founder CEOs, I'm always like, yeah, so like, how much do you hate? I find it shocking that people don't find speaking to VCs all day exciting, but I will take your word for it. It's been like an amazing experience for me that I, to be able to form a team around the company and to just see everybody flourishing in the things that they love to focus on. So for example, now that I'm the CTO, one of the main things I'm focused on is what does rapid data center deployment look like at the company?

56:53And, you know, kind of working to say like, hey, I want Lambda to be this sort of vertically integrated high velocity powerhouse that. So when you look at the world, you say, all right, there's two people in the world that can and two companies in the world that can do high velocity deployments, SpaceX AI and Lambda, where we're just extremely focused on how do you cut every little piece out of the process to stand up compute faster. and that's like something I've just been diving into and really enjoying with my new time as a CTO. What was XAI's record like when they launched? I think it was like 200 and something days.

57:39Yes, and you think that can be matched or exceeded at a repeatable pace? I think it can be matched or beat, yeah. And that's process mostly? I think it's everything from the site selection process, the set of constraints that you use in a site selection process, the MEP pipeline, the way that you construct the data center. How do you make it so that the end customer will consume that compute? And how do you cut out a lot of stuff out of the process? Because oftentimes, you know, the people who've been designing these data centers have really kind of been real estate people, as I've mentioned, who've been kind of grabbed by the scruff of their neck by a hyperscaler.

58:22They're like, go and build this design. Here, go. Go get a GC runoff. And they don't know anything about what goes inside of it. And so, and the hyperscalers, on the other hand, have been really building towards traditional cloud services. I mean, if you look at a modern region in any of the clouds, they have hundreds of services. I mean, everything from satellite base stations to tape storage to spinning disk to face recognition APIs. I mean, these are all the services and each of those services requires a different SKU and has different parameters about what you're kind of servicing. and in fact, you know, you might have somebody who's trying to run an ATM backend on one of these things.

59:08That's a pretty different design space and design constraint than an AI data center that could maybe, you know, have a lower availability and uptime, right? And so that's kind of where I think Lambda is able to build a lot of really unique value and, you know, through this kind of targeted AI first approach. You had a quote where you said that AI won't write software, it will become the software. What do you mean by that? So that's in my sort of like idea around what I call neural software and or, you know, a neural computer, neural operating systems. And the best way to kind of get this experience is to go to your chat GPT or your cloud and say, hey, just, you know, render for me an ASCII art desktop interface.

1:00:03Okay. So you're working in purely in the domain of text. And I want you to just pretend to be an operating system for me. I'm going to say, click on this, you know, open up this and I want you to just behave like a computer. So give it that prompt. Okay. And what you're going to see, I think, is that you're going to see that sort of future of the large language model becoming the software and not generating the software. And this results in an extremely sort of squishy and flexible way of interfacing with a computer where it's not possible to have a bug, only a misunderstanding about the prompt and what you've asked for.

1:00:48And I think that for a lot of the pieces of software on your computer, you might see that taking over where, you know, you can get the glimpse of the future with this ASCII art, and then eventually it'll also have a multimodal network that's generating every pixel on your screen, as well as every audio waveform that comes out of your speakers. The advantage to this is that you can really sort of dream up software. The only the part that is being experienced by you and is actually implemented, if that makes sense. It could have whatever feature it is if you ask it. And that's like a really powerful way to interact with a computer, I think.

1:01:31So it's not like you give simple instructions to the LLM, suddenly the LLM is the software. I guess make the analogy, Like, you know, vibe coding takes in a prompt and then outputs human readable, writable, compilable code that runs on normal human software programming language substrate, right? You know, it outputs C code, which gets put through a compiler, or it outputs Python code, which gets put through a Python interpreter. That software is static. Once it's been generated, it can't change, right? you can vibe code it again and maybe you know vibe code on the fly there's like a couple different stages of the gradient between traditional human written software and then you go to like maybe vibe coded software then you go to just in time vibe coded software where like it's a live creation of the software application but still software but then you go to the next step which is just you're interacting with the LLM and it is emulating kind of how software might behave.

1:02:33And that's the difference between Vibe coding and a neural operating system or neural software. Neural software, there is no code that's running. It's just modifications of the feature activation space and the context in the mind of the neural network. And how far do you think we are from that? Is that something? I mean, we have prototypes of it today. So we have prototypes of it today. And when you say you, is that Lambda? Or is that... Yeah, Lambda's developed a prototype. There are multiple other companies that develop prototypes of this. There's academic research that has outlined what this might look like.

1:03:16And how far are we from mass adoption? I would say that generally speaking, when I'm early on something, I tend to be about a decade to a decade and a half early. So I would say that between a decade and 15 years, we will see mass adoption beginning or otherwise happening for neural software. I mean, you already have it. So here's another example, by the way, you already have, So you can think of a Tesla self-driving car or any type of end-to-end neural network and then large model that's doing autonomy as a form of neural software. And people understand that aspect, which is it's seeing video, it's making decisions about what to output.

1:04:09Now, the user experience is the driving experience. that said that is an example of neural software i would i would argue and so we already see that today now the question is when is your everyone's computer is going to adopt that i'd say a decade do agents change anything from your perspective as a compute provider and if so in what way to understand what needs to change on the computer we understand need to understand what's changing with the user so when you're doing vibe coding with agents one of the things you'll notice is that your wall clock time in the world is mostly spent on running tests, gathering data, searching through code base.

1:04:54A lot of the time is spent not just inferencing a neural network, but it's actually spent doing other things. And it's actually very much similar to how software engineers spend some of their time, right? You know, the old XKCD cartoon of compiling where they're sword fighting on the office chairs. And someone says, what are you guys doing? Compiling. And so now there's a bunch of time spent compiling. There's a bunch of time spent running tests because the part of the way that the agent 24-7 loops really work well is when you are constantly banging against a nice suite of automated tests to make that the code you're writing is good.

1:05:38And so, well, what does that mean? It means that every single cloud service needs to start doing a lot more traditional CPU workloads. They need to focus on a great environment, a secure environment to host your Claude code instance on. And then you need think about security from the perspective of you need to think about how this massive influx of new applications are going to be secured how do you use the agents internally well i mean a lot of the engineers at lambda are already you know doing a fully agent driven workflow i mean if you just go to clod code and say hey use advanced workflows or you know spin up agents, you can do that.

1:06:33So that's like step one. I've demoed internally and some folks have adopted what I kind of call self-assembling software. And so self-assembling software is this idea where you kind of tie in to a 24-7 running agent fleet product requirements and constant user feedback that's coming off of the system. So you have a very clear and tight loop to go from submitting, hey, this is a bug or this is a feature request. And there's a fleet of agents who are implementing that live for you. And that sort of cycle I call self-assembling software is because you kind of say, hey, this is what the software is for.

1:07:18But most of the development for it is going to happen after the software is launched and the users start to interact with it and customize it for themselves collectively. And I think that that is kind of maybe the future paradigm of where a lot of the agent-driven development is going to go towards. The other side of that, eventually, once the models get smarter, I think that they're quite not quite there yet. But, you know, tying that back into, hey, I need help. And I'm not talking about the human, I'm saying the agent's going, hey, I need a human to help me. I need you to plug in a thousand GPUs for me, or I need you to give me an API key to a particular service.

1:08:01I need you to go sign up for something for me. Can you go please negotiate this? And I think that that's actually how you're going to start to see it happen, which is product user feedback gets implemented by the agents. The agents then also ask the people at the company to go and do things for them, all in the service of delighting customers and making money. You've talked about gigawatt scale factories. Is that what you were describing earlier around like setting up a beginning super good at creating data centers very quickly, but it's also making them bigger? What is that concept? It's an AI factory, which is a basically land data center servers inside that is generating tokens.

1:08:42And a gigawatt scale means that it's consuming a thousand megawatts or a billion watts which um is a lot of power sort of like maybe you can think of it for context new york city is something like five gigawatts you also talked about one person one gpu is that uh your vision for the future unpack that for us so um you know before people really believed in the ai thesis i you know when i was pitching our series b and c I would kind of talk a lot about the similarities between, let's say, the computer industry and the AI industry. I really felt like AI was forming a set of generational companies.

1:09:26And there was going to be a set of generational companies that got minted with the changes that were coming with AI. And this is like in 2020, 2021. And if you read about the history of Apple, for example, in the early days, the motto and the sort of the credo at Apple was one person, one computer, one person, one computer. And, you know, there's a sense of humility that's embedded in this one person, one GPU, which is the one person, one computer. You think about how visionary Steve Jobs was. That was, you know, Apple was, you know, what founded in 1976 or something like this. The Macintosh came out in 1985 or 1984, excuse me.

1:10:081984, you know, whatever. Eight years or so after founding. Is that one person, one computer yet? No, not even close. All right. So 1984 to 1994. All right. Well, is it one person, one computer? Well, we're just starting to have the internet boom. So we're not quite there yet. 2004, we finally have broadband internet access. And maybe for the first time in the United States, there's not quite one person, one computer, but there's certainly like one person, one family, one computer, you know, or, you know, something like this. It's like getting close to it. You don't have until 2014. So 74, 84, 94, 2004, 2014, 40 years after one person, one computer, do you have probably truly one person, one computer?

1:10:57And you get actually beyond one person, one computer because people have laptops and cell phones and I would consider a cell phone a computer. So and then and then finally, you didn't have e-commerce penetration until 2024. Fifty years after the founding or so of Apple Computer, when e-commerce starts to actually penetrate because of COVID. I think that the reason I really wanted to choose that one person, one GPU is because, one, I believe that in the future, everybody in the United States will need the computational power of one GPU or more to just do their daily work, enjoy life, whether it's getting access, whether it's getting entertained, whether it's being productive, whether it's being creative.

1:11:45And I also recognize that it took Steve Jobs and Apple, one of the best companies in the history of capitalism, half a century to accomplish their goal. And so I think this is not it's not just like an overnight. Let's, you know, quickly get to one person, one GPU. So that's that's that's what that means to me. To close, are you ready for a couple of quick hot takes? Sure. What is one idea in AI that is overhyped? I think a lot of the sort of agentic workflows for things that are not software engineering, I think, tend to be overhyped. And I'll tell you that the reason for that is because one of the ways that you get an agentic workflow working really well is that it needs to have very concrete feedback mechanisms, which are done brilliantly through automated testing.

1:12:32It's not at all done brilliantly for like going by a site. there's no traction to give a model to go and iterate over a long period of time on. So I think agentic workflows for things that aren't readily verifiable. Now, I wouldn't say as far as everything that's not software engineering, because there's plenty of readily verifiable fields. CAD, computer-aided manufacturing, finite element analysis, computation fluid dynamics. There's a bunch of fields where you can really do a great agentic workflow. I mean, and simulate it and then go and iterate. It's not the case for, hey, Claude, make me a billion dollars.

1:13:12Make no mistakes, you know. Sadly, or maybe not. Okay, fascinating. Inflation would be, it wouldn't be inflation. It would actually be just value creation in the economy. Deflationary even. What is one idea in AI that is underrated? Yeah, I really think that the neural OS thing and, you know, also some of the aspects of self-assembling software. I still think people... You know, the funny thing is I'll give the same answer. Agent-based workflows for software development. I think that most people don't understand. They literally don't understand because they've never tried it. They've never gone to Claude.

1:13:48Go to Claude, go say maximum effort, use the latest model, and then go and build whatever you wanted to build and say, you know, spin up 10 agents to go and do it. I think a lot of people still haven't done it yet. Well, Stephen, it's been wonderful. Thank you so much for spending time with us. Matt, thank you so much for having me. Appreciate it. Hi, it's Matt Turk again. Thanks for listening to this episode of the Mad Podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from.

1:14:22This really helps us build a podcast and get great guests. Thanks and see you on the next episode.

From the publisher

Many people said GPU compute would become a commodity. The opposite happened — and a new category of "neoclouds" is now racing to build the physical backbone of the AI boom. Stephen Balaban, co-founder and CTO of Lambda, explains why the conventional wisdom was exactly wrong, why we're still massively underbuilding compute, and what it actually takes to stand up a gigawatt-scale AI factory: land, power, cooling, networking, and a financing stack most people have never heard of. We go deep on the physics of how energy becomes tokens, NVIDIA's real moat, why a 2023 GPU can lease for more today than the day it shipped, and Stephen's provocative vision of "neural software." Plus the wild Lambda origin story — from a facial recognition startup to a camera in a baseball cap to a near-billion-dollar cloud business. This is the state of AI compute in 2026, from inside one of the companies building it.


(00:00) — Cold open

(01:21) — Why GPU compute was never a commodity

(02:45) — The H100 price index and what it gets wrong

(04:02) — The real moat: technology or financing?

(05:57) — Winner-take-all, or room for many neoclouds?

(06:48) — Are we overbuilding or underbuilding AI compute?

(09:26) — What if AI gets 10x more compute-efficient?

(10:44) — The real bottleneck: land, power, and shell

(11:38) — The backlash against data centers — and the misinformation

(15:00) — Opening the hood: from photons to tokens

(17:11) — Extracting more value from the same chip

(19:26) — Frontier inference and distributed training, explained

(23:26) — What actually drives compute cost

(25:21) — Lambda's chip stack and the NVIDIA relationship

(26:17) — A multi-silicon world? CUDA, CUDNN, and NVIDIA's real moat

(28:59) — Networking, storage, and the one-click cluster

(34:46) — Renting vs. owning, and full vertical integration

(36:24) — How global is Lambda? Does location still matter?

(38:44) — The financing stack: off-take agreements, SPVs, and credit

(41:16) — Why a 2023 GPU leases for more today

(42:36) — A futures market for compute?

(43:54) — Origin story: facial recognition, Perceptio, and Apple

(47:03) — The Lambda hat and Dream Scope

(48:59) — The $60K bet that became a cloud business

(52:00) — Holding the team together through the hard times

(54:30) — Bringing on a new CEO; Stephen as CTO

(57:33) — Matching xAI on high-velocity deployment

(59:29) — "AI won't write software — it will become the software"

(01:01:30) — Neural software vs. vibe coding

(01:04:25) — Do agents change the compute layer?

(01:06:14) — Self-assembling software inside Lambda

(01:08:18) — Gigawatt-scale AI factories

(01:08:57) — One person, one GPU

(01:12:04) — Hot takes: overrated and underrated in AI

More from The MAD Podcast with Matt Turck

All 44 episodes
The GPU Myth: State of AI Compute 2026The MAD Podcast with Matt Turck · 1 h 14 min
Listen in VO