In short
Base10 CEO Tuhin Srivastava explains why AI inference is the “last market,” how workloads are shifting toward custom/post-trained models, and how Base10’s “inference cloud” tackles capacity constraints across many clouds. He argues the application layer will persist because unique user signals (e.g., clinician edits) enable differentiated post-training. He also discusses multi-chip futures, runtime bottlenecks, and the economics of compute scarcity.
Guest background
Tuhin Srivastava is founder and CEO of Base10, an AI inference cloud company. He previously led infrastructure/product efforts and recently acquired a post-training-focused research team (Paz).
Key claims
Base10 grew ~30x in a year; expects >$1B revenue this year. ~95% of tokens served are on dedicated custom inference; customers rarely run “vanilla” open weights. Compute supply is extremely constrained (mid-90s utilization) and hyperscaler inference has operational/runtime edge-case limits.
Notable examples
Abridge (ambient scribe integrated into US hospital clinician workflows); Stripe as an analogy for building for frontier customers; DeepSeek as an example of open model capability and cost; BrainTrust evals and “sandboxes” for coding agents.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOThe AI Inference Landscape
0:45 to 2:52
Discussion on the growth of AI inference and its significance in the market.
“I think what's happened over the last, honestly, 24 months, but it just kind of keeps getting bigger and bigger, is that I think everyone is realizing that you can put AI everywhere.”
The Role of Application Layers
2:52 to 6:20
Exploration of the importance of independent application layers in AI.
“Sorry, Abridge is an ambient scribe that is used by physicians in almost all hospitals in the US.”
Enterprise Adoption vs. Application Companies
6:20 to 9:45
Insight into the current market split between application companies and enterprises adopting AI.
“I think firstly, you just learn a lot by building with the company's greatest scale, doing the most interesting things.”
Open Source Models and Their Evolution
9:45 to 12:10
Examination of the evolution of open source models and their significance.
“Yeah, look, I think these models, firstly, are fantastic.”
Geopolitical Considerations in AI
12:10 to 14:00
Discussion on the geopolitical implications of using Chinese models in AI.
“You know, like, and like, you can argue whether it's at the absolute frontier or not, but like, let's, let's go back three months and it's there.”
Post-Training Customization and Acquisition Rationale
14:00 to 15:10
Learn about the importance of post-training customization and the rationale behind a recent acquisition.
“the customer is making some modifications to the model with their own data, specialized for the use case.”
The Relationship Between Inference and Post-Training
15:10 to 17:20
Explore how inference and post-training are interconnected and their implications for model optimization.
“So they were post-training models and running them on base 10.”
Navigating the Supply Crunch in AI Infrastructure
17:20 to 19:30
Understand the current supply crunch in AI and how companies manage capacity across different clouds.
“So between that and post-training, these are very difficult to gather capabilities.”
Challenges with Cloud Suppliers and Operational Capacity
19:30 to 21:30
Discuss the challenges companies face with cloud suppliers and managing operational capacity amid the crunch.
“things that we think are going to be very important for very mission critical use cases.”
Strategic Factors in Dominating the Inference Market
21:30 to 24:30
Learn about the key factors that contribute to becoming a dominant player in the inference market.
“crunch we're supplier and operationally crunched onto people who can who can run these data centers as well.”
Show all 20 chapters
The Future of Chip Technology in Inference
24:30 to 27:30
Examine the future landscape of chip technology and its implications for inference and AI workloads.
“Is it, as you mentioned, cost of capital?”
Adapting to Emerging Workloads in AI
27:30 to 28:01
Discuss how companies need to adapt their investments based on emerging AI workloads and trends.
“And like, it just like, given the scale that they operate at, given the scale that they operate at, it's hard to see.”
Investing in AI Runtime and Workloads
28:01 to 29:50
Explore the crucial factors influencing AI runtime efficiency and workload management.
“a bunch of the other chip providers have done it's actually hard for that ecosystem to form.”
Discovering Edge Cases at Scale
29:51 to 31:38
Learn about the unexpected challenges faced when scaling AI services.
“We will partner with or on the sandbox side build the best sandboxes experience that will exist.”
Concerns of Capacity and Compute
31:39 to 33:37
Discuss the critical issues surrounding capacity and the demand for compute power.
“And like, you know, and I'll give you a few examples here.”
Recruiting and Building Effective Teams
33:38 to 36:46
Uncover strategies for recruiting top talent and creating a strong team culture.
“to get the amount of value that we want to get out of our limbs in the next five to 10 years.”
Operational Culture in AI Infrastructure
36:47 to 38:37
Examine how operational culture impacts AI infrastructure and team dynamics.
“And so, you know, I think that is, you just have to get used to it.”
The Impact of Lowering Inference Costs
38:38 to 40:46
Analyze the effects of decreased costs on AI model usage and consumer demand.
“Like the personal or business ROI of it, the demand for it goes up, not down.”
Envisioning the Future of AI and Intelligence
40:47 to 42:00
Explore predictions for how AI will evolve and affect consumers and developers alike.
“We have this massive shift where we're moving from software and seats and digitization into actual intelligence, selling units of cognition, selling agentic workflows.”
The Future of Design and Workflow in AI
42:00 to 42:34
Explore the shifting landscape of design tools and workflows in AI.
“it's the extinction moment for a bunch of folks, which is like, everything needs...”
Transcript
Automatic transcript. May contain errors.0:05Hi listeners. Today, Elad and I are here with Tuhin Srivastava, the founder and CEO of Base10, the AI Inference Cloud. We're here to talk about capacity constraints for AI compute, why inference is the last market, how the workload is changing, the open source and perhaps multi-chip future, and what 30x scale in a year looks like. Tuhin, welcome back. Hi. Good to see you. Thanks for having me. All right. You are in one of the craziest markets, AI inference. It's very important. There's a lot going on. You guys have grown 30x over the last year. And I think I can say you're expecting to do more than a billion dollars in revenue this year.
0:47What's going on? Tell us about scale. Yeah. No, it's been nuts. I think what's happened over the last, honestly, 24 months, but it just kind of keeps getting bigger and bigger, is that I think everyone is realizing that you can put AI everywhere. You have all these great options available from closed source to open source models. The open source models have crossed some sort of chasm in terms of their baseline capability. And then I think RL techniques and post-training for specialized models has become mainstream enough. and there's enough examples of it working, the customers realizing they can kind of own their inference more and more.
1:34And what that's meant for us is more the long tail models coming true, customers in-house seeing a lot of that intelligence themselves. And as the application layer just gets bigger and bigger and bigger, and that's growing, we are just someone index on that and we've been around to be able to collect the demand. There's an existential question in here that I think everybody is continually asking of, does the independent application layer get to exist at all versus the labs? Like, how do you, you have to believe this. Why do you believe it? Yeah, look, I think it'd be, it'd be a sad thing if it didn't exist in general.
2:11And I think that's like my, but you know, sadness is fine. I said all the time. Sadness is fine. But that's not the reason why I think the application layer will exist. I think the application layer will exist for a number of reasons. One is because, you know, I think this idea that what is valuable to a company is, you know, the user signal that they can gather, that only they can gather. And to the extent that that is encoded in a model, I think a lot of their business will be at risk. But to the extent that it is encoded in workflows, that is where they will be able to develop mode. So a good example of that is, say, a company like Abridge, where the clinicians edit off the notes and what they do with those notes after the fact, and the thing that happens inside the EMR three steps down, that becomes a workflow that only...
3:11Can you explain what Abridge does? Sorry, Abridge is an ambient scribe that is used by physicians in almost all hospitals in the US. I think a large investor, great company, great team, great product. And they've basically got this very, very deep integration into hospitals, into clinician workflows. and my argument would be here is that actually it's very, very hard for a frontier model company to go to EWOA because they just don't have access to that user signal. And what will happen over time is the folks who have access to that user signal can start to post-train models on that reward signal and start to get long horizon agentic models running that.
4:01And I think to the extent that that is possible and that signal is differentiated and unique and is somewhat rare to get access to, there will be an application layer. And I think, you know, support companies is another example of that where, you know, a support task isn't one-shotted. Usually at a company like Base 10, when a ticket comes in, there's like, what, like one, two, 10, 20 actions that get taken. And that is where, you know, someone can develop a specialized model. So there's almost two versions of this then. There's new companies like Abridge or Decagon or some of these other things that you mentioned that are doing these new types of applications that are using AI and they sell it to customers.
4:44The other is enterprises building things in-house or building their own models. What proportion of the market today do you think is these new application companies versus enterprises just adopting AI? And how do you think that looks in a couple of years. Yeah, I think that's, I think you asked me the same question two years ago. I had to be repetitive. It is crazy. At least I'm consistent. The answer is just that it's crazy that the answer is still, I think, I think if you look by inference count, it'd be 99 % the full. Yeah. And that kind of represents the scope of the opportunity here is that the majority of the market hasn't come online and added AI into the world.
5:26Yeah, most of enterprise adoption is well ahead of us. And I think that's one of the very exciting things about AI. Yeah. There's just so much still to come and people are underestimating that, I think. 100%. But what's cool is that we're seeing the transition happen, right? Before it was like, hey, are they using AI tools? I don't think that was immediately obvious two years ago. I think that's obvious now. But yes, they are. Are they using closed source model APIs? I think they're starting to get there. And then once you do that and then you kind of see what is possible, then comes the whole custom model adoption.
5:54I think that is all that is ahead of us. today. So if the majority of your customer base today is, as you described, the former application companies, AI natives, the fast-growing, I mean, some of them are at considerable scale now, like the abridged cursor, open evidences of the world, what do they teach you? What does that push the company to do? How do you think about serving them versus evolving for the enterprise? Yeah. I think firstly, you just learn a lot by building with the company's greatest scale, doing the most interesting things. We think of it two ways. I think there's the most obvious way, which is just build for the highest scale.
6:41the customers that will push you the most from technologically and everything kind of will fall into play. I think the Stripe evolution as a company showed that was like Stripe now like serves so many enterprises, but 12 years ago, that wasn't the case, but they just built for the frontier and kind of went with them. And the second way we think about this is to just think about building for companies that are serving enterprises. So yes, we don't serve the enterprises, but our customers serve enterprises. Abridge serves every device, OpenEvnance, Decagon, all these, Ryter, Gamma, all these companies serve enterprises en masse.
7:19And what we actually get is like a translation of the requirements from them, which is like, you know, they're like, hey, we need this data retention. We need this web models need to be deployed. This is the types of GPUs or the latencies they're okay with. This is the model requirements from like a transparency perspective that they care about. And so I think that is actually the more nuanced answer is that if you listen to what their needs are, we actually get a full translation of what the enterprise was required. I would say that by serving companies like Abridge and Open Evidence, we're probably pretty well suited to go serve the healthcare system and latent health given that they are selling to them.
7:55How much of a shift are you seeing in terms of the types of open source models that are being used? And so I think we've seen an evolution where two, three years ago, I think the main thing was kind of Mistral and then a few other things. And then Meta kind of came along with Llama and then it kind of really shifted in terms of the misperformance models or of Chinese origin and different ways. Do you see that sort of mix reflected in terms of what's being used by our customers? Yeah, I think customers, at least the customers we are serving, are very, and these are like the fastest growing AI companies in the world that are very forward thinking.
8:25They want to use the best model. And they are optimizing. I think there is a subset of tasks, which I think is small today, where people really start to start with cost. But everyone comes from capability first, because that's really where the economic growth is being unlocked, where the value is being delivered. And then they optimize. And I think that's like actually been, you know, and so with that in mind, you know, you name, like you name it, everything from GPT OSS all the way to Moonshot, to DeepSeaks, to Canopy, Orpheus, which is like really good text-to-speech models. Customers generally want to use whatever's at the frontier.
9:11And I think the difference has just been, I think we have a lot more visibility into how to run these and how to run these really well. And secondly, that they're good now. There have been a number of different concerns raised about the use of Chinese models, in particular security, or is there something embedded in the models or Trojan horses or other things. A, do you think there's any real concern there? And B, people often talk about how there should be U.S. counterweights to this. from a geopolitical perspective, do you think that's something that's legitimate or something we should be worried about?
9:42Or how do you think about the sort of origins of these models versus their uses? Yeah, look, I think these models, firstly, are fantastic. They're amazing. We work with these teams. They're truly awesome. I'd say, look, I don't, it is hard for me. It's hard for me to see, and I could be wrong, but if I network bound these models that they're not magically going to be able to cross those network boundaries. And so data is there. And I've never seen any real evidence, except from some very early models that I think people picked up on very quickly that there is some agenda or bias built in this.
10:25I do think that to some extent is, I think there is importance to the US that we develop our own models. I think that that would be a massive loss if that there are five companies, you know, five different labs in China that are creating open source models. And we're struggling to get one set up. So it's necessary. I also think it's inevitable. And, you know, like the DeepSeek moment a year ago, I remember someone saying to me, and I thought it was like very well said, which is like, and the world's changed a lot. But they said, hey, you know, we should kind of just forget. that this is a Chinese model, we should just act like this came from meta and build with that in mind.
11:13It's like, you know, I think you're kind of missing the forest from the trees. Like there's two scenarios, right? Either America does not ever come up with good open source models. I think there's probably a fundamental problem there or we will get there and we need to be ready for that world. Yeah, that makes sense. It's interesting because, you know, like you, I think it's very important for the US to have a strong open source footprint here. At least for now, it looks like effectively the Chinese government is subsidizing at least a large subset of these models. And that subsidy or surplus is effectively just being passed on to US enterprises who are adopting these models.
11:46In other words, it's a way for the Chinese government to effectively subsidize US enterprise in an indirect manner. And I think that's a little bit lost right now. But it's always interesting to weigh that against some of the other concerns that are raised. I appreciate your comments on this. I think the concern also there just becomes It's like, what happened if we aren't able to, like, if it is fun, like, I think if you think of the economics here, which is DeepSeek by most, DeepSeek is a very good model. You know, like, and like, you can argue whether it's at the absolute frontier or not, but like, let's, let's go back three months and it's there.
12:22And so think about everything. We were doing a whole lot of things three months ago. And so let's just think about that. Well, you know, if you could run DeepSeq, probably 20 % of the cost of running open-handropic models in production with comparable, better latency, probably better reliability. If we don't have access to that intelligence in that form, I think it's just a massive loss. And as a country, we won't be able to innovate as fast because the cost of intelligence going down and control of intelligence, what we have seen just means more intelligence. Intelligence being embedded in more places.
12:56Yeah, an important note here that we didn't mention explicitly is that the state-of-the-art models, the ones that are most far ahead on the frontier, are actually still the closed-source, Anthropic, OpenAI, Google, etc. Yeah. What has been, actually, maybe you can just characterize workload a little bit, like how, of tokens being served on Base 10, like how many of them are from custom models of some kind versus like vanilla open-source today? It is all custom. It's basically... Okay. So like 95 % point. 95%. And I think that's really cool, to be honest. Look, we have two businesses. We have three businesses.
13:33We have three businesses right now. Should we help you count? No, no. So we have like dedicated inference, which is basically custom model inference. Your SLA is your SLA. Then we have shared inference, which is a shared inference endpoint, shared SLAs. And then we have a training business. I'd say 95 % of the tokens today are on the first business. And almost all of them, there's probably, yeah, for almost all of them, the customer is making some modifications to the model with their own data, specialized for the use case. And I think what's even more important is they might be compiling in different ways.
14:14No one is just running the vanilla open source weights. Like you might be customizing it for quality, but you mostly might be customizing it for performance. You made an acquisition of a research team a few months ago. You mentioned post-training customization. What was the rationale behind the acquisition? What is that team doing today? Yeah. So the rationale around the acquisition was, you know, we are infrastructure and product people. We are product people and now are really good infrastructure people. And we didn't have much of a research capability ourselves. And what we saw was the market moving heavily and heavily, like that we could accelerate the market itself with post-training resources, either productized or honestly, even just as resources for that market.
15:09So Paz was a company that was a base 10 customer. So they were post-training models and running them on base 10. And I think what they realized was that they would eventually need to become an inference company. And what we realized was like, hey, we really needed that expertise because it represents a way for us to get closer to the customer earlier and be able to support them all. and just made sense as a, like pairing them together. And just as I said in the opening statement here, which is, you know, as more and more post-trained models have come up, we've realized that the demand for people to either for software loops to do post-training or for post-training expertise is very high.
16:03And we're really, really investing in that. There are also a bunch of Australians, And, you know, I like to think that we had a bit of alpha there. But yeah, that's been fantastic. They're working with all sorts of customers. And it's also very interesting when you start, you know, we were doing a lot of research on the performance side and less so on the post-training side. It's interesting as we've started to do a lot more research on the post-training side, you start to see how linked inference and post-training are. And even when you think about stuff like quantization and when you should do that and how you train the model affects how you need to quantize for inference and how paired these problems are has become very apparent.
16:53And more and more, the post-training inference are both sides of the same problem. So because inference ideally will get more post-training where inference creates data, you do evals, you can now post-train on that reward function that you found with those evals and hopefully just set up that entire look. Plenty of folks from Ant and OpenAI, Sam, Greg, et cetera, have said in recent months that inference is super strategic, inference talent is strategic, capacity is strategic. So between that and post-training, these are very difficult to gather capabilities. I imagine that lots of your customers go to you guys for advice on how to do this progression of moving to custom models.
17:39Like, what do you tell people about the lifecycle and when they should invest in that? Yeah, I think it's, hey, go find, go prove to yourself with the best in class model that you have something worth optimizing. And I think, you know, a lot of, you know, if a customer comes to us, was that meme, which was like, it was like two years ago. It feels like no GPUs, pre-product market fit. It's like no post-training pre-product market fit is what I'd say. Yeah, yeah, yeah. It's what I'd say. So people that you're working with here are very at scale first. Yeah, they have a user signal that they know how to optimize.
18:13And they've shown that they can, you know, they can serve customer value and that they have something special around that value. And once you have that value, it's like, okay, now how can I do that better, faster and cheaper? With the idea being that, hey, if you need to be very good at customer support, you maybe don't need to be that good at coding and that a specialized model might be a better fit for that problem and you can do it better, faster, cheaper. What about the capacity side? You started with unifying capacity across all the clouds and new clouds. How do you think about this when everybody keeps talking about a supply crunch and a multi-year supply crunch?
18:47I think there's so much narrative around the supply crunch. And no matter, as much as we hear about it, I don't think people realize how bad it really is. There is very, very little Slack compute available. We run pretty large clusters ourselves, and we run them in uncomfortably high utilization. When I'm saying we're mid-90s utilization most of the time.
19:21We sit in 18 different clouds now. We have 90 clusters around the world across 18 different clouds. And like, you know, initially we started, we built this technology to be able to like kind of create one runtime fabric that spans all these different clouds and try to abstract that away from our customers as a way to think about reliability, latency, failover, all these things that we think are going to be very important for very mission critical use cases. that same technology, like just our ability to get compute wherever humanly possible has been really, really helpful in our ability to get supply.
19:58And what I mean by that is we can be introduced to a new provider in a different country and have it up and running with the whole base 10 inference stack. As part of the fabric. Half a fabric in half a day, maybe less. and that gives us enormous flexibility. Even for us, it is hard for us to grow. We have a, I think it's, yeah. We have a 4 p.m. standing meeting for the company where we basically like, how do we manage capacity for the demand right now? I think the second part, which people don't really, the second part that people don't really understand is that there are also a lot of suppliers right now that it's kind of grifty.
20:56You know, like I think, you know, they haven't run data centers before. You know, they don't understand SLAs, especially for inference. And so like, you know, even when there is capacity available, there's a lot of, like there's probably we run a lot more and we have redundancy so it's fine but if you you know there's probably like a dozen good like clouds and I'd probably like put like three or four of them in like the the gold tier and I think that just means that like supply like not only are we supply crunch we're supplier and operationally crunched onto people who can who can run these data centers as well.
21:39How far ahead can you actually buy capacity right now? In other words, is there any slack in the market if you buy two years ahead or five years ahead? You mean like contract length or actually like, hey, I want this in January 28? Either one, yeah. I mean, it's more the I want this in January 28 or at least I have some visibility into my future supply. Yeah. You could buy that, but you've got to also remember how quickly the market is moving. And that gets balanced somewhat off the fact that the H100 is such a great chip.
22:17And it's crazy. It's four years, four and a half years old. The price is going up still. Maybe it has a useful life for nine years. So that's good. But at the same time, yes, you can do that. But you're making a lot of bets as part of that. And then in terms of, I think that's the big thing that's changed over the last six months is that the term length that people want has just gone up. So if you wanted 1 ,024 B200s, which is, you know, from a good cloud, right now you're not getting that less than a three to five year contract right now with probably a 20 to 30 % TCV prepay. um so like actually what becomes important when acquiring capacity um is you need to have enough demand to supply it um to serve but then you also need like a low cost of capital um which is which is actually changing the dynamic pretty significantly does that does that impact how you think about going public as a company because arguably yeah i think you'd go sooner yeah exactly yeah i think you need like i i think the and i think there is demand for that um but i think you the pool, the, it also, you know, one of our, one of the, one of the realizations that we had recently and we're, we're software people.
23:42Um, and so we don't, we don't think like this all the time is that, you know, our business has like very interesting work and capital, um, requirements, like, you know, um, and, and, and I think, you know, even, and that as a result of that, it has very interesting financing, um, requirements. And we're not, at least right now we're not even going down to the down to the dirt there's also things you could do in terms of debt or other structures that yeah yeah and yeah i've learned a lot about debt yeah recently given the uh supply crunch uh inference being one of you know the top couple markets you're going after you have um plenty of people who understand this problem and therefore you know some competition how do you uh how do you think about like what are the factors that create a dominant player here or a winning player?
24:30Is it, as you mentioned, cost of capital? Is it access to supply? Is it software? Is it demand? Yeah. Just being excellent at everything. Yeah, it's... Look, I think what's so interesting about inference is... Is it operations? I guess it's actually cloud. Yeah, I think so. Yeah, I think like GPUs as a service is not sticky. I think that's been seen. Like customers generally just see that as commodity. Inference with the software layer included is incredibly sticky. You know, like just like, you know, none of our top 30 customers have ever churned. You know, we're talking like 400 % annual NDR around our business.
25:12And so it's like very, it's very, very sticky. So I think that software layer is very important. The optimist in me is like, oh, there's so much value in the software. And we will build the best software layer for inference that exists. I think, as I think is becoming clear now, access to inference compute is a strategic advantage. And I think that is the strategy that even the labs are going after, which is if we have all the compute, good luck running inference. Yeah, yeah. In a world of constraint compute, the number one thing to own is compute. And so just owning it in and of itself is an asset.
25:51And I think people underappreciate that. Yeah, you can't make a good hot chocolate without milk. Unless you're a vegan. Unless you're a vegan. No one wants a vegan inference. Well, I've got to ask you, people might want alternative milk, right? So, okay, like when you, the H100 is a great chip. People want a B200, they want a GB200. They want, of course, tons and tons of NVIDIA. When you think about making a bet, you know, several years in the future, Do you believe that there is a multi-chip world? What do you think happens from a compute perspective on the chip side? Yeah. I think diversification everywhere is the same way I want to water many models.
26:38We want to water most things. You'd be sad if it didn't happen. Yeah. And I think everyone would be sad. I will say to some extent, which is, yeah, and I think there will be inference-specific chips. I think you have like decode-specific chips, I think. And we're looking at - And NVIDIA said this too. Yeah, yeah. I mean, that was a whole Grok LPU thing. It's like, you know, I think that is very straightforward and makes sense. I think people really, really, really underestimate supply chain stuff with NVIDIA. Like how good they are at that. CUDA, how good CUDA is. the developer ecosystem around it.
27:16And, you know, we, the ability, like, to me, like one of the most important things as an infrastructure company in this moment is how fast you can move. And you can move fastest with NVIDIA today. And I think that is the reality. And like, it just like, given the scale that they operate at, given the scale that they operate at, it's hard to see. it's hard to see the I'm not saying it won't happen like the short term like in the next couple years how anyone's going to be able to compete for that. Especially with, you know, so much of the other players like what you need to be able to compete here is the ecosystem to form around you.
27:57And if you tie up all your supply with one buyer which, you know, a bunch of the other chip providers have done it's actually hard for that ecosystem to form. You know, like if you think about if you're a big lab and you have a proprietary deal with one chip type where you get 90 % of the supply, it's actually in your best interest to make sure you get 95 % of the supply and everything that's built for you and no one else could ever use it. When you think about reacting to the market, what do you think is happening with the actual workloads that you have to go invest in, right? Like obviously code agents and long horizon agents over time have become a big deal.
Read the full transcript
28:33People talk a lot more about CPU compute, video inference is different. I don't know if it's that. Sandbox is like, what's important for you guys to invest in now? Yeah, look, I think for us, all the runtime stuff is obviously very important. And what that means is like, what chips we run on, how we run, what kind of workloads we support. Do we get very good at diffusion transformers? Yes. Coding agents need sandboxes. We should code with sandboxes. There's all sorts of new speculation techniques to get faster imprints. We need to do that. even stuff like KV cache away routing and that stuff's a bit old now, but continuing to be very good at that and somewhat disentangling pre-fill and decode and starting to treat them as separate problems.
29:17I think that's something we are very focused on and we're seeing massive gains. That's at the runtime level. I'd say beyond that, everything we think about is how to create more of that loop between inference post-training because we think that just begets more inference. And so we will build a partner on almost everything there. So we're going to work with the best evals company in the world to make sure that's very well integrated, like BrainTrust into and around base 10. We will partner with or on the sandbox side build the best sandboxes experience that will exist. And then we'll create the best training APIs APIs to make it so continual learning becomes somewhat of a solved problem.
30:06It's not just like a discrete thing. That's, I think, the core base 10 product thesis. It's like, how do we build that loop? And then everything around that becomes, how do we make sure that we can do everything we can to ensure that gets as big as possible? That's access to compute. That's an infrastructure. Make sure we can get compute anywhere. Make sure we have access to our own compute. and then I think it's all the primitives that come after that just that just become incredibly like margin of creative both for us and our customers which is you know stuff like you know sandboxes and like the async batch inference like how we drive utilization by having a first class batch inference experience to me this is like what an inference cloud looks like it's that you are very good at inference and then you you start to do all the things tangential all that loop into inference and partner and where necessary and build where necessary.
30:59But we really do want to own, like start with that core inference story and then go down to unblock supply, create margin and go off the stack to unlock value. What would surprise people about some of the issues you discover only at scale? I'll give you an example. I was surprised when you guys ran into scale limitations, like fundamental limitations with some of the hyperscaler products that you were consuming. Yeah. And I'm because I kind of think of, you know, the AWS GCPs of the world as supporting infinite scale. Yeah. I mean, I think you just, and like, again, like I think very, very large companies that run services of big scale is probably the same stuff.
31:40All the edge cases just become... You actually experience them. You experience them. And like, you know, and I'll give you a few examples here. Like you see, you know, you start seeing, you know, yesterday we had, for the first time ever, we saw some kernel panic. And that only happened because some fluent bit worker was creating too many logs and the scale was too big and it was all into one node. And it was happening two terms at the same time by two different workers. So you see all like the systems level and kernel level problems. but then you start to see I think the craziest stuff is you start to see with LLMs that these run times are pretty immature even how we use KVCache is probably a little less sophisticated than most people see and we are starting to see the limitations of the current and the next set of primitives that need to be built from a scale security performance perspective But I think it's really at the runtime level and the systems level.
32:47And then, but the edge cases are, I'd say a lot more systems level than they are LOM specific. What are the things that keep you up at night? Capacity. I think, you know. Quick answer. Yeah, I think capacity. I think the other one is probably just this market's so big. And so like, it represents a moment when you should be as aggressive as possible. and really, we've grown a ton obviously over the last 12 months, the last few months, but the answer's always just go bigger, go faster. And I think that's really, really fun. It's also a little exhausting. And it's also like we are all in somewhat uncharted territory in terms of how fast and how big you can go and how things can get.
33:34But I think the big one is compute. I think there's no world in which there's enough compute to get the amount of value that we want to get out of our limbs in the next five to 10 years. Or we have to invent a lot of new stuff. Yeah. If we just talk a little bit about what you're learning scaling, 30X is like an aggressive thing to go through as a company. You've brought in a lot of really amazing talent, like Danny and Samir and Stephen Day, folks on both the technical and the go-to-market side. What do you think is working about how you are recruiting and scaling or what's your philosophy on that?
34:19We were very, very flat like until, I don't know, 12 to 18 months ago. I remember I went on a walk with a lot actually and a lot of it's like, you just need leaders. And like, it's actually like so contrary to everything. you know as engineers you're like oh yeah you overhead it's all it's all everything is overhead everything is overhead um and you once told me i think that you you didn't you're like hey sarah sarah what about we just have engineers instead of sales people yeah yeah yeah bad everybody learns it everyone's the same and we're all about i remember like you know you you said it so clearly at the time a lot and i think that's what we've noticed which is like actually having a leadership team that you can trust, that you can trust is so important.
35:07I think the two or three things that I would say is like, you want people where you can give them whole problems. And so like, you know, if you feel like you are micromanaging, if you feel like you need, if you feel like, you know, you have to be involved in everything, I think that's a bit of a cop out as a founder. Because you're just like, I just need to be involved in everything. It's like, no, you probably don't have the right people. I think the second thing is be very, very clear what you're optimizing for. Because I think when you're very, very clear what you're optimizing for, the people on and like, if it's something generic, like we want the smartest, hardworking people, like you can't do much with that.
35:45Like with us, what we cared about was, hey, actually, we don't care about a lot of people who have done this before. We care about people who think first principles. first principles. Work has to be a high priority, but they also have to be very kind and nice and care about the collaborative environment. We don't have a hero culture, very low ego. And if you need a manager, it's probably not the right place to be. But I think when you have that clear rubric the the people become very apparent that will fit into it and the people that don't um fit into it also become very apparent i think what's more like we've hired amazing people like you mentioned but i think what's a lot more interesting is like i think we've we haven't had a ton of like turnover there unnecessarily like people tend to work um because we because we have a very we are very clear what we want it took us a while to get there though what about the idea of like an operations culture and we were talking to alice and henry about this and she's like well the hard thing about cloud is actually just operations i slipped with a pager under my pillow for a decade i don't think i've seen you detached from your slack channel yeah for my phone is buzzing right now yeah i'm getting anxious so um and and you've been concerned before like do people get it like you know what is distinctive about that i i think I think like one, I think if you've worked at an infrastructure company, like we were once in a meeting with a bunch of AWS execs and this was, you know, like very senior AWS folks, all their pages went off multiple times during our 45 minute meeting, you know, like it's a, I think like it's very much like just a cultural thing.
37:33But yeah, like I don't, you know, our like inference can't go down and like, you know, we, you know, you learn to like, Like, you know, I think Amir, my co-founder, when his pager goes off, his seven-year-old said, is that a P0? Oh, is that a P0? And so, you know, I think that is, you just have to get used to it. That's the culture you live in. And it just changes the speed. But also it's, you know, becomes like, you know, a cultural thing. I think it's very, very, it rejects people that don't fit into it very, very quickly. Like engineers who avoid pager do. Yeah, you know, when we have pagers, we're like, everyone on the call.
38:15Like, you know, like there's been a joke that there may as well be a siren that goes off in the office. So people have been talking ad nauseum in the AI community about Jevin's Paradox. Yeah. Where if you decrease the cost, it's really a question around price elasticity and availability. If you decrease the cost of a good, say intelligence as a good, people actually consume more of it. Like the personal or business ROI of it, the demand for it goes up, not down. Do you see this? And are you working against yourself trying to make these models more efficient? Do people just use them more or less?
38:54Yeah, I think you think about this from a developer's perspective and a consumer perspective. I think consumers just want the best answers and the best experience that's somewhat governed by more intelligence to some extent. I think when you go to the developers, from the developer's perspective, they would insert more intelligence if you make it cheaper. Like that's, you know, and they will insert more intelligence anyway. But if you make it more cheaper, they'll insert a hell of a lot more intelligence. And you see this with agents. Agents are just longer running now. And I think that's what we have seen with the cost of inference going down, which is, you know, folks are just like, okay, we can run this for longer.
39:37or we can make it do a bit more work and we'll get to a larger end. I think like compute scales from an inference perspective as well. And, you know, I think we are seeing that with almost a lot of customers, which is, you know, they either start with like, this is the quality of answer I need to get to, and this is the amount of inference I need to do to get there, or this is the base level model that I can start with, that I can work with to get there. And I think the more we drive down the costs, what they realize is more intelligence just means better user experience. I just want a better answer.
40:12Better answers, better experiences, more dollars. More actions. More dollars, even more revenue. So yeah, I think inference going down just begets more. It is truly like, I think we're kind of in a world that is, you know, it is the lost market, right? Like even if there's AGI, all that's left is inference. Yeah. So you do not see in your customers a, this answer is enough and this action is enough dynamic. No. Yeah, it's going to keep going for a long time, it looks like. How do you view all this kind of evolving towards the future? So basically, this is one of the, it seems like it's going to be one of the biggest markets of all times.
40:47We have this massive shift where we're moving from software and seats and digitization into actual intelligence, selling units of cognition, selling agentic workflows. What does this all look like in a couple of years? What is your view of this future world? I think for consumers, it's the best possible thing, right? Like everything is somewhat smarter. You know, you get better care because your doctors have access to better tools. There's more, you know, like there's all this stuff about there being less software engineers. I think we just build more software. I think we just build a ton more software.
41:22And like, you know, I see, you know, we're not slowing down hiring software engineers. We're just building more things. and for the consumers that just means better tools more software all those good things it's almost like everybody has their own team for everything right you have an agent which helps with your doctor you have an agent that helps you learn stuff you have an agent that helps you organize your life it's a concierge it's a concierge yeah it concierge everything for everyone yeah and I think like what that means that's amazing I think that's great and education same thing you have concierge education you get personalized access to everything I think then you go one step back in how it affects developers, I think, and companies, I think if you don't embrace this, I think it's the extinction moment for a bunch of folks, which is like, everything needs...
42:11And I don't think that means that forward design needs Figma. I don't think that's a thing. I think what's more interesting is just like, all these workflow and software companies need to figure out what is the intelligent or intelligent inserted versions that drive the amount, all that user value for those end consumers that we talked about. Yeah, very exciting. Thank you so much for joining us today. Yeah, thanks guys. Find us on Twitter at NoPriorsPod. Subscribe to our YouTube channel if you want to see our faces. Follow the show on Apple Podcasts, Spotify, or wherever you listen. That way you get a new episode every week.
42:50and sign up for emails or find transcripts for every episode at no-priors.com.
From the publisher
Baseten CEO and co-founder Tuhin Srivastava sits down with Sarah Guo and Elad Gil to discuss the rapid growth of AI inference demand, Baseten’s 30x growth, and why inference is becoming the strategic “last market.” Tuhin Srivastava argues the application layer will persist because companies with unique user signals can encode value into workflows and post-train specialized models, citing examples like Abridge and support workflows. The conversation covers GPU capacity constraints, Baseten’s multi-cloud fabric across 18 clouds and 90 clusters, long-term contracting dynamics, the importance of the software layer for stickiness, evolving workloads, multichip possibilities, and operational lessons at scale.
Sign up for new podcasts every week. Email feedback to show@no-priors.com
Follow us on Twitter: @NoPriorsPod | @Saranormous | @EladGil | @Tuhinone
Chapters:
00:31 Baseten growth
01:55 Why the app layer wins
05:57 Serving frontier customers
07:55 Open source model mix
09:21 Chinese models and geopolitics
13:07 Custom inference dominates
14:22 Post training acquisition
17:10 When to invest in custom models
18:35 Supply crunch and data centerse
22:25 Longer GPU Contracts
24:09 What Makes a Winner
26:07 Multi Chip Future
28:19 Runtime Roadmap
31:08 Scaling Edge Cases
33:48 Hiring and Leadership
36:44 Operations Pager Culture
38:19 Efficiency Drives Demand
40:41 Concierge Everything Future
42:34 Conclusion




