In short
Base10’s CEO explains AI inference (serving models in production) and argues every company will want to “own its AI” via a continual learning loop. He details Base10’s approach: infrastructure/compute procurement, a core inference software stack for speed and reliability, and inference-adjacent primitives (post-training, RL, evals/routing, and execution sandboxes for agent code).
Key claims
hyperscalers can’t match specialized software; compute is necessary but software differentiates; “owned intelligence” requires running open/open-weight and post-trained models with customer data and feedback; inference will remain the biggest market even with AGI; agents increasingly become Base10 customers by deploying/debugging models in production.
Notable examples
DeepSeek V3 as a turning point for enterprise openness to open-weight models; customer use cases mentioned include healthcare (Open Evidence), legal/support (Sierra/Decagon/Bland), and go-to-market/coding agents (Clay, Cursor, WhisperFlow).
Guests
Tuhin (Base10 CEO). Interviewer: Alex (podcast host).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Chapters
Tap a time to open that second in VOUnderstanding Inference and Base 10's Role
0:45 to 2:50
A deep dive into what inference is and how Base 10 facilitates its execution.
“You know, they have the model, but they vertically own everything from the model all the way to the chips it runs on all the way up to the APIs that you are using to call that model.”
The Layers of Inference Solutions
2:50 to 6:20
Discussion on the different layers of problems Base 10 addresses for customers.
“We do this using like a pretty complex architecture where we're able to pull together compute from all sorts of different regions of cloud.”
Market Potential and AI Clouds
6:20 to 8:19
Exploring the vast market potential for AI and how Base 10 compares to major cloud providers.
“that's going to be divided between the people who are capturing inference.”
The Importance of Software in AI
8:34 to 13:20
Discussing why the software layer is crucial in the AI inference market.
“And you just said the software is what matters, and you have an acquisition that you just announced that I want to talk about today on that.”
Shifts in AI Capabilities
13:20 to 14:00
Tuhin reflects on recent advancements in AI technology and their impact.
“where in the last like eight to 10 months, it just really feels like AI has started becoming useful in a way that it was not before?”
The Rapid Evolution of AI Capabilities
14:00 to 19:33
Explore the significant shifts in AI capabilities and their impact on businesses.
“No company would be running LLMs in production even in like 2021, 2022.”
Founding a Company in the AI Space
19:33 to 27:08
Learn about the journey of starting a company focused on machine learning.
“from thinking our open weight model is going to be a thing is like to, oh, wow, this might be a majority of token volume.”
Vertical Integration in AI
29:26 to 30:15
Discussion on why companies are moving towards vertical integration in AI.
“It feels like everyone who has a stake in the compute build out is going vertical.”
Base 10's Strategic Decisions
30:15 to 31:29
Exploring Base 10's growth strategy and focus on controlling their tech stack.
“It's either to deepen customer value and ideally add more value to them at the end and hopefully skim off some of the top for yourself as you add that value, like that's one place.”
Scale and Growth Metrics
31:29 to 32:57
Insights into Base 10's token volume and revenue growth highlights.
“We do like 40 to 50 trillion tokens a day.”
Show all 19 chapters
Agents and Execution Environments
32:57 to 34:01
Discussion on the rise of agents using virtual machines for productivity.
“And speaking of that, a trend I'm obsessed with right now is the rise of these agents using virtual machines and browsers for people to get things done.”
Integration of AI Models
34:01 to 35:56
Exploring how Base 10 integrates AI models into their execution environments.
“And then we start to think about, hey, it's so interesting when you start to think about all these problems together because a lot of these sandboxes or these execution environments run on CPUs, not GPUs.”
Future of AI Agents in Society
35:56 to 37:50
Discussion on the potential widespread use of AI agents in everyday life.
“Do you think that billions of people are eventually using agents with virtual machines?”
Scalability of AI Infrastructure
37:50 to 39:24
Addressing concerns about the scalability of AI systems and data centers.
“think like you know could i see my my um my mom and dad being you know massive instinct users or or town users or 100%.”
Ethical Considerations in AI Development
39:24 to 42:01
Exploring the ethical implications and societal impacts of AI technology.
“A big theme on my show so far, both with Sam Altman and Zuck, was the data center backlash and I'm curious how that touches your world.”
The Importance of Grounding AI Discussions
42:01 to 43:20
Explore how AI's impact on various professions shapes our understanding and responsibility in its development.
“But also, that's the amazing thing about academic research and policy is that there's a counterbalance to these things.”
The Future of AI Inference
43:21 to 45:00
Analyze the potential shift from AI training to inference as industries evolve.
“Speaking of things coming fast, I've heard you say that inference may be the last market post-AGI, which is an interesting take.”
AI Agents as Customers
45:01 to 46:34
Discuss how AI agents are beginning to operate autonomously and interact with customers.
“I want to end here because I think it's a really interesting look at where things are going.”
The Human Element in AI
46:35 to 47:02
Contemplate the irreplaceable role of human interactions amidst rising AI capabilities.
“What is the last thing you and your team are going to be delegating to agents?”
Transcript
Automatic transcript. May contain errors.0:00Alex Heath:Tuhin, for people who don't know what AI inference is beyond it's how you serve the models, I'm really curious to hear you explain it because it feels like the most exciting, frenetic part of the AI stack right now. And you're thick in the middle of it with Base 10. And I'd love to know what it means for people who know the term, but they don't know much else. Yeah, absolutely. So what is inference? So inference is the way you serve models. Quite often, customers will come to us. They'll say, hey, we have this model. We want to use this open weights model. We have post-trained this model. We need to figure out how to actually use it in our applications.
0:42And so when you work with a closed frontier model provider, you kind of get this for free. You know, they have the model, but they vertically own everything from the model all the way to the chips it runs on all the way up to the APIs that you are using to call that model. When it comes to post-trained models or your own models or open-weight models, you don't get that for free. And so you'll come to an inference provider and you'll say, hey, I need to run this in production. Here is some scale, approximate scale I need to run this at. And then you'll come to company base 10. now there is so much that goes into actually running a model um and you know for us like that the infrastructure level problems so that is like hey where is the actual compute going to come from to run this model um so we take care of that um there is the orchestration piece which is hey like how do i um get this user request to the right chip with the right model on it we solved that problem and then there's like the chip level details where it's like how can i make this model run very, very fast.
1:47And so really, when it comes to inference, you're really like, you're not just looking for an API. You're not just looking for GPUs. You're looking for like a performance, reliable system that puts it all together and puts it behind an API. And Baystone kind of takes care of all that. All right, that was kind of marketing. So let's double click on that. Let's actually double click on what does that actually mean. So Baystone's like three different layers of problems that we're solving for our customers. So the first one is this idea of infrastructure and compute. So you need chips to run these models on.
2:19So for our customers, if you're a customer running this model at decent scale, you'll probably need thousands of chips to run this. If you go to the market right now and say, I need 1 ,000 B200s or GP300s, good luck. It's not happening. As they're gone, someone's already taken them. And so you'll come to like, what we do is that we kind of take care of this compute procurement from the standpoint of what you need or that's necessary to run inference on for our customers. We do this using like a pretty complex architecture where we're able to pull together compute from all sorts of different regions of cloud.
2:57So we sit on top of 20 different clouds in 90 different regions, which is completely abstracted away from our customers. So that's the bottom heterogeneous compute layer. I didn't even know there were 20 different clouds. Yeah, exactly. Well, you call them clouds, neoclouds, data centers, hyperscalers. You call them to us. They are sources of compute to some extent. Yeah. Okay, let's go up one layer. That becomes the core inference layer, and that's kind of what I described, which is like, hey, I care about running these models really, really fast and really, really reliably. And so that's kind of like what Base 10 has done for a long time now.
3:30That is what we are known for. And our customers come to us not only saying, hey, we want to compute, but hey, we need the whole software stack on top of that, on top of this, so we can run this so it doesn't go down. We can run this so it's very fast. We can run this with the right numerics and so on and so forth. That's the second layer. We can talk as much as you want about that. The third layer is around all the primitives that come on top of that. And that's kind of like, these are inference adjacent primitives a lot of the time. What I mean by that, these are things that either make your inference more powerful, They give you different paths and different types of inference.
4:04They make your models more powerful, but they're all related to inference in some way. So that might be something like post-training models and RL on top of those models, which not only helps you train more models, but has a bunch of inference baked into it. It might be something like sandboxes, which is how do you actually execute code. We're going to get into that.
4:22Alex Heath:Yeah. How do you actually execute code generated by these models? It might be stuff like evals and routing. when people talk about inference, they're kind of talking about all these problems together, all these problems together. But there are lots of different parts to it. It's everything except training, right? It's serving AI in production. It's when you get a chat GPT response back or you use Muse or GrokBot or any of these things, they're all inference. And inference is growing like crazy, it seems like as a market. but it also feels like it's hard to make sense of it. There's a lot of players.
4:58Alex Heath:They're all approaching it in slightly different ways. And I'd love to better understand how Base 10 differs in its approach from not only your direct competitors and inference, but the traditional clouds. Because I think someone listening or watching this who is maybe not as in the weeds as we are would go, well, why would someone not just use AWS or Azure or Google Cloud to run this stuff because they have AI? Why do we need these AI clouds doing inference? It's a good question. Look, I think the most important thing with all this is that the market is just ginormous. And this is important for a reason.
5:37But if you kind of project forward four or five years and there's some silly number associated with how much AI spend is. Let's call it$5 trillion is what they say in five years from now. It doesn't matter what that number is. It's somewhere between$2 and$10 trillion. Let's say AI runs about like a 50 % gross margin. So that means there's about, you know, $1 to$5 trillion of inference spent in the market. And, you know, if you think about how that's going to be divvied up between hyperscalers, between closed frontier models, between inference providers like Base 10, So any which way you cut it, that's somewhere between$10 to$50 trillion of market cap that's going to be divided between the people who are capturing inference.
6:26If you then start to break down how different folks are approaching inference, everyone has a slightly different take, I'd say. So there's the hyperscalers who have historically been really good at infrastructure-like things, but they've also been really good partners. So I think the hyperscalers will have inference solutions and they do today, but they're also really good partners to the inference companies themselves. We have really strong relationships with all the folks you mentioned there. The reason why you'd use us over hyperscale, really a lot of it just comes down to building specialized software.
7:02And it's like, hey, we do one thing, they do many things. And so when you're in a market that is moving as fast as this market, where you need to run stuff in production and when your inference provider is down, your product is down, you would generally tend to who will move as fast as you in the most customer aligned way and give you the most attention. And that is base 10 today just because we do fewer things and allows us to be best in class at those few focused things that we really do. Yeah. And I think it goes beyond that because I think there's just so many layers of this problem that I described.
7:39And we're kind of solving it layer by layer there. Whereas I think a lot of other folks and like, you know, whether it's hyperscalers, whoever else are like, kind of like, hey, I've got some basic inference solution that works off the shelf, but isn't really dialed in to serving users in production at scale.
8:08Alex Heath:to Granola, the AI notepad for people in back-to-back meetings. It works everywhere you do and lets you focus on what matters. Try it at granola.ai slash sources and use the code sources for three months off. This episode is also brought to you by Jira Byadlassian, where teams and agents get the context, coordination, and control to move work forward. Try it free at jira.com. That's J-I-R-A.com. The criticism maybe of the market you're in is that these inference clouds, they're just reselling GPU capacity at a very thin margin, and they really have to prove the software layer still. And you just said the software is what matters, and you have an acquisition that you just announced that I want to talk about today on that.
8:51Alex Heath:But can you explain that a little more, why the software matters so much? Because as you said, the chips are hard to get. And if you've got chips, you see the base 10 ads in San Francisco and the billboards, and it's very clear. It's like you're saying, I've got compute. Then that seems valuable. But at the same time, you're saying, well, the software is actually what matters. Yeah, look, compute's necessary, but the software is what differentiates, right? So it's like, I don't want to downplay the role of compute, or compute's obviously the market constraining thing right now. At the same time, there is so much value to be driven, to be built on top of that.
9:24The way we think about this is if you wind back a few months from now, there's been this really big push towards owned intelligence. And what is the idea of owned intelligence? The idea of owned intelligence is like, hey, because of cost control, because of data sovereignty, because of, you know, credible permanence, which is the idea that you don't want models just to disappear in one day, is very, very important for every company in the world, to some extent, to own their own intelligence. Cool. What does that mean? Well, what that means is that they need to have the capability to not only run open source models, but an open way models, um, in addition to close frontier models, but to figure out like when to use, which, how to use their data and the, and the user feedback, um, from their customers to improve these models.
10:12What customers have realized is like, like inference is one part of that. And I would argue it's probably like the backbone of that to some extent, but there's all these things that you need to do to unlock that continual learning loop that makes customers kind of internalize their intelligence. So that might be inference. That might be how do you eval each of these models as they come out. It might be how do you post-train models with RO environments to make these models better. It might be like, hey, the execution environments for when these models are producing code and doing their own things for them to do this work.
10:48And then eventually it's like, how do you get a better model out of it all that you run it again and kick off this whole cycle again. That whole thing I just described to you is what we call a continual learning loop to some extent. And that is what the software layer that we are... That's the software stack. The software stack we are building is all those things I just described to you and all of them in the entirety, which in a lot of ways become their own hyperscaler or new cloud when we build it. And compute is necessary underneath that until you have that thing, until you have that loop working, the compute's kind of useless.
11:23Or it is exactly what you said. It is reselling of, you know, taking access to it and putting it somewhere else. But the minute that we build the software layer for that continual learning loop, what we are doing is like providing every enterprise, every customer in the world, some credible advantage they are getting from running their intelligence. And what that does for them is it protects their margins. And as a result of them protecting their margins, they will interpaste a lot of that margin to us because we are kind of making them independent again to some extent.
11:56Alex Heath:Yeah, that's interesting. I haven't heard continual learning framed in that way. I mean, when you spend time in the labs, you know, the frontier labs, they're talking about continual learning in the sense of the models themselves learning and getting better constantly and the research loop closing and being aided by AI. But you're talking about it in like a sovereign company AI sense of controlling your own destiny. which I hadn't heard before. Yeah, well, I mean, it's exactly that, right? This idea of owned intelligence, like Satya has talked a lot about this, Jensen's talked a lot about this, which is like, hey, your user data, your signal from these models, the improvements to these models themselves will become the core IP of every company in the world.
12:40Every enterprise, the extent to which you thrive in an AI first world is the extent to which you own all those things. in order to own those things, you need to own that loop that I just described to you. And that is what we are building for. That is like, you don't want all that data to be the input to someone else's continual line loop, which is the loop you described about how the model provider companies get better at training their models. And that to me is the software stock. That is the differentiated value that you're providing. That is when this gets a lot more interesting than just bare metal compute.
13:19Alex Heath:Does it feel to you like it does to me where in the last like eight to 10 months, it just really feels like AI has started becoming useful in a way that it was not before? I mean, I think that, you know, post-ChatGPT, obviously like that product took off and kind of kickstarted this whole boomerang, but it was this kind of like better chat bot, Google search thing for a while. And I mean, even in my own work, both, you know, as a podcaster, but investor now too, its utility is incredible. And I'm curious if you've been feeling that as well because you started Base 10 in what, 2019, which was well ahead of all of this.
13:58Alex Heath:The models were not capable at all. No company would be running LLMs in production even in like 2021, 2022. So yeah, I wanna go to the founding story and how you saw this when you did, but also do you feel that what I'm describing about the shift in AI right now? I think there was something in December that happened where these models just kind of had a capability shift. I think it's happening again right now, like in the last month where there's like another capability shift. But the last like six, seven months have been insane, right? Because so much has happened. Like obviously it felt like coding models finally got really good.
14:29The compute shortage happened. Open weight models started getting very, very good. And like, you know, it's like release after release after release. It's just like, oh, there's another one. There's another one. There's another one. And I think what has happened as a result of this is that one, people are paying attention to intelligence. So intelligence, people using intelligence and enterprising using intelligence has skyrocketed. As a result, earnings calls have become about using AI. Copies are being judged by how much they're spending on AI. That in turn is like making them rethink, hey, what is our long-term cost strategy and how do we, one, get AI in more places and two, how do we do it in some way that doesn't blow up our companies?
15:14And I think that is just taking over the entire narrative, but also just pushed adoption even further.
15:19Alex Heath:It's wild. It's wild. And I can't imagine being inside it the way you are. I mean, before we started this, you were saying, I may need to step away because I'll get paged because one of our customers needs help because everyone's growing like crazy. Before we get into more stuff, go back with me to 2019. You're starting this company with your co-founders. You were in banking before, is that right? How did you see this? because this market didn't exist then. Yeah, my background is that I grew up in Australia. I moved here for uni. Graduated uni in 2009. I did what everyone does in 2009. It's like, oh, you go work in banking.
15:58That's right. So I moved to New York. I worked in banking. I worked in infrastructure finance. Basically doing project financing for infrastructure projects like toll roads and bridges and airports, which is interesting. They're all data center people now, funnily enough. After that, I started working in technology. I ended up working at a lab in Boston, working on using machine learning to predict and diagnose neuromuscular disease. That was in 2012. From 2012 to 2019, machine learning was very, very early. It's kind of more classical models than large language models. If you go think about what machine learning was being used for 15 years ago, it was stuff that like paypal was doing and stuff that facebook was doing and a lot of it was like fraud classification it was recommendations and like we were kind of working in that paradigm for a long time i started a bunch of companies didn't really go anywhere 2019 um started another company and we were really thinking at the time like two things one is like you know my two co-founders of phil and amir um and then soon after that punkage um but phil and amir you know they're my best mates and so really the idea was like how do i how do i start a company with my friends that's number one number two number number two which was just as important was this idea that in 2019 we kind of knew that machine learning was going to be a big deal like opening ad had been around for a few years at that point um they were doing stuff like you know smart people were going to work there um but it was still very early like there's no there's no commercial proof what to happen but we're like look machine learning looks like it's gonna be pretty pretty big we don't know what that's going to look like.
17:37We had no idea. We're just like, oh. That's why I'm so in awe of all of our customers because I think application layer work is so challenging. Actually understanding that end user problem in a domain requires so much product expertise and domain expertise that frankly, we just didn't have in that way. But what we knew was a mission. We knew a bunch of our machine learning. We thought it was going to be a big deal. We're like, oh, well, if we believe in this and we know a bit, let's go to build a picks and shovels business alongside to support machine learning. And, you know, as long as the market is big enough, we'll have a shot of building something big.
18:12I think that what happened in 2022 was the market just accelerated in a way that none of us expected. And then like quarter after quarter, it just feels like, you know, we're like, surely this is it, right? And then, you know, and then it takes another turn. I was chatting with some folks that, you know, you work with, Guy and Effie and Jill about this Monday. And basically this idea that even in 2022, when they started thinking about machine learning and AI, it felt really early. It didn't feel like it was obvious then either. And it just feels like quarter after quarter, it's like it just come online in a way that it's kind of gone beyond all our wildest dreams.
18:51And so in 2019, when we started the company, we were definitely building an infrastructure company. I think that infrastructure company had many different facets. One facet of it, which seemed to us at the time was the least important facet, was being able to serve models. And then that just became the entire company.
19:11Alex Heath:And you raised at$5 billion valuation in January of 2026 and then$13 billion valuation in June. So in six months, you almost tripled your valuation. What changed in those months? Is it what we've been talking about? Yeah, it's all those things, right? Which is like, I think in the last like six months, what has changed is that we've gone from thinking our open weight model is going to be a thing is like to, oh, wow, this might be a majority of token volume. Do you think open will be that? I think open and custom and owned intelligence will become, I think the future is mixed intelligence where you're using close frontier, using open-witch models, using post-trained models.
19:56But I think custom frontier models will probably be some version, some percentage of the models that are uniquely served to needing the most powerful models in the world at all times. And you think about right now, think about what AI could solve three months ago versus what AI can solve now. And AI could solve a lot three months ago or two months ago even. If you can run that two-month-old intelligence for 90 % of your tasks at a fraction of the cost and own all that data sovereignty and ownership narrative that I talked about earlier, like that's a no-brainer and it's very rational.
20:31Alex Heath:How are you feeling about the state of open source though? Because the US is so behind, China is still dominating. I think a lot of people, a lot of tech leaders are worried about what's happening. And a bunch of CEOs were recently advocating for the government here in the US to protect open and open weights and open source. But yeah, curious what you make of what's going on right now. Yeah, look, I think it's both a very exciting time and an interesting and challenging situation, right? So look, we want more intelligence everywhere. And I think, you know, open weight models are fundamentally a very important part of the ecosystem.
21:09Like, you know, in absence of open weight models, you know, all that ownership narrative that I just told you about is gone. Like you need it. They're a necessary precondition. So we understand their importance to the independence and sovereignty of American enterprise. So that's why I'll say one. Two, so I feel really good that there are options now. So that's amazing. On where they come from, this seems temporary to me. I think there will be American open-way models. I hope from what we can tell, there'll be really good American open-way models in weeks, not even months or quarters. Soon enough, there'll be really good American open-way models.
21:44And I think, you know, once that flywheel kicks off of like, hey, we have stuff that is competitive with Chinese open-weight models, I think a lot of the narrative that you're seeing right now, which I think is a little bit more fear-driven and like, are we falling behind -driven than there's capability-driven, if that makes sense, I think that will kind of wash away as well. I'll give you a good example of this where all the open-wind stuff really kicked off 18 months ago with DeepSeek, when DeepSeek V3 came out, which was a Phanteks model came out over Christmas of 24. The first, you know, we were obviously serving that to a bunch of customers, but the first thing I did was I was like, hey, okay, we've met so many enterprises over the last six or seven years.
22:26I was like, let me go and talk to 10 of them right now and see how they feel about them. And the reality was that I'd say 90 % of them, if not more, were like, look we would never use this as a production this is you know like like we don't know where these are comfortable we don't really understand how this works like you know what are they being trained on um yada yada yada and and i'd say they were like deep tv 3 was interesting because i think it was like the first time you had something which approached the frontier but was still behind you know it's still like three six behind and like it wasn't clear how you'd run them and there's all this like uncertainty around uh was most of like and even though it's behind it was such a big
23:01Alex Heath:moment everyone was freaking out the air race has been reset yada yada yeah 100 and it was like it was like the paper that they trained it on like a fraction of the cost and blah blah blah fast forward a year now you have all sorts of models you know you have like you you even have some really good american models like inkling but then you have like um you have glm you have kimmy you have deep seek and you know hopefully a nematron the nematron family's getting really good from nvidia as well um all that narrative from 18 months ago is completely evaporated every enterprise we talk to now is very excited to run open source, open way models.
23:34And I think even, even like debating where they're coming from isn't really relevant. I think they're just like, this is just the new world. And turned out that like fear and certainty and doubt with more just capability driven as opposed to actual, what does this mean driven? So the minute they got good enough, everyone's like, can we use them? And, and like, you know, the infrastructure needs to exist. Open way models need to exist to, to run them. But like, we're in a really good spot right now in terms of owning that entire stack that's necessary to be able to run these models of production.
24:04I think as that stack matures, and obviously we are to some extent building that stack, we are building that stack until that stack matures. And as that stack matures, I think the barrier to adoption of open-way models is going to go down as well.
24:19Alex Heath:Does open have to keep growing for base 10 to keep growing? Are you really kind of tied at the hip with this? I think that's a good question. Like the intellectually honest answer is yes. You know, there's probably people who don't want me to say that, but the intellectually honest answer is yes, which is like we, you know, we believe in the open-weight ecosystem and we want to give back to it and push it as well. Like, you know, we launched something called Base Labs a few weeks ago, which is our research lab, which is like, hey, how can we make it? You know, we acquired a research team at the end of last year.
Read the full transcript
24:55They're doing really good work and we're just like, how do we put all this stuff out as well? But we are fundamentally, you know, behind the open-weight ecosystem and like we, the extent to which, I don't think open-weight models need to exist, but I think customers need to want to own their own intelligence and that needs to be an idea and concept that matures aggressively. And I think open-weight models are just an accelerant to that. So like, does that make sense? which is like, you know, they're like a core ingredient of that.
25:26Alex Heath:And does BaseLab signify that you're going to start doing your own models, BaseTent? If we need to, like we will do whatever is necessary for customers to own their own intelligence. We post-train models extensively with customers today. So what that means is that we take a base model and then we take their data and we set up their learning loops to be able to, you know, have those models to run them in production. Today, we are not doing large-scale pre-training. there's folks like Nemotron who are doing a lot of that stuff and we're building and we're part of the Nemotron coalition and we're doing work with them there.
25:58Look, at some point, if it's like, you know, we feel like we have a, we have a differentiated advantage to training models ourselves. And now, you know, the compute situation starts that it allows it. Yeah, a hundred percent. Like I think it'd be, you know, that is just our way of giving back to that ecosystem that we believe in and we're trying to power.
26:15Alex Heath:Mercury is a modern take on banking built for startups like mine. When I decided to start my media business, Mercury was by far the most straightforward, full-featured banking solution for me to set up quickly. The interface is intuitive and simple, saving me valuable time every day. I use Mercury to track my spending, bills, and invoicing. I love that I can delegate permissions to my team so they can keep things running for me in exactly the way I want them to. My favorite part is how forward-looking Mercury is with AI. Legacy banks are stuck in the past, But Mercury is built for how modern software works today.
26:49Alex Heath:I use its built-in command assistant to analyze cash flow and help me move money. And Mercury also connects to other AI tools like ChatGPT and Cloud. I use this feature all the time, and the folks at Mercury actually let me know that I'm one of the top users of it. So trust me. It's finally easy to get real-time financial data about your business wherever you need it. Visit Mercury.com to learn more and apply online in minutes. Mercury is a fintech company, not an FDIC-insured bank. banking services provided through Choice Financial Group, and Column NA, members FDIC. I spend a lot of time context switching between meetings, often with no time to process one before the next starts.
27:26Alex Heath:Thankfully, Granola runs in the background the whole time. It's an easy-to-use AI notepad for meetings that works everywhere, even on phone calls. I use Granola to recall what was said in meetings and create helpful summaries. I use it every day to stay on top of what I need to get done with my team. It connects to my email and suggests follow-ups for me to quickly review and send, saving me valuable time. Granola isn't just a core part of my workflow, it's basically my second brain. Try granola at granola.ai slash sources and use the promo code sources for three months off. AI is only as useful as the context it has.
27:59Alex Heath:But when that context is scattered across tools, threads, and DMs, your team and your AI agents are flying blind. That's the problem Jira by Atlassian solves. What's the goal tied to your project? What got decided last week in Slack DMs? Atlassian's teamwork graph pulls all of the valuable pieces together, from Jira, Confluence, GitHub, Slack, and more, so nothing falls through the cracks. You get 44 % more accurate results with 48 % less token usage. With Jira, you can easily share your work context with AI agents you already love, like Claude, Cursor, and GitHub Copilot. Assign them work directly or connect your tools through MCP.
28:36Alex Heath:All of this lets you spend less time digging through endless links and messages, chasing down what got decided and by who, and spend more time actually shipping. Learn more at jira.com. That's J-I-R-A dot com. Framer is the AI website builder that brings agents into the same canvas where your website is designed, managed, and published, so you can move faster without giving up your taste or control. Framer powers the Sources podcast website at podcast.sources.news, where you can find new episodes, transcripts, and a lot more. Learn how you can get more out of your site from a Framer specialist or get started building for free today at framer.com slash sources for 30 % off a Framer Pro annual plan.
29:18Alex Heath:That's framer.com slash sources for 30 % off. Framer.com slash sources. Rules and restrictions may apply. It feels like everyone who has a stake in the compute build out is going vertical. You mean you've talked about NVIDIA, which I know they're an investor as well. But like they've got Nemetron now. They're doing models. They just bought Hugging Face. They're going more vertical in the compute stack up and down. And it feels like everyone is realizing this is such a big market. It's so important to control your own destiny that you've got base labs. You know, you're thinking about the model layer and how base 10 can play there, not just at the compute and the software stack for inference.
29:58Alex Heath:So I'm wondering, do you agree with that take that kind of the incentives of everything right now are driving companies in your position to try to own as much of the stack as possible? Yeah, it's a good question. I can just speak for ourselves. Look, the reason why you own more and more of the stack is for two reasons, right? It's either to deepen customer value and ideally add more value to them at the end and hopefully skim off some of the top for yourself as you add that value, like that's one place. And the other one is like to the extent that you feel blocked by the rest of the ecosystem.
30:31The reason why like, you know, we think about going like building up and down the stack is like, you know, we build up the stack and that's all the primitives on top of it first I talked about, that third layer. Because we think that is how we accrue more value and add more value for the customer. And like we own more of that learning loop. That's like a very value-driven thesis. I think we go down the stack because, you know, we need to be in charge of our own destiny. And we need to know that, you know, we think that in 12 to 24 months, we will be, you know, one of the largest individual users of compute on the planet.
31:10Like, you know, like that is the way that that is, you know, there'll be the labs and then there'll be based in, you know, like that is the way that this is trending. in order to make sure that we can continue to grow at the pace that we want to grow at we just need to at least believe that we have control over some of those things that sit downstream of us or upstream of us however you want to think about it and like that is why we
31:32Alex Heath:vertically integrate in that way can you give me a sense of the scale of base 10 today you mentioned you work across dozens of clouds can you talk about revenue growth custom number of customers anything that kind of puts a weight to what you do? What's the scale that we can do? We do like 40 to 50 trillion tokens a day. That's a lot. That's bigger than a lot. And how much were you doing six months ago? I think token volume has grown 40x year on year. Wow. Yeah, it's kind of insane. And then, so that's one. Two is like, you know, we have thousands of customers, I'd say, you know, the way to think about base 10 is that if you have used any of the application layer companies, and we've probably touched part of your workflow today, probably in the last few hours.
32:22And tokens that you have consumed have probably run through Base 10 through all the apps you use, through all the products you love. And then in terms of revenue, I don't know what we published today, but I think we have 10x in the last 12 months. It is everything about our revenue. about our business is just kind of multiplying as the market is accelerating. And like, that's, it's a weird thing. It's a weird moment that we're like, we're very, very blessed to build in this market. And then obviously there's a ton of like macro tailwinds that are continuing to push.
32:57Alex Heath:And speaking of that, a trend I'm obsessed with right now is the rise of these agents using virtual machines and browsers for people to get things done. I think, you know, very at the consumer level, there's Muse, which meta just shipped. I had Zach on the last episode talking about that. Instinct is taking off in Silicon Valley world. There's startups like Town and then there's GrokBot and ChatGPT work, obviously. But the idea is you take a model and you give it a harness that includes a computer and a browser and the ability to log in and do things. And most people in the world who use AI have not experienced this yet, right?
33:35Alex Heath:They're still on that chatbot paradigm of before. But this is really starting to take off. And it also implies that token volume continues to go through the roof, I think. Yeah. And you just made an acquisition in this space. Yeah. So I want to understand that and how you think that is going to start intersecting with what you do at Base 10. Yeah, look, I think this just becomes, you know, these runtimes and execution environments for these, the place where agents do work. That's the way to think about it. It's like, you know, and like for us, like this is like our first version of like time to build stuff specifically for this idea of like agents, you know, like instinct, like town like grokball like where they're going to need to do their work and we've been thinking about this space for a long time now um we met um paul and black soul and the entire team and we were just blown away by the technology that they've built and like really the idea was was like hey like i don't think inference and sandboxes are different problems i think they are extensions and related like you know where these models run um is where the execution environment should be and that makes a lot of sense to me.
34:41And then we start to think about, hey, it's so interesting when you start to think about all these problems together because a lot of these sandboxes or these execution environments run on CPUs, not GPUs. Then you start to think about that own infrastructure thing again. It's like, oh, should these GPUs and CPUs be co-located? Should they, you know, how do you, and like really we start to think, it is both, again, the extension of the value that I talked about earlier for our customers where it's like, you know, if inference is primarily going to be driven by agents in the future, we're also going to need to understand where the agents are doing their work and communicating themselves.
35:19And that comes into the execution environment, it comes to the browser use, and that's where we start to think about Blaxel and how that fits into base 10. It's very, very natural. But then we start to go down the layer again and think about the layer of heterogeneous computers. Like, how does that change when these two things are co-located? And like, that is why this is such an interesting company. I think like it's so complex. And I I think, you know, I just want to bring this back to the first question you said about the competitive environment. Like inference is a very, very, very special workload and it needs like new hyperscalers and new clouds to be built around it.
35:48And that's like, that's really how we think about not only what we do, but how we act acquisitive and how we join forces with folks like BlackSoul.
35:56Alex Heath:Do you think that billions of people are eventually using agents with virtual machines? Like, do you think this gets really big? I mean, you kind of seen it with like instinct and, you know, like. And love to instinct. I love it. I use it, but it's, I don't think it's huge yet. I don't think it has millions of users. Oh, no, no, no, totally. But you could see it, right? Which is like, you know, we've been trying to solve this do work for me problem for so many years. Like think about how many chat spots there are where like, you know, can you make this reservation for me? Can you pay my taxes?
36:26Can you, can you, can you, um, book a flight for me? I think this fundamentally changes how we do work. I think this is what's so interesting, which is, you know, I think about our customers a lot. And I think about customers like open evidence, which is like, you know, basically brings intelligence to the fingerprints of frontline physicians. Or a bridge, which is like making healthcare providers be more present by like, you know, being the ambience to cry rather than making them take notes. Or I think about folks like Clay, who are like kind of changing how go-to-market teams, go-to-market teams.
37:01work or even what the texture of go-to-market looks like or you obviously think of stuff like cursor um which is like coding or i don't know if you're familiar whisper flow whisper flow is awesome um when i work walk around our office in jackson square i see a bunch of engineers whispering into their microphones to their agents yeah um and what i see with all these things is like we in a lot of ways like tech and like the communities that we are in are really really early adopters but you see the impact it's having on healthcare you see the impact it having on the legal field or with support with companies like um sierra and decagon and bland and so on and so forth there's just no way that it would make sense for all those productivity gains and the changes in how they work to not flow to the consumer experience i think like instinct in town and these are just like the first example these are just the first examples of that and i think i think like you know could i see my my um my mom and dad being you know massive instinct users or or town users or 100%.
38:01And I think that's where things get really exciting for how much more we have to... Again, this goes back to the first thing you said. It's like, Inverge is about to get 100 or 1 ,000 X bigger. And that's the crazy thing. It's like the amount of penetration we actually have in the market right now, it feels like 1 % or 2%, not some saturation point.
38:23Alex Heath:Is it scalable though? If hundreds of millions, eventually billions of people have agents or multiple agents running computers in the cloud, don't we just need a lot more data centers? Like, is this actually scalable? Yes. I've answered that. Yes. We need a lot more data centers. We need a lot more compute. But you also got to remember that we're aggressively lowering the cost of inference as well. And so I think both things are true. One, we need a lot more inference optimization. And two, we need a lot more power to show and land to do data centers. and we'll get more efficient in building data centers as well.
38:58But the good news is that, you know, there's a lot of land. It's like, you know, like the country that I grew up in, it's like, you know, like the majority of it's unoccupied. And, you know, there's so much space. In the U.S., we're obviously dealing with stuff, but, you know, we are thinking about and there are really interesting people thinking about how we create energy in a sustainable way. How do we figure out land? How do we make these data centers more compact? and I think that will just become a very big topic of conversation in the coming weeks, months and years.
39:29Alex Heath:A big theme on my show so far, both with Sam Altman and Zuck, was the data center backlash and I'm curious how that touches your world. Is it something that impacts you, Base 10 and your customers? Is it something you're thinking about a lot? People are really reacting viscerally to what's happening and I'm curious how you think about that. Look, does it affect our business? yeah like yes like you know like we we need a lot of compute and we need to scale and we need to figure out how to how to get intelligence in the hands of more people and that's a real lot of compute and i think you know all these things to some extent just slow us down i think at the same time like you know there are there are there are both valid concerns about like you know what this means for these talents and and and the countries where we're putting these things up but also like a lot of education to do i think we've done a pretty terrible job at the industry so far kind of bringing everyone else along and describing, hey, what does this actually mean?
40:22How does this change things? To me, it's more just, it just shows to some extent the responsibility that we all have to be able to make sure that this thing that we have seen the value accrue from grow to the potential that we see it has. I don't know if it helps for us to be going and spreading FUD around these things. But I think to me, it's a lot of just like education driven as opposed to anything else?
40:47Alex Heath:Well, I think it's very tough because, yeah, I think a lot of researchers at Anthropic especially, but also OpenAI, they really do believe that what they're building has potentially negative effects. I mean, everyone has seen recently the ex-Anthropic researcher who said, you know, 10 % chance it may kill humanity, like most viral tweet of all time in terms of views within like 24 hours. Is that true? That's wild. It's true. Yes. I think Elon confirmed that. And so you see that and then that researchers on Anderson Cooper later that evening. And it's hard to balance all these things because you just listed a bunch of your customers who are doing genuinely impressive things with AI, helping doctors, helping lawyers be better at their jobs, et cetera.
41:29Alex Heath:But at the same time, there's this existential just like fear and dread about what AI will do. And I'm wondering if a company like Base10 that sits in such a critical part of the stack, what role do you have to play in this conversation? Have you thought about that? Yeah. Look, I think a lot of it for us is just two things you can do, right? There's one, make sure we are listening and we understand the actual environment. We can't live in a bubble. This is kind of the thing I said to you earlier as well, which is living in San Francisco, we just have like a very specific view of the world or we the way we see it and like seeing how it's affecting doctors how it's affecting lawyers like that's that helps ground you to like hey this is thing and then two it's like telling those stories is very important and and highlighting the use cases and three you know i think the biggest thing is like look we have had lots of powerful technology for many years you know like there's been um and the way we the way we work around them is to figure out like what are the guardrails and what are the balances that we put to make sure that they don't cause undue harm.
42:37But also, that's the amazing thing about academic research and policy is that there's a counterbalance to these things. It's like when we encourage security hackers to find day zero vulnerabilities, and that is so we can build an ecosystem around that to patch them and fix them and roll them out. I think the same thing will happen in AI. I think the biggest challenge right now is just like things are moving very, very, very fast.
43:03Alex Heath:Yeah. And we're not having that time to react. I think that is like probably where the responsibility comes in, which is like take stock of what's happening and just be thoughtful about the future we don't live in and then make sure that we bring people around and along and, you know, maybe don't drink too much of our own Kool-Aid. Yeah. Speaking of things coming fast, I've heard you say that inference may be the last market post-AGI, which is an interesting take. the idea that maybe the labs just reach their training goals and AI training stops and the billions of dollars spent on training shifts to inference.
43:39Alex Heath:Do you really believe that? How do you see that practically playing out? I just spent a lot of time in OpenAI and they were talking about AGI with Astra. Sam and Greg told me they basically got it. They feel like they have AGI, but the world is continuing. New models are still shipping. And so I'm curious if even Astra and what AopenAI said recently has changed your view on that or not. All that claim is that we just have a lot of AI. And whether it's the final market or not, it's definitely going to be the largest market, that's for sure. Running these models in production and serving and user value and agents doing work this way is going to be massive.
44:17The final market claim is really just that even in a world in which you have AGI, And again, I'm not educated enough to be able to speak to that.
44:30Alex Heath:The researchers who are building it can't speak to it either. Don't worry. They all have different opinions. I can't speak to that. But what I can tell you is, look, without AGI, Infrance is going to be the biggest market ever. With AGI, well, the only thing left to do is for these models to run. It'll be the only market left. Because these models are going to run in loops and do everything for us. And if that's the case, it's like, well, we're going to need a lot of computing and a lot of software to be able to run these models really, really fast. And that's what that claim is. Yeah, I think I believe that.
45:01Alex Heath:I want to end here because I think it's a really interesting look at where things are going. You all shipped, I believe it was an MCP server and a skill so that coding agents can use Base 10 directly. So does this imply that agents are becoming your customers? Yeah, agents are deploying models in Base 10 for sure and running them, which is kind of wild. agents are deciding which model to use when. And how recent of a phenomenon is this? I think this has been happening in some form for a while. The agents are just getting good and the models are getting good. Are you imagining a future where your customer support, instead of you getting paged mid-interview to go deal with a customer that needs something fixed, like base 10 agents are talking to customer agents?
45:44If you go into our incident channels when things happen, the agent is debugging for while we are working on figuring stuff out, the agent's debugging by itself. It has access to all our code. It has access to everything about the model. It knows everything about the architecture. It knows about the customer issues come up and it's like, hey, there might be something unique about this model that's using this data in this way. You should go check this out. And then the extent to which we give it power to do that thing itself and patch it, and today we do, and we run stuff in production and these are really important things and we don't want to delegate all that such large impact actions.
46:24But I think for sure, I think all these problems just start getting solved by agents. And what happens is the human skill just has to change to figure out how to interact with the data to give them the most leverage and vice versa.
46:35Alex Heath:What is the last thing you and your team are going to be delegating to agents? The last thing? Hanging out with each other? I don't know. The human contact? I went to dinner with a few of my colleagues I'll just say that. That was a highlight. There's no time I'm delegating my social interactions. But is that maybe all we've got left? That's all we've got left. I mean, that's all we ever had, is what I would argue with you. All we had was people. All we had was people. That's a great place to end, I think. This is getting way too philosophical. We got to end this. No, Tuan, I really appreciate it.
47:07Alex Heath:I appreciate you chatting through all these things with me. Thanks for joining. Thanks, Alex. Thanks for having me. Banking should feel like modern software. Get everything you need in one place. Visit mercury.com to learn more and apply online in minutes. Mercury's a fintech, not a bank. Check the show notes for details. Granola is the best AI notepad I've tried. It works everywhere, on a video or phone call, in person, or an Apple Watch. Try it now at granola.ai slash sources, and use the promo code SOURCES at checkout for three months off. Jira by Atlassian is where your team and your agents work from the same context.
47:42Alex Heath:Try it free at jira.com. That's J-I-R-A dot com. Framer is the AI native website builder that lets you build faster without giving up control. Visit framer.com slash sources for 30 % off. Rules and restrictions may apply.
48:14Thank you.
From the publisher
Tuhin Srivastava is the CEO and co-founder of Baseten, an AI cloud startup recently valued at $13 billion. He tells me why every company will want to own its AI and what’s driving the explosion in demand for inference.
We discuss the rise of agents using browsers and virtual machines to get things done and Baseten’s acquisition of Blaxel to power that shift.
We also talk about the data center backlash, why companies are embracing Chinese open models, his plans for Baseten’s new research lab, and why he thinks inference becomes the only market left after AGI.
Thanks to the show's premier sponsors: Atlassian, Granola, and Mercury.
This is a public episode. If you'd like to discuss this with other subscribers or get access to bonus episodes, visit sources.news/subscribe




