In short
Eye On A.I. Podcast Episode Notes
Episode Overview
- Title: #239 Tuhin Srivatsa: How Baseten is Disrupting AI Deployment & Scaling in 2025
- Host: Craig S. Smith
- Guest: Tuhin Srivatsa, CEO & Co-Founder of Baseten
- Release Date: [Insert Date]
- Sponsor: Thuma
Episode Summary In this episode, Tuhin Srivatsa discusses the challenges of AI deployment and how his company, Baseten, is providing innovative solutions to streamline the process of deploying and scaling AI models in production environments. As enterprises increasingly adopt open-source AI models, the episode explores the hidden costs of AI inference, the shift from closed-source to open-source models, and the future of AI infrastructure.
Key Concepts and Discussions
- AI Deployment Challenges
- Current Landscape: Many enterprises face significant challenges in deploying and scaling AI models, particularly with existing platforms like AWS SageMaker and OpenAI.
- Bottlenecks: High costs, complexity, and lack of transparency are major hurdles in AI deployment.
- Baseten's Approach: Baseten simplifies AI model deployment, making it faster, more affordable, and more efficient.
- Transformation of AI Infrastructure
- Inference vs. Training: Baseten focuses on optimizing the inference aspect of AI deployment rather than just model training.
- API Integration: The company provides a comprehensive suite of tools that allows enterprises to manage compute resources and deploy models seamlessly.
- Shift to Open-Source Models
- Market Trend: A significant shift from closed-source to open-source AI models is emerging, providing enterprises with more control, security, and cost-effectiveness.
- Model Performance: Open-source models have begun to rival closed-source models in quality, making them a viable option for enterprises.
- The Cost of AI Inference
- Hidden Costs: The episode highlights the often-overlooked costs associated with AI inference and the importance of optimizing performance for better efficiency.
- Why AI Models Fail in Production
- Common Pitfalls: Discussed the reasons many AI models fail when deployed and strategies to prevent these failures.
- Future of AI Infrastructure
- Growth Forecast: Tuhin predicts dramatic growth in enterprises using AI, especially with the increasing adoption of open-source models.
- Emerging Technologies: The role of new chip technologies (beyond NVIDIA) and their potential impact on AI deployment is discussed.
Actionable Insights
- For Enterprises: The episode emphasizes the importance of considering open-source models in AI strategies for better cost efficiency and control.
- Developer Experience: Baseten aims to significantly reduce the time required to deploy models, contrasting the lengthy processes of traditional platforms.
Key Moments in the Episode
- (00:00) Introduction to Tuhin Srivatsa and Baseten.
- (01:50) Discussion on the importance of AI infrastructure.
- (09:17) Insights on why most AI deployments fail and how to address these issues.
- (20:44) The future challenges of AI scaling anticipated in 2025.
- (37:05) Examination of the reality vs. hype surrounding AI technologies.
Conclusion This episode of Eye On A.I. provided a comprehensive insight into the future of AI deployment, highlighting how Baseten is positioned to disrupt traditional methods of AI infrastructure. With actionable insights for developers and enterprise leaders, the conversation emphasizes the growing importance of open-source models and the need for efficient deployment strategies.
Stay Connected
- Host Twitter: [Craig Smith Twitter](https://twitter.com/craigss)
- Podcast Twitter: [Eye on A.I. Twitter](https://twitter.com/EyeOn_AI)
---
This comprehensive note format captures the essence of the podcast episode while organizing key information for easy reference, ensuring listeners can derive value from both the discussions and insights shared.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00You're a company that is realizing that AI is a big part of your strategy, and you've been using OpenAI or Anthropic to get started. Now, all of a sudden, you see a lot of adoption. And you're like, wow, this is cool. Everyone's using AI. And based on that, you say, okay, I really need to scale this up. As you scale, you start to hit some limits. So one, you want to run this side of your VPC, which is very expensive. Two, you're now beholden to OpenAI. You don't have much transparency in terms of where your data is going, what does this model look like, and so on and so forth. What DeepSeek provides an open source model to be able to do that.
0:30And so you can run this model faster, cheaper, better with a bit more control and all of a sudden become someone irresponsible not to consider something like that. Hi, my name is Tuhan. I'm the founder and CEO of Base10. You know, my background is in machine learning engineering for the past 15 years. You know, prior to that, I studied electrical engineering at USC. I think, you know, I've been really focused on bringing machine learning products into production and, you know, making sure that models get used in good ways. And that kind of led to the genesis of Base 10. Base 10 is a machine learning infrastructure company.
1:11We're about five years old, based in San Francisco and New York. and we focus on the ability to run, giving customers the ability to run their models in production performantly, reliably, and securely. And what that really means is that, you know, you come with your model or the model you want to use. We kind of give you all the tools to be able to run that as fast as possible, as scalably as possible, and kind of come with all the batteries included so you don't have to deal with issues. So, you know, it's a combination of workflows, to accomplish your compute and then a bunch of optimizations to allow that you know you can serve your customers well yeah and and when you say ai infrastructure um what exactly is is base 10 uh a cloud is it a um a private cloud or or yeah yeah 100 so you know i i would divide i would divide infrastructure up into a couple of different problems so there's like the training problem and the inference problem um they're the the two broad categories and look there's a bunch of subcategories alongside those we're wholly focused on infrastructure on the inference piece which is running those models yeah what do you what do you need to run models so you kind of need like two or three things right so you need the ability to acquire compute um because these models require gpus um and so we allow you to either come to us and we will provide a compute or you can run it on your own compute and that brings up to the second piece which is you know we give you a lot of software to manage the model the running of those models on that compute so you have kind of like compute on one side apis that run models on the other side the thing that maps those two things that's what base 10 is and then you know when you move to production when you move to production and when you work to move to real workflows, you have all these needs.
3:11So think about, you know, when you run a web service, you know, you obviously have, you know, where does the where's the web service run, you have your service, and then you have all the observe observability and CICD and alerting and whatnot that sits on top of that and basically provides all that kind of glue as well. If that makes sense. Yeah. And it's a software layer that then orchestrates compute resources from whatever cloud it's connecting to. And when you say the tools to, you know, manage and operate your model, what does that mean? What kind of tools? Yeah. So all sorts of things. So like, you know, to start with, it's like, hey, you know, how do you manage different versions of this model?
4:06So, you know, we give you, you know, version management. Then there's like the CICD part, which is like, hey, how do you roll out models? Then you have, so that's one piece, which is like, hey, you know, A-B testing, production rollout, all those things, blue-green deployments. We give you access to tooling for that. But then there's like the observability piece, which is like, hey, my model's running. How well is it running? and so we have a whole observability suite that sits within base 10 and then you know obviously the management of compute as well so one thing we provide is the ability to run your model across multiple different clouds so that's one one workload that's sitting across a bunch of different clouds and you know it's like how do you orchestrate that and we give you kind of software layer to say hey i want this model to run in this place and overflow to that place and this the amount of replicas i want running here um and how you know what are the scaling settings for that so we give you an abstraction on top of that as well and then you know we go all the way down to the performance level which is like so i've talked a bit about the software and the workflow piece i've talked a bit about the scaling piece then there's a real performance piece which is how well does a model run on the individual gpu and then we give you a bunch of software to optimize that itself yeah uh and uh why would someone use base 10 as opposed to whatever else they would use what would they use without base 10 um yeah so you know it's funny like you know we spend a lot of time you know selling up market and like talking to customers who are either in the enterprise or we're selling to the enterprise and what we find is that actually for that set of requirements there's not that many tools that a customer has to at their disposal like you know this market the really cool thing about ai is that everything is so new the really the really bad thing about ai is that everything is so new and so what that means is that the the set of tools required to run models in production um is pretty underdeveloped so most of our customers just end up building them themselves.
6:11So what that means is usually going to their cloud of choice, so whether that's AWS or GCP or Azure, and starting to piece together all the different services that sit on top of that to construct something like Base 10. And honestly, that's where we see that we have the biggest value add, where if you are now sitting inside AWS or GCP and trying to connect Vertex to the observability suite to the compute suite, all of a sudden you're fighting these internal organizations at GCP without really knowing it or at AWS without really knowing it. And what Base 10 does is kind of, you know, packages all those things together and provides one comprehensive end-to-end solution.
6:56And that's really where we see our value. So I'd say most of our customers that we talk to are trying to cobble it together themselves using a vast number of resources at their disposal. Yeah. And the interface for Base 10, is it all the tools are together in one place? And yeah. It's one consolidating experience, yeah. Yeah. And then do you actually, could you share the screen and just walk us through? Is that okay? That is okay. I'm not ready for that, but I can do that. Okay, at another time then. And so for someone to use Base 10, walk us through a use case. Yeah. So why don't we talk about something that's very top of mind, DeepSeek.
7:56Okay, so DeepSeek is an open source model that was released. DeepSeek is an organization that released a bunch of open source models. you know, at the end of December and a couple of weeks ago, respectively, a base model could be three and a reasoning model called R1. Now, you're a company that is, you know, realizing that AI is a big part of your strategy and you've been using OpenAI or Anthropic to get started. Now, you know, all of a sudden you see a lot of adoption and you're like, wow, this is cool everyone's using ai and um based on that you say okay i really need to scale this up as you scale you start to hit some limits so one you want to run this side your vpc which is very expensive um two you know you're now beholden to open ai you don't have much transparency in terms of you know where your data is going um where your data is going what does this model look like um and so on and so forth what um deep seek provides the open source model to be able to do that.
8:56And so you can run this model faster, cheaper, better with a bit more control and all of a sudden becomes someone irresponsible not to consider something like that. So what BaseStand does, okay, now go to DeepSeek. Okay, I want to run DeepSeek inside my VPC. This is a 671 billion parameter model, which is very big. So for a number of reasons, this is very hard to run. It requires a lot of compute. So you can't just run it on, you can't run it on your laptop to start with. So you need to go and find compute. For something like DeepSeq, you actually need a whole lot of GPUs. So it's not just, it's just not like one H100.
9:40You need either 16 H100s or eight H200s. So all of a sudden you need to, so let's talk about the H100 case because not many people have H200s today. So you have two nodes of H100s that you need to run. So multi-node inference is now a problem where you need to shard this model and its weights across 16 different GPUs. And then you have it running on the GPU. And then as your product scales, you need that to scale up and down. So let's map that to base 10 now. So you come to base 10. It's as simple as writing a simple Python file that pulls down the weights from Hugging Face. And then you write an inference path in the predict path.
10:19it's a Python file, honestly, probably 20 lines of Python code. We have a lot of boilerplate, so you don't have to do this too much either. You come down, you go to your CLI, and you run trust push, trust push. And all of a sudden, this model with all its scaffolding and the compute resources under the hood gets spun up on base 10. What you get with that is a visual interface for being able to manage everything around that. But then you also get an API endpoint. You also get an API endpoint to run that. So all of a sudden you have this model running, but then you're also getting all the scalability.
10:59So now when you start to get traffic, cool, it's all working, but then your traffic doubles. So what do you do? Well, in the old world, you have to figure out how to scale that up. Base 10 automatically scales for you. And so all of a sudden it's running a twice as replica and then scales back down where you don't need that traffic um you know something you hit a hitch you know and you're like what's going to run you you log into base 10 you hit the metrics tab and all of a sudden you have all the metrics on your on your service okay the uh and then the model
11:36resides on on the virtual machines in whatever cloud you're using yeah and and base 10 is is a layer on top of that to to manage the model is that right you're not providing your so we we we can provide compute but they're not our compute you know we sit on top of um all the clouds you might think about so like azure opening azure um aws gcp crusoe oracle so we can provide compute and we have you know we're putting great pricing and all that stuff if necessary we like to think of ourselves as a software layer and we're not compute providers you can get that if you want but you know what we're seeing from most enterprises today is that everyone wants self-hosted and somewhat of it's somewhat of a non-starter so even you know enterprises that are using open ai and anthropic at scale they're using it deployed within their vpcs and so we can sit on top of your vpc yeah uh you know i'm familiar with sage maker at amazon it's some of these uh some of this functionality sounds like sage maker how how um do you do you differentiate or am i completely off base no that's right i think you know say sage maker you know was built for the old world when there were small models there's no small models and no easy to learn in memory but again like stage make like yes like in in the um in the best case scenario SageMaker would give you a lot of this functionality what we actually see from our customers is that SageMaker is that you know you're fighting within fighting with software to make it work and really like you know we the thing we haven't touched on right now is the developer experience that sits on top of this which is that you know we try to make it easier so we've heard from many customers that you know hey it can take me anywhere from 40 to 60 hours to deploy a model and SageMaker, our goal at base 10 is 10 minutes to deploy a model.
13:34And that's how we think about our differentiation, which is for AI teams today, your number one differentiator factor is your time to market and how quickly you can go to market. Your tool shouldn't get in the way. The tool should be an enabler for that. And we're really trying to speed up those iteration cycles between you and your customers. So you can, you know, roll out pretty advanced AI initiatives, um, as opposed to, you know, spending your time fighting with SageMaker, um, or Vortex, um, you know, amongst others. Does that make sense? It's just like, it's like the, um, you know, this is like, I really like the analogy of something like a data dog, um, where look, Amazon and GCP have a bunch of observability suites.
14:25yeah you know there's probably two dozen observability products in each of those data dog is a really great business because what it does is it kind of links all those tools turns into one um provides that end-to-end developer workflow so you don't have to fight with those tools yeah and and base 10 is itself is not open source but you employ a lot of open source tools Yes, that's right. And so like one thing is like, you know, when you're writing code in base 10, you're just writing Python code. There's nothing really proprietary there. And so like our core technology is closed source. But you know, we have we have trust, which is a way to package your models, which is an open source technology that, you know, again, it's an on ramp to base 10.
15:13But if tomorrow you say, hey, I actually want even more control and I don't want to be beholden to base 10, you know, you can pull your models. And we bank on ourselves to keep you with the value we're providing, but we don't want to have you stuck. And honestly, the closed source nature of Base 10 is for no other reason except speed. It just allows us to move faster and move with the market. With open source stuff, we're built on the shoulder of a bunch of open source giants. And we can do a lot of work with open source, but for us, it's more about how do we enable our customers to move as fast as possible.
15:51Yeah. If that makes sense. Yeah. And you're serving. I mean, it's for the deployment primarily of open source, large models. Right. Yeah. 100 % open source and custom models, I would say. So, you know, like, for example, we do work with foundation model companies that have their own model, then we are the serving layer for them. Right. And really like the narrative there or the way I would think about that as a customer is that you know your your core differentiation is in the application layer for the most part or the model like the infrastructure is pretty much undifferentiated um you know you can either go the open ai or anthropic round and build out you know a three or four dozen inference team to run these models um or you can ideally just use someone like us and kind of abstract that problem away allow you to move faster on the the experiences that differentiate you yeah yeah when you say uh abstracted way i mean if if you're if you're using one of the big foundation models uh like open ai uh is is base 10 you are you do you do you create a private instance of that model uh that then you can fine tune and exactly yeah or Or are you just hitting that model through an API and base 10 is operating at the application layer?
17:22Yeah, so it's a bit of both, right? So you can kind of choose, like, you know, there are definitely customers that deploy a model and use this as API. And, like, it's an API management product. But, you know, a lot of our customers go a lot deeper. And, like, you know, you're able to tune the parameters and around which. because one thing you have to remember is that um all customers it's not like a one-size-fits-all solution for all customers if that makes sense like sure the um the core serving technology um is not going to be that differentiated between our customers but what will be differentiated is how could how they configure that serving technology if that makes sense um and so like you know people have different traffic patterns so they need different scaling stuff they they have different quantization methods so they you know they care about like the different pro the different um uh like quality of the model and like some people care about quality because i'm talking about speed and so that's the control we're giving you while abstracting out the call service lay if if you would like that yeah um the um you're also more cost efficient aren't you than
18:38than sage maker or or it's another uh system like that can you talk about the pricing yeah um look there's a couple different ways we do pricing depending on like where you're running the compute um what we find is that you know when you're running on base 10 because we have this elastic compute model where you know um you know what our goal is to go and negotiate compute on on behalf of our large customer base and then amortize that over our customers. It's almost like a collective, we're the union of unions of startups in a way. But what that allows us to do is it allows you to, let's take a customer, you know, typical web traffic patterns, kind of like this, right?
19:22Where it's like, you know, in the day you get a lot of stuff and now it's pretty quiet. The default right now, if you run your own compute with a SageMaker is to just have it on all the time. and what that means is that you know you're the gaps in traffic you're paying for base 10 actually allows you to scale with your traffic so you you end up actually just saving a lot of money with that auto scaling behavior but then we do a lot of performance tuning as well so we can you know we employ stuff like distillation and speculative decoding to be able to run your models faster which again results in more compute better computerization and cost efficiency.
19:59And then in terms of actually how we think about pricing the product, look, we want to be charged. We want to be seen as a software layer, not a compute markup layer. Does that make sense? And so for all intents and purposes, you can think of it that we pass through the value of computer, then we charge you a software layer on top of that. And ideally, the way we think about that is, especially the enterprise, is that we are saving you employing a team of you know anywhere from three to 30 people to do this and we just want to have some um portion of that at our i think it ends up with a lot of cost savings for our customers and it's kind of cost savings on three fronts uh yeah and and this uh it seems that you're in the right place at the right time because open source is exploding right yeah yeah Yeah, 100%.
20:54And so are you seeing a spike in demand as people turn to open source models? I mean, again, not that they can't use proprietary models with Base 10, but it seems geared toward open source deploying open source models. Create an oasis with Thuma, a modern design company that specializes in furniture and home goods. By stripping away everything but the essential, Thuma makes elevated beds with premium materials and intentional details. I'm in the process of reorganizing my house, and I'm giving Thuma a serious look for help in renovating and redesigning. Thuma combines the perfect balance of form, craftsmanship, and functionality.
21:47With over 17 ,000 five-star reviews, the Thuma Bed Collection is proof that simplicity is the truest form of sophistication. Using the technique of Japanese joinery, pieces are crafted from solid wood and precision cut for a silent, stable foundation. With clean lines, subtle curves, and minimalist style, the Thuma bed collection is available in four signature finishes to match any design aesthetic. Headboard upgrades are available for customization as desired. To get$100 toward your first bed purchase, go to Thuma. That's T-H-U-M-A dot C-O slash IonAI. IonAI all run together, E-Y-E-O-N-A-I. So for$100 off your first purchase, go to thuma.co slash IonAI.
22:53That's T-H-U-M-A dot C-O slash IonAI to receive$100 off your first bed purchase. 100%. You know, like I think the narrative, the narrative, the high level narrative has been over the last, you know, let's call the startup time is the chat GPT moment. You know, that feels like a decade ago. It's crazy to think that was two years ago. But, you know, the narrative has been over the last two years that, you know, there's been two prevailing narratives. One has been, look, closed source is the future. Anything behind that's not a frontier model doesn't matter. And, you know, that's largely been pushed by the large labs who will be developing these closed models.
23:46There's been the second other narrative, which is, hey, open source models are incredibly important to the future. These are very powerful models. Having more things out in the open is going to make it a big deal. So we definitely have conditioned our company on the existence of both. But what's happened over the last two years is that there's been a convergence of the quality of open source models with closed source models. And what that means is that as a enterprise, you now have a lot more options in deploying these models and what models to use. And again, we're talking about anywhere from like a 30 to 40 % decreased baseline in costs to at times an 8 to 10x decrease in costs.
24:29And with that in mind, and with the security and privacy concerns where these models run, and they're not being beholden to one provider as your model provider, and what that means is that enterprise have more options. And frankly, it's somewhat irresponsible for you not to be considering how to factor open source models into your strategy. And so, look, the business has grown remarkably. I'd say like 18, if I looked at 18 months ago to now, we're talking about, you know, a 200x increase in revenue and usage over that time period. You know, I'd love to tell you that we are special and, you know, we have, but I'll tell you we're the right place, right time.
25:17And that's really it. But I think especially over the last two weeks or three weeks as this barrier of can open source ever be as good as closed source has been breached. So we've probably heard from, I want to say, three dozen enterprises saying, hey, we need to be briefed on the strategy here going forward. Everyone is trying to figure out what does this mean. And that to me is very exciting. and that that's where that i i can only see that increasing accelerating over the next 12 months yeah and uh you you you're focused on inference um do you offer uh uh other chips than nvidia yeah because i i've had uh uh rodrigo leung of samba nova on and andrew feldman of cerebrus and i haven't had the grok uh q i haven't had him on yet but uh but these guys are blindingly faster than gpus for inference so yeah yeah it's a really fascinating question it's amazing it's a it's a really fascinating market so look firstly we are agnostic we're cloud agnostic we're chip agnostic um that that's the reality you know we we you know we're we're right we have run on tpus we've run on trinium we've you know we've looked at we've looked at amd chips we've spent time on this that being said like that now that's great and like that is the the narrative going forward um any work we have done on anything except in video chips has been somewhat painful somewhat painful and you know we just don't see you know kudos is old technology that has you know be you know i was when i was downloading video game drivers 20 years ago sitting in my in my bedroom my childhood bedroom i was using kuda drivers to get fifa running and so you know the this is all technology that is a lot a lot um easier to work with now if i if i go back to the original premise of base 10 which was like if you're if you're a car if you're a company trying to move as fast as possible um you don't want to be sitting there fiddling with kernels um you know the day before a big release uh and nvidia just provides that and like they're from what we have seen they're hands down better than those other chips or faster and easier to use than those other chips i mentioned um in terms of the kind of the the new chip the new chip companies it's so exciting we're very excited about those and seeing where they go i still would say you know there's a you know i don't know if any of them have proved or shown the unit economics to show that that is actually just you know that is actually sustainable at any major scale um and you know i i think that will change but i think today um they are fast but the abstraction hasn't built been built on them nor do customers have clarity around, you know, what are the unit economics here.
28:28Yeah. Okay. So how does somebody use Base 10? Yeah. And so it's pretty straightforward. You know, you sign up, you can either sign up and go through our self-serve flow or, you know, get in touch about it. We'd love to help. what we believe is that um you know we we really believe in like accelerating time to value for our customers um and so like we almost take like a boutique like approach to working with customers so when we work with enterprises we'll actually give you a team that will help you get up and running really really fast and so what that means is that we'll you we'll talk about the use case you identify the model we'll work with you to optimize the model for your use case and we'll deploy it within base 10 so there's kind of two separate approaches to using basic there's the more self-serve option um and there's the more high touch option and you know we're happy to support folks uh whatever journey um you you want to go down but there's a the the trust is a really great entry point that i talked about earlier it's our open source packaging library that you can use um but you can also just sign up and get started it's a model library that you can deploy dozens of pre-existing models.
29:41It creates an API for you, and you're almost ready to go instantly. Yeah. And on the models that you've deployed or you see people deploying, you look at the hugging face now. There's, I don't know how many there are. Hundreds of thousands, I'd guess. Hundreds of thousands, yeah, of open source models. So is it pretty concentrated on like the top three or four? And what are those right now? I mean, I think like the truth is, is that, you know, just with anything in life, everything just follows a parallel. and what i mean by that is that you know the the top 10 model models account for 90 of the usage and you know for us there's you know those models are the llama the llama family the mistral family of models quen which is from the alibaba deep seek has um deep seek and all those distillations are a big deal that's on the language model side on the on the image model side fluxes to stable diffusion models are very very popular and then on the audio side whisper is still you know for all variants of whisper which is the text um texas speech to text model um is very very uh uh very very popular yeah yeah um for your um growth i mean who is who is your competition is competition uh and how do you see your your company growing yeah i mean i mean look like i'll be honest with you like i i i am you know somewhat convinced that this is like one of the largest value creation opportunities of our lifetime these models have to run somewhere um to me like basehead can be one of the most important companies um in this space um the scope of the opportunity also means that you know i think other people see us as competition um you know our core differentiation from those customers is kind of across the the three pillars that i described both pillars that i described so um and i'll go over again which is infrastructure like i think we have some of the best infrastructure that's been built here that's scalable runs wherever um you know you name it like any cloud any chip uh scales scales kind of forever there's a massive differentiation there and we've built a lot of software there and interesting things there the second piece is performance which is like you know we think performance is necessary but not sufficient to win these models have to run fast we provide you with a swing of tools to run them as fast as possible the third one is developer experience you know we we think you know we differentiate pretty aggressively with most of our customers in terms of ease of use and the last piece which i you know cannot understand is our forward deploy engineering team which is like we we want to work with our customers we enjoy working with our customers and we don't think of it as like a cost of goods sold we think of it as like a value creation opportunity and that is a massive differentiation so like you know again when you when you go and sign up with one of our customers and when you are competitors and when you go and sign up with us i think the thing you hear time and time again is you know these people felt like an extension of my team uh as opposed to like you know customer support yeah and and when somebody deploys a model using base 10 then then you um yeah i mean you help them deploy but then are are you there for the long haul with you said that you know people can uh port their models uh elsewhere but but yeah talk about that 100 like you know we we have like a sub 10 minute response time to any time a customer needs anything you know we we you know we we want to scale with customers forever one of the big reasons for you know the self-hosted offering we have is that we believe that it helps align incentives between us and our customers like i don't i think if we were if you had to run a base 10 cloud and the markup was all in the base 10 cloud um i think that'd be pretty that's pretty you know i i it'd be hard for me to fathom at scale as a buyer and so we want you know started like our our journey hopefully looks like you start on the base 10 cloud you know we we sit across a bunch of other clouds you don't have to think about it you hit some substantial scale we're happy to port it over to your cloud and change the nature of the deal so we can be supportive for the long haul.
34:29We're pretty proud that in the five years of our existence, we basically have never lost customers. Obviously, it's not zero, but it's pretty close to zero in customers who have left us yeah and and it's not only uh cloud you can um this can work on premise yeah yeah and do you see what is everything primarily on cloud right now or is there i mean as people get their hands on these open source models uh in their certain industries where things have to be on-prem. Do you see that as a trend? 100%. I think there's many trends here, which is like on your cloud, on our cloud, on your cloud, on your VPC, on your FedRAMP VPC, and then in our data centers where we're, like for our customers in their data centers.
35:28And I think, you know, this is like, I think this is like one of like the most interesting trends, and I'm sure, Craig, you you have thoughts on this which is like the reverse migration from you know it was kind of like everything on pram to everything in in our cloud to hey how about your cloud um and i think this is the future of software the byoc approach yeah yeah bring your own cloud yeah interesting the uh uh and and how how are you scaling i mean can you talk about uh the number of uh models deployed or the number of customers how it's growing it just seems again that you're in a in a unique uh position given what's happening in the industry yeah um you know i think there'd be other people would be very upset at me if i disclose exactly i was like tell you know we have we have thousands of customers and you know tens of thousands of models deployed and you know the again like the growth is the growth is astounding to some degree you know like again like we're we are you know this market i i think what's amazing is like we've grown so much over the last 18 months but what's very clear to me is that over the next five to ten years um the like i don't think we've scratched we're just scratching the surface right now of what's possible you know most enterprises have not come online yet i think the opportunity in front of us is a lot more exciting than where we've got problem if that makes sense yeah uh and beyond And inference, people are now focused on building agents.
37:06Yeah. How does that work with East End? Yeah, well, you know, it's very funny. I think there's two new approaches, two new changes in the industry over the last six months, or like 12 months, which actually I think are pretty conducive for us. One is that agentic applications and the other one is reasoning models. And both of those actually just require more inference. So an agent is just a chained number of inference calls. And in a lot of ways, all these performance issues become a lot more important when you get to agentic models, agentic applications, because you're not just making one call.
37:53At times you're making hundreds of calls to the model. and so agentic models drive more inference and i think like the optimization piece gets a lot more important with that and then the second piece is reasoning models which is you know scaling um test type compute and or just inference scaling inference um again more inference and so again like these are great macro trends for us but i think they're also like very hard things to deal with as a customer and i think you know we really do want to help people through these through this journey if that makes sense yeah well let's talk about the the agents first does does the application the agent application does that sit outside of base 10 and then and then the it's it's calling or going through base 10 exactly to get the inference yeah exactly the application is our customers you know we just make it easy to do the inference itself so like all that like look like we have a product called change which allows you to build a lot of these compound ai systems which you know we mentioned become ai agents but i think higher level than that you know we don't want to get intertwined with our customers application logic that is theirs to us like those applications those agents have a bunch of inference calls in them and you know we want to be that delegated inference for them yeah yeah uh and uh just sort of backing up a little bit where where do you see this going i mean you know these reasoning models are getting increasingly powerful the the agents are becoming increasingly powerful people are now talking about society of agents or yeah only humans only interact with at one end or the other um i mean how just can you talk generally about how you see ai developing yeah going forward yeah i think i think that's right like if i'm being really honest like uh again craig you you're you're in this as much as i am from the high from like the the macro level which is like i mean every six months i'm sure every prediction you've made over the last three years has been has been proven drastically wrong and i think you know that that's the truth here it's like look we're here riding the wave uh i think what's more important is that what we're seeing is that inference becomes a lot more important reasoning models again these models get a lot stronger uh um and the inference time compute gets a lot more important um but what's more important to be honest is like what i think like separates out it's how this ai hype from past bubbles is like the value is so real you know we're fundamentally redefining like you know when people ask me about is this a bubble probably a short term yes but long term no right it's like you know yes we're probably over like overvaluing a lot of things over the last next few months and um but it does seem like over the next 10 years, everything will fundamentally change.
40:58And AI is going to be a big part of that. And inference is going to be a really big part of that. And whether that's through scaling inference time compute, whether that's through reasoning model, whether that's through agentic applications or compound AI systems, inference has a big part in that. But again, you'd have to be crazy to make a prediction in this market how things mature over six and six and 18 month six to 18 month timeline but i think over five to ten years what we can what i can tell you is that um i'd be very surprised if you know like 10 times or 100 times as many enterprises aren't using ai um in their in their core flows than they are today yeah i want go back to these new chips um and and you were yeah uh nvidia has has a lock because of cuda uh but on the inference side yeah i guess you're you're deploying models but uh cerebris and sambanova and grok they all have their own clouds uh you know pending a time when the hyperscalers sort of install their chips, if that happens.
42:23And for an application that's looking for inference, they can hit those clouds through an API. So they don't have to fiddle with... Autoscaling and what? Yeah, exactly. Right. But so how does that relate to base 10? i mean can it yeah i think it's a really good question so like look like at some point you want more customization and like you know what you don't want in your production application is being throttled by grok's own scaling limits or cerebrus's own scaling limits i think that's when base head becomes really important it's like hey you know you get started with the fastest easiest way to run these models you know and that might be through a opening eye it might be through a grok it might be through a cerebrus um you know you get a lot of tokens per second it'll be fast but all of a sudden you know you when you are running your model and you need to scale that 100x you know you really need a lot more control and options in terms of where that compute lives and a bunch of workflow tools that sit on top of that and that's what that's where i think that transition from off the shelf fast apis to something like base 10 happens when you have control when you need developer workflow it's like imagine imagine giving your your web application to a third party and have it just being a black box you're not going to do that you need observability you need control you don't want to be beholden to someone else and that's really where i see that transition point you know we say that to customers all the time it's like hey sound like you need an api right now and not a base set yeah you're not at that scale yet or you're not at that maturity with something like base 10 makes a lot of sense um let's be design partners and help you through this transition as you get through it um but then that answer your question in terms of like that trade-off like when you need when you need something off the shelf and when you need something that's a bit more custom bespoke and gives you more flexibility and and the knobs to turn what is uh scale to zero functionality yeah scale to zero it goes back to auto scaling which is like kind of what i said to you which is like uh a lot of folks they don't need gpus all the time so we give you the serverless type approach where you can basically use use the gpus when you need to or compute when you need to and when the models aren't being used it just turns off so you don't gain charge now this only works if you have really really fast cold starts which is like when i need it you know it spins up very very quickly and you know we've done a lot of work on that to make that as fast as possible so that if you know at times you know the wall time of a depending on the size model, it can be sub 10 seconds from the time where you're not paying to its consuming resources running at scale.
45:04For people to use Space 10, they just go to the website? Sign up. You can sign up to talk to us. Honestly, if you send me an email at tuin at basetend.co, I'd love to help. And I'm happy to push you towards the right people and get you started. But once you go to our website, it'll be very, very easy to figure out how to get started. and if you fail at that and it's not as easy as i'm i'm selling it to be send me an email and i'll figure you out yeah and it's base10.com is that right yeah i i we have we have base10.com and base10.co um so yeah so you know we we acquired the dot com later than we wanted at that point is there something that i haven't covered that that you think listeners should know no no this this is very comprehensive like you know depending on who you are look like open source models are going to be a big part of the future and we'd love to figure out how to help you use them in your applications and you know that's you know we're happy to help as thought partners as as software providers um and yeah hit us up but otherwise craig this was fantastic It was very comprehensive and thank you for having me.
From the publisher
This episode is sponsored by Thuma.
Thuma is a modern design company that specializes in timeless home essentials that are mindfully made with premium materials and intentional details.
To get $100 towards your first bed purchase, go to http://thuma.co/eyeonai
—————————————————————————————————————————
AI deployment is broken—can it be fixed? In this episode, Tuhin Srivatsa, CEO & Co-Founder of Baseten, reveals how his company is DISRUPTING AI infrastructure, making it easier, faster, and more cost-effective to deploy and scale AI models in production.
As enterprises increasingly turn to open-source AI models and grapple with the high costs and complexity of scaling, Baseten offers a game-changing solution that eliminates bottlenecks and simplifies the process. Discover how Baseten is taking on AWS SageMaker, OpenAI, and cloud-based AI deployment platforms to reshape the future of AI model deployment.
What You’ll Learn in This Episode:
-
Why AI deployment & scaling is one of the biggest challenges in 2025
-
How Baseten enables enterprises to run AI models faster & more efficiently
-
The shift from closed-source to open-source AI models—and why it matters
-
The hidden costs of AI inference & how to optimize for performance
-
Why most AI models fail in production and how to prevent it
-
The future of AI infrastructure: What comes next for scalable AI
Whether you’re a machine learning engineer, AI researcher, startup founder, or enterprise leader, this episode is packed with actionable insights to help you scale AI models without the headaches.
Don’t miss this conversation on the next era of AI deployment!
#AI #ArtificialIntelligence #MachineLearning #Baseten #AIDeployment #AIScaling #Inference #MLInfrastructure #TechPodcast
Stay Updated:
Craig Smith Twitter: https://twitter.com/craigss
Eye on A.I. Twitter: https://twitter.com/EyeOn_AI
—————————————————————————————————————————
(00:00) Tuhin Srivatsa’s Journey in AI & Baseten
(01:50) What is AI Infrastructure & Why It Matters
(03:30) How Baseten Optimizes AI Model Deployment
(05:19) Why Most AI Deployments Fail (And How to Fix It)
(09:17) The Future of Open-Source AI Models in Enterprise
(11:01) How Baseten Automates AI Scaling & Inference
(14:12) Why AI Developers Struggle with Cloud-Based AI Tools
(18:47) The Real Cost of AI Inference (And How to Reduce It)
(20:44) Why AI Scaling is the Biggest Challenge in 2025
(26:55) Can AI Run on Non-NVIDIA Chips? (The Hardware Debate)
(31:23) The Future of AI Model Deployment & Inference
(37:05) How AI Agents & Reasoning Models Are Changing the Game
(40:39) The Truth About AI Hype vs. Reality
(45:04) How to Get Started with Baseten
(45:48) The Future of AI Infrastructure




