In short
Practical AI: Episode Summary - Self-hosting & Scaling Models
Podcast Overview Podcast Title: Practical AI Podcast Description: Engaging discussions on artificial intelligence (AI) focusing on practical implementations and real-world scenarios to make AI accessible to everyone.
Episode Details Episode Title: Self-hosting & Scaling Models Episode Description: Features Tuhin Srivastava from Baseten, discussing trends in model deployment, monitoring, and the impact of generative AI on open access models.
Key Participants
- Tuhin Srivastava - Co-founder of Baseten
- Chris Benson - Tech strategist at Lockheed Martin
- Daniel Whitenack - Founder and CEO of Prediction Guard
Key Themes Discussed
Current Trends in AI and ML
- Open Source Models: The rise of platforms like Hugging Face has made it easier for developers to access and utilize open-source models, akin to GitHub for software development.
- Generative AI: The boom in generative models, particularly through platforms like ChatGPT, has shifted user expectations for performance and speed in AI applications.
Shifts in User Demographics
- Transition from data scientists being the primary users of AI tools to a broader range of engineers and product developers who need to integrate ML into existing applications.
Infrastructure Challenges
- Model Hosting: Running AI models in production requires complex infrastructure, including considerations for latency, throughput, security, and data privacy.
- Evolving User Needs: Users now demand faster, more efficient, and scalable solutions due to the accelerated pace of AI advancements.
Baseten's Role in AI Deployment
- Deployment and Monitoring: Baseten offers a platform for deploying and monitoring models at scale, addressing the complexities involved in running models in production.
- User Experience: The platform aims to simplify the deployment process, allowing developers to focus on application-level concerns rather than underlying infrastructure.
Differences Between Hosting Solutions
- Base10 vs FastAPI: While FastAPI allows for model deployment, Baseten provides a more streamlined workflow with built-in features for versioning, A/B testing, and observability, thus reducing the boilerplate work for developers.
Future Directions for AI Infrastructure
- Multi-cloud Solutions: There is a clear trend towards multi-cloud infrastructures, allowing companies to leverage different cloud providers while maintaining control over their data and models.
- Fine-tuning and Optimization: The discussion also covered the potential for fine-tuning models and the need for better control over the fine-tuning processes as organizations seek to customize models for specific use cases.
Key Takeaways
- Model Deployment Is Complex: Successfully running models in production involves significant infrastructure challenges, including security, scaling, and observability.
- Growing Interest in Self-hosting: Many organizations prefer self-hosting models to maintain control over data privacy and security.
- Demand for Faster Solutions: As AI technology evolves, there is a growing expectation for fast and efficient solutions that enable rapid deployment and iteration.
Conclusion The episode provided valuable insights into the evolving landscape of AI deployment, particularly the shift towards self-hosted models and the infrastructure challenges that accompany them. Baseten's approach aims to simplify these complexities, making it easier for developers to integrate AI into their applications.
Listen to the full episode [here](https://changelog.com/++).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:28Welcome to Practical AI. and database close to your users. No ops required. Learn more at fly.io.
0:43Welcome to another episode of Practical AI. This is Daniel Witenak. I am the founder and CEO at Prediction Guard, and I'm joined as always by my co-host, Chris Benson, who is a tech strategist at Lockheed Martin. How are you doing, Chris? Doing very well today, Daniel. How's it going? Oh, it's going great. I spent the afternoon in sort of a brainstorming session with a couple of our team members here at Prediction Guard. And it was a ton of fun. So talking about a lot of prompt engineering things and how different models perform and that sort of thing. So it was a good time. I'm glad you're doing that because you know what?
1:20I just want things that just work. You know, I don't have to think about it. I'm glad you're thinking about it. I think we might have someone else to talk to who knows how to make things that just work. Yeah, yeah. Well, a lot of the models that we're running sort of just work for us in terms of inference because we're hosting some of our models in Base 10 and we've got Tuan joining us from Base 10 today. How are you doing, Tuan? Hi, Dan. Hi, Chris. Nice to see you guys again. Thanks for the kind words, Dan. Yeah, yeah, for sure. Well, it's exciting to have you actually back on the show because it was, I believe I looked it up, it was like May, June 2021 when we recorded and released the last episode with you.
2:06So how are you doing? What's new and how is Base 10? How's the ride been? Yeah, it's been crazy. I feel like the last, I think it was like May 2021. That's like a millennium in AI time, you know? Oh my God. if it actually does feel like it was a different job ago yes it feels like the job before the last job if that makes sense no i think um being crazy i think the last two years for everyone here have probably been a bit of a whirlwind and you guys are pretty on top of current things and machine learning and ai and i think um i imagine just like for you guys it's hard to keep up at times with you know what's going on i think you know we only do one show a week and you know We almost need a daily show.
2:53There's so much content now. It's not enough. Yeah. Don't give our listeners ideas because I don't know if I can do a daily show. But it is a lot. And I think, so I'm looking back previous at our episode. And last time we talked about sort of the easiest way to create ML apps. That was kind of part of how the conversation was framed. And I know just from working with you and talking with you as friends that a lot has changed and you've seen some things within how people are deploying machine learning AI systems that now Base 10 is really focused on. Could you give us kind of like the high level view of Base 10 and the type of problem, the type of solution that you're offering?
3:41Yeah, yeah, yeah. I think it's just worth probably pointing out, before I go into base 10 specific things, what are the key things that changed since we last talked? I think if you think of the year of 2012 to 2020, data scientists were the ones doing a lot of machine learning. I think that's changed for a number of reasons. I have a lot of thoughts on that. But probably the bigger changes are the emergence of good open source models. And I know you do a lot of work with that. And we've seen Hugging Face as a community evolve into this really, really vibrant place where if you want to get a sense of how fast things are moving, it doesn't take long to take screenshots of Hugging Face every Monday morning and see how the trending is changing.
4:24And you'll see that things are pretty different every week. I don't know if it was Daniel that said this or another friend of mine, but the analogy was that, you know, hugging face has become to AI kind of what GitHub has always been for software developers over the last decade or so. It's just, you know, it's the place to go to find it. Anyway, I didn't mean to interrupt you there. Yeah. And the good about it, but also the confusing, too. It's like, you know, you have like random person clones model or copies model and uploads random version of that model that like maybe works. Yeah, I understand.
5:01I feel like the game you have to play with hugging faces like, but does it run? Does the model run? But does the model run? but you know i think open source emerged and i and i think like stuff like whisper showing up and you know some of these ocr type replacement models showing up you know they're probably the more interesting ones to me not because what they do but because they end up just solving a lot of open problems you know you know if you think about transcription as a problem of think about nuance and like how long they were working for that yeah like literally 20 or 30 years of of work just kind of, all right, that's a solved problem now.
5:38Let's move on. Oh, we've actually solved multi-language with the same model too. Business models come and go, don't they? Yeah. Yeah. It's wild. And I think the last piece is just around the chat GPT moment for AI, interesting for a number of reasons. I personally think it's someone who built infrastructure that it's most interesting because if you want to call that the iPhone moment of AI, I think it's a bit different to that because it's so early in the journey. it's like if the iPhone showed up and we were all using 5110 from Nokia the world would be very very different and I think because consumers and developers their first taste of machine learning and AI was through chat GPT and GPT APIs the stakes are just you know it's just harder to build something good like you know people don't want to use a model that takes 12 seconds to run you know like high speed production inference is you know taken for granted when you're using um this model and then when you kind of combine that with okay open source models need to be run somewhere all right we we personally think that's like okay well there's a massive infrastructure opportunity there um well maybe not even opportunity that just say a fact that a whole new stack will be built to support uh models to be able to power these end user experiences and i think like that's kind of the core insight kind of going into base 10 and talking a bit about base 10 and like what's changed is that you know we kind of two years ago when we were talking about data scientists we weren't talking about engineers i think that's pretty key to our story which is that i think we came to the realization that every engineer needs to grapple with machine learning now as opposed to maybe a smaller market of data scientists i think going from smaller models that run in memory to larger models is another big you know focus change we've had i think there's a bunch of language stuff and nltk stuff and you know you were doing all that work daniel but for the most part everyone was using you know psychic learn and that and psychic learn models for the most part run in memory on cpus on cpus yeah if you think just in that time period the amount of maturation you know that's occurred in this industry it still bought you said something a second ago which i just kind of hit me and that is most people out there in the general public you know are really just getting into this you know with chat gpt and stuff and we've come a long road already.
8:02I'm listening to you and it's amazing how far we've come in such a short time. Yeah, it's insane. A hundred percent. And I think that going from small models to large models as well, it's just like that change kind of happened pretty quickly. I saw one who did a bunch of work with small models. They have their time in place, but they're just not that fun anymore. Something that runs in memory just doesn't give you the same feeling. And then I think the last one is just so much of the stuff that was happening with machine learning outside of fang i'd say was like fang had some production use cases around like ad serving and search and whatnot but outside of that it was mostly just internal workflows you'd go and work on fraud and content moderation and recommendation systems and i think you know going from hey every product is going to have some sort of machine learning in it and every existing product will definitely and 90 of new products will be built with a new pillar that is machine learning and AI, which wasn't the case, I think, two years ago, which is just crazy to think about.
9:03Yeah, it's completely changed in that time. I think one thing that was brought into focus for me while you were talking was that as soon as you make this leap to kind of larger models and you make the leap from some closed API that's very fast to maybe running your own model. There's two things that become like immediately clear. One is the infrastructure challenge around that, which I think is the workflow around that and the model hosting. Base 10, of course, is an expert in that. And the other side of that is like the product sort of concerns around running these things, which I feel like we always have great conversations too.
9:47And because like you're on the one side of that and I'm probably on the other side because yeah, like when you're using chat GPT or even the OpenAI API, they have layers of protections on the prompts and like on the output, they have filters to make sure they're not responding in certain ways. And there's all these product concerns that people don't think about. And then they take like a Llama 2 model or something, they run it and then there's like, oh, this doesn't respond, like this is not a product, right? And so like the infrastructure is a piece of that. the ability to iterate very quickly with models of a variety of types, I think is part of that infrastructure challenge.
10:26How do you see that infrastructure piece of this kind of playing out? Yeah, maybe just before I go into that, when you were saying that, it reminded me of like, I bought a drone in 2014 and I was so excited. It was a DJI Phantom 3. I was so pumped and, you know, I flew it around a bit. And then I basically had an autopilot mode where you didn't have to do anything. It kind of just stabilizes itself. and then there was this button that said manual mode and it had like all sorts of warnings and i remember saying like being like oh how hard can it be right oh i really like joking like i was in a safe space but i put it into manual mode while i was up in the air and it just fell to the ground very very fast i don't know chris do you have some experience with this well i'm i'm in aerospace uh professionally so i know a little bit about that and yes chris weren't you like the tv host of the drone racing competition or something like this?
11:21Yeah, a few years ago, I was one of the hosts of the first drone racing league. They had a championship series and they were using, instead of you as a human, they were navigating obstacle courses and stuff like that. And what I learned through that experience is when you have how much autonomy is required to make even small things fly well. yeah and yeah so i i sympathize with you for going on manual there oh no don't do it don't do it and the analogy i think holds here it's like the closed the closed api it seems so great and then you you see something like a llama 2 or a mistrawl and you're like okay i'll just rip and replace this and it's like nope that's not gonna work um for a number of reasons and i think that's kind of how we see about it like think about infrastructure-based sandwich is that running models in production is very very difficult it's difficult for a number of reasons so i can decide but like from like a user requirements perspective like latency and throughput paramount costs are something you want to optimize data privacy is you know a whole another beast security comes into play then orchestrating this across a bunch of different hardware orchestrating this across clouds becomes a problem benchmarking these things isn't easy and that's even before you go into all the evals and the you know the kind of like the guardrails you want to put around this thing to get it running and i know some of the stuff that you think a lot about dan so just as an extension of what you're saying and daniel you mentioned it at first but if you could also talk a little bit about what the difference is about just hosting like like just having the model hosted and kind of the idea around what you have to put around it as to make it a product because i don't think most people talk about i don't hear a lot of conversations about that and it's a big set of gotchas on what to do you know and kind of what's involved in that what's your thinking around that when people are looking at doing that?
13:10I can talk well about the first one and I have thoughts about the second. Dan's going to be the expert around the latter piece of that. But to get a model running in production, there's actually a ton of work you need to do from a infrastructure perspective. And this is before we talk about the workflow stuff. You need to figure out how to containerize this thing and get this image running. As we alluded to earlier, just taking a model of HikingFace and expecting it to run is not a thing. There's a bunch of requirements that these models have. there's quantization inside the code there is um different base images that they might need based on torch and pytorch and python versions and then you know you can really find yourself in a bit of a pickle just trying to dockerize a model so the first thing is you need to figure out how to get this in some sort of containerized form so you can run it elsewhere once you have that the truth is that you need to spit up some sort of service that can deal with variable traffic and the reason why that is is that traffic you know these things tend to be expensive they tend to be bound by compute so if you get smashed with a bunch of requests it's not like you can just have one model it will queue out there will time out the whole thing will slow down your product won't work so you need to figure out how to scale up and down with traffic and then you need to figure out all the security concerns that come with all that that's just on the serving layer now you need to start thinking about the workflow layer that sits on top of that and i I think version management is non-existent for them.
14:31If they're hooking up into CICD, to really treat this like a service or a microservice that has putting your model behind an API, you need observability and logging. Another whole set of features. And what you realize is that taking a model and getting it working in production behind a reliable, secure, performant API and maybe cost-efficient, as someone who's done this myself, as someone who's built a company to try to abstract this way, it is easily for one model, a couple of people have had counts work for a couple of quarters, if you're lucky, at scale. And I think that is the most efficient organizations that can hire people with Kubernetes experience to be able to do this.
15:14So that's the type of things that we try to abstract away from our users, where it's like, you know, you figure out the Python code, we'll figure out everything else. And we'll give this model this first class treatment so that you can version around, that you can log around and that you can observe it and you can call it. But you don't need to think so much about that. You get that. That's now to the point where you have something behind an API and ready to consume. Now, there's a bunch of stuff that needs to happen to make sure that, you know, it doesn't start saying random stuff that you protect against hallucinations.
15:46It's not just ingesting PII all the time. Dan can probably talk really quickly about that as well. I'm sure. Yeah, I think part of the reason why I'm always excited to talk to Tuen and his team at Base10 is because they are experts in this. All of those layers that we just talked about, I was actually on a call with someone the other day and we were talking about spinning up some microservices or something. And I think my comment was, I just really don't want to care about Kubernetes because I don't want to wake up lying in a ditch crying in the fetal position. Like that's how I view that like whole world.
16:25So props to you and your team for dealing with that side of things. I think that's what's allowed us then on like the prediction guard side in a lot of ways to like bring up a model quickly and then have the time to think about some of these other things too. And I don't know if you can comment on, like, I have my own perspective from trying to run models for my company. But it would be interesting to hear the perspective of different personas that are coming into Base 10. Like, are they people that are sort of application developers that are, you know, not infrastructure people? Are they like data scientists?
17:13Like, what are the types of people that are coming to Base 10? And maybe along with that, as you mentioned, closed APIs are getting used a lot, but still people are coming over to think about hosting their own models. One question would be, why? Who are these people and why? Totally. I'll answer the second one first. No, I think I'll add them together. is that it is more and more just engineers. I'd say like, I don't know if there's any distinction now between like, I'd say it's less and less data scientists, your traditional data scientists. It's more and more people with some ML exposure, product engineers, infrastructure engineers who have tried to build it themselves and have really felt the pain.
17:58I think from a product engineering perspective, like why people want to use open source APIs, I think cost is one big thing, is that OpenAI has to stack up. over time i think or anthropic um i think the other one is data privacy and security is that you don't want to just be piping over all your data to open ai today um and especially when you start to talk about b2b use cases and enterprises and i think there's like probably the more interesting one is that there are just like a long tail of people working on weird models people are fine tuning models fine tuning open ai models is you know not great um you get a bit more control with that manual mode with open source models.
18:38And so it's kind of like the long tail of use cases, I'd say, are coming more and more. And these can be engineers, they can be machine learning engineers. They can be honestly like a lot of audio models, like different modalities that there's not that much exposure to with closed APIs and a lot of custom stuff as well. You mentioned, you know, like shipping data over to OpenAI. And I have talk to gazillions of people who have that as a constraint in their businesses, you know, because the attorneys for the business are like, nope, you don't want to send, you know, your proprietary information and stuff over that.
19:11I guess you would not have that issue at all with base 10, would you? I mean, that kind of goes away altogether when you're hosting in that way, right? Yeah. You have ownership of your data. We don't log any of that data. You're treating the model is just like a map of input, the output, and nothing else. Yeah, that would really solve a lot of people's problems by taking an approach like that. Yeah, and I think the second piece there is that once you adopt the base-in approach, you can then start to think about self-hosting, deploying it within your own VPC. It's like, you know, we have customers that deploy base 10 within their own AWS account, and data never leaves kind of their boundaries or their accepted boundaries.
19:49Yeah. And we've kind of, I think you've framed the concerns that you're looking at with base 10 very well, these sort of infrastructure scaling concerns of hosting your own model. Could you maybe take a step back and just describe like, if I go into base 10, like how have you architected the approach like to, I'm an application developer, I want to run, you know, some random fine tune of Llama 2 that I've created somehow. What is it like for me? What does that look like with the way that you've structured this? And what's some of the thinking behind that in terms of the workflow and how you want it to be for people so that they can treat, I guess, that model as a first class thing that is a first class asset in terms of what they're monitoring and logging, that sort of thing?
20:42Yeah, for sure. I'll go with just to make it and try to take away as much of the complexity, but maybe more importantly, it's still that you'd have a bit of control. I think base 10 data is like a one line to deploy your models. We don't believe that anymore. Either we think that actually having a little bit of structure around it gives you a bit of structure up front, gives you a lot more flexibility a bit down the line. So we have an open source library called Trust, which is basically it's abstraction that if you write your model in, you get kind of everything free. And so basically you need to write two things.
21:19You need to write one Python class with a load function and a predict function. And this is vanilla Python code. It can sit within your monorepo. You can specify requirements as you want. There's nothing base 10 about these files. You know, you could run them outside of base 10. I think that's very important. But once you write that load and predict function, it does two things. One, it tells us, hey, what are you trying to do here? And, you know, we can load that up. And when we deploy a model, we load that function. When we infer, we run the predict function. But more importantly, within those functions, we allow you to compile stuff down.
21:52We allow you to kind of do the tricks that you need to do so that you still have that control. And within that, you can write preprocessing and post-processing functions that allow you to maybe strip out some data, log something, monitor something. But really, it's still giving you that control at the product and application level while still abstracting out the thing we want, with trust once you have a trust developed and you can go to trust.basedend.co and check out a bunch of these um it's a pretty simple abstraction and you can just push that up and we kind of give you all the work on version management around that deployed trust yeah could you speak then to like that's like the prep kind of that goes into oh i've got my weird model i'm writing this python class.
22:38I'm going to deploy it on somewhere. And I know that like one thing that I think is really cool how you've made base 10 trust is like you mentioned, it is open source. And so you can run trust things and deploy in a variety of ways. One of those being like base 10's hosted infrastructure, which is of course easy, but it's also like generally a great, great sort of framework to package your models. But let's say that you do kind of go the base 10 route, you deployed this through the base 10 client to base 10. Could you kind of compare and contrast, like, let's say I just tried to run my model in a fast API API in a EC2 instance or ECS or like whatever that is in my cloud, what is going to be different about what I look at when I kind of look at my model in base 10 versus running this API somewhere else?
23:37And how does that make a meaningful difference? Or what are you trying to do in terms of making a meaningful difference for the day-to-day for people? Yeah. Well, I think what you're doing is that, so besides, you know, you can run that model in fast API. Great. You got this model, you give it an input, it gives you an output fantastic let's carry on but it's like the depth of features and the creation of workflow which are really important here and so like the depth of features is that hey if you can do that with fast api great you're still gonna have to set up auto scaling you're still gonna have to set up observability you're gonna have to set up logging and whatnot hardware management but i think the workflow is probably more important to be honest which is like you we're creating a defined way for you to publish new versions of this if you want to a b test two models you can have two models running at the same time you know it's really that the removal of boilerplate and the addition of some workflow so that you know when you were deploying this in production and you need to roll back a version you don't need to go and scramble to find that fast api file that you're using before and we've all been there before and so it's that creation of workflow is probably i think what a lot of our customers probably use us for besides you know i think The production grade inference is a given, but a lot of the differentiation comes from that one.
24:53Just to totally boil it down, you're saving them a lot of work right there. Yeah, 100%. There's a lot of grunge work. It reminds me of Dan liking to do his data massaging that I'm always teasing him about. Joking aside, you're basically saving us all sorts of work so we can get into production faster, get it up and running, and know that it's production grade all the way through. with a minimal amount of effort and know that it's just there. A hundred percent. We're working with this customer right now. This is a pretty late stage startup where it's hundreds of millions of dollars, AI native product.
25:30They've got a team of four AI info people to manage this. And they've been working on this for about two years. We were able to replicate and get a more performant API up and running in two days. Wow. I think that is kind of what we are trying to, is the performance, the workflow, is the maintainability, but it's also just the speed to prod. I don't know how many of your users do this, but the fact that this might be sort of revealing about me as a person and also reveal some utility of base 10, but like I can literally like last night I'm sitting on my couch and I can log into base 10 on my phone.
26:09And that just changed the auto scaling from like two replicas to like five max replicas and the timeout and like all those things of like the auto scaling of my llama to fine tune from my couch like in between halo games so that was that was terrifying i don't know what that reveals about me as a person but certainly like that ease of use i think is really interesting it it's like that uh that proverb when you like try to solve a problem and then you're like i'll solve this with regex and then you just have another problem to solve it's like kind of like that it's like you want to deploy your model and then you want to deploy it with kubernetes and then you have like a whole nother problem to to solve and like solve the auto scaling stuff and and all that and then like i think on top of that it's just all the sre work that you have to do for that service like you know what happens when it goes down you know what happens when you need to migrate something over what happens when there's a new gpu you want to use there's just so much i feel like we've really turned the corners and i think stuff like ai and ml um that really has helped here because people want to move fast it's like i feel like we put the build versus buy debate a little bit to rest for a bit where we just don't hear it as much it's like hey we we want to build it ourselves Like people are just like, we want managed solutions.
27:35Yeah. You got to go fast these days because if you don't, somebody else is going to get there first and you're not going to have a business. And the market is remarkably talent constrained. Like, you know, again, like Dan, you're saying this and this makes me happy because like, just like, so you know, like Dan's background is in data platform and dealing with all these things, you know, like it is. I've cried myself to sleep. In fetal position. Yeah. and so really it's just a ease of use and like it's ease of use and the ability to scale with you like and that's probably like the two things which we try to bring to our customers and I think just even outside the base and it's probably the the biggest opportunity I'd say in machine learning infrastructure right now is think of all the user stories that are important now that weren't important 12 months ago and maybe just take a long a slightly longer term view than And what can I build around OpenAI APIs, which has like stopped a lot of attention.
28:29These are all places where BaseSand is thinking about going. Or like, you know, we will partner with people who are doing. If you think about people at the emails layer, think about people at the fine-tuning layer, you think of people at the training layer, at the observability layer, the logging layer. It's like there's an entire new stack here to be built. And that's a massive opportunity. And Zake's a really interesting company to look at, to be honest, because they were training their models. They were helping people train them. They were remarkably early. But like the value is very, very clear.
28:55to a pretty sophisticated buyer in Databricks. And so to any folks building tools around here, like so much value to be added. It's a green field. I have another question to ask you because you are living in kind of that world of model deployments and the various ways that people are doing this. People are fine tuning their own models or just using open models. I'm wondering how you see the trends going. So you've already talked a little bit about open models being available, small to big models and how people are hosting them. There's tons of people that are also exploring this area around running models kind of at the edge or in various environments or on laptops.
29:43And also people that are exploring kind of, you already mentioned quantization and running models on CPUs, potentially instead of GPUs. I'm curious for as like someone who hosts like a lot of models in the world like what are you seeing in terms of this trend like because you hear about these topics but I don't really have a good sense of like obviously people are exploring those things but are those people just like extra loud on the internet or like certainly there's use cases Chris knows them well for running certain things at the edge. But for many people out there that are like maybe building a SaaS platform or something, you know, that's less relevant, although maybe the costing, you mentioned cost optimization as well around like that sort of thing.
30:34So yeah, how are you seeing as someone, I guess my question is, is someone who hosts a lot of models for a lot of people? What are you seeing as people's concerns, both in terms of that costing, optimizing models, and like deployment targets, I guess. Totally. I think what we're seeing is it's remarkably early, to be honest. Like there are opportunities, like people are deploying stuff on edge, but I think there's not enough of them today. Just kind of think about like a generalized opportunity there. And so I think, you know, there's a company called OctoML that started with edge deployment and then things that were just like, let's move this out because that's where the opportunity today is.
31:14I think all the stuff that's happening around running these models on less and less hardware or optimizing them in x-way or y-way. It's remarkably intriguing, right? It's pretty crazy that we can get a model that we can barely run on, you know, the biggest GPU we can find and start posting it out how to compile it down or rerun it with C++ kernels and all of a sudden it runs anywhere. That's pretty fantastic. But I do think that, you know, we're still in the research phase there. We're in the experimental research phase and, like, you know, we've seen a lot of people deploy those models. Yes, we've seen a lot of people play with those models.
31:49We've seen a lot of interest in those models. But I can't really think of too many examples of people running those models in production just yet. But it seems inevitable that over time. Oh, it will happen. That is the arc that we are on. Yeah. These models are getting smaller and smaller. Yeah. When I'm not on the podcast, I'm in a world where it's all about things moving around in time and space. And a lot of those things will have AI capability on board going forward. So that's, I agree with you, there's a lot of research going on and no one, it's not a solved problem. And when, you know, there's not a set of best practices yet, if you will, but it's an area that is majorly ripe.
32:25I'll be really curious to see if you guys or another company is able to leverage all the expertise you've built up and experience you've built up in the cloud and kind of move out into those areas. Bring your own device, yeah. Yeah, there's millions of devices out there just waiting for you. Yeah. I'd say the challenge there is just around how I'm guessing Lockheed's devices are a bit of a snowflake and you can't build for one type of device and just go and apply that to the next device there. I feel like there's probably some generalization that needs to happen at the OS layer before we can do that.
Read the full transcript
33:02But I am also completely uneducated on Edge stuff. So you probably have a lot more to say that than I do. Well, and it's interesting, too, I sort of ask this in a leading way, because one of the people that we're talking to kind of genericize this, but they run some equipment at the edge in manufacturing, and they have like a hub at the edge, which is air gapped, which doesn't talk to the internet, but their whole, like next generation of things is going to be internet connected. And when I was talking to them about like doing some things with large language models in that environment, you know, essentially where we got was, hey, well, it's going to be more hassle for you to figure out some of these like model optimization things and all of that than to just like set up an API in base 10 or something like that and and just connect it out if that's where you're headed anyway.
33:57and it's not like a military security concern like one of these situations. So yeah, I think we're probably see both and, but for a lot of people, it's like kind of how I've categorized these things in my mind is, yeah, some people will want to run Kubernetes in their own infrastructure and they have the expertise to do that. And if that's you, then like, great. You're one of maybe a few people on the planet. I don't know. Good on you. Good on you. And similarly, like with Chris, if you're running a lot of models on the edge, which I know certain people are and in certain industries, it's really important.
34:34That's great. And that expertise will be there. But I think for the, I don't know, my sense is that for a lot of majority of people, like separating out that infrastructure concern of model hosting is really, really a useful way to think about things. So I don't know, we'll see. I'm always bad at predicting the future. but we talked to this one customer it was saying that for some reason the ceo bought a bunch of gpus and they literally have machines in the office they're like oh well we're gonna do the i think amazon has like the kubernetes anywhere or something where you can basically uh like a hybrid sort of thing yeah yeah i've you know we ran away from that opportunity as suffice to say But I think these people are thinking about these problems.
35:20I don't know if there is a solution here just yet. Yeah. Well, as you kind of think about, so obviously, and I know just from our discussions, like you're helping a lot of people and really significant use cases in the space with already with what you're doing with the infrastructure side of model hosting. But as you look to kind of the next, I can't even say like the next years because things move so quickly. But as you look towards the future and like what is not yet solved on the infrastructure side with model hosting and what are like you and base 10 really excited to dig into, what comes to mind and what are you thinking about?
36:05Yeah, I think within the containers, like that layer gets very interesting. I think, you know, VLM, which I'm sure you've played around with, TGI, which, you know, they're both great. I think they're still far from ready for prime time just yet. VLM, TGI, and CRCLM that NVIDIA just put out. I think there's going to be more and more of these frameworks and supporting these frameworks, I think, is going to be very, very key for us and what we're really excited about. So we're going deeper at that layer. So you can kind of bring your own framework on your container and really benefit from that. So we're going to have first-class support for TRT-LLM pretty soon.
36:45And we already do it for TTI and VLM. And I think that side's pretty interesting. I think we have a big launch coming up, I'm happy to talk about it right now, actually, but around multi-cluster. So that's basically being able to, one, use your own compute to bring your compute to base 10. so the the control plane sits on base 10 and the workload plane considered gcp azure aws or some combination of the three and then we'll keep adding clouds to that and so i think that's very very exciting especially in the enterprise that because you know people want self-hosted that's huge it is it's really nice that's gonna be big and then kind of just beyond serving and i think with really side that you know we've already had one foray and we learned a lot and we end up retiring, but we're going to get into fine tuning at some point.
37:34I think that's, we keep seeing, just like I said about the edge device stuff and the compilation stuff is that fine tuning is still an art. I'm a little bearish on APIs that say, give me your data, I'll give you a model. I think you need more controls. As someone who built an API that said, give us your data, we'll give you a model with Blueprint. I think you need more control. You need control over your base models. You need more control over even the fine-tuning scripts to customize that model. We'll start to think about that very soon. We're already doing a bunch of work with customers to make sure that we're marching in the right direction.
38:09So I'm very excited about that, which is that over time, Base 10 becomes this place where you can run your models great. But then you can also start to collect data sets around your model. Just imagine if you could just give your model to Base 10, and then you either opt in, and we basically write all your input and output data to S3. that's beautiful yeah like essentially a level of uh caching for model inputs outputs yeah exactly yeah the multi-cloud thing really will be big for enterprise by the way just to that point because i think most enterprises across many industries are recognizing that their future is a multi-cloud world it's no longer tied to one and if you have base 10 able to do that hosting and run a control plane and you can deploy into any of the cloud clusters that you happen to be and maybe different parts of the company emphasize one or the other, then that takes a lot of challenge that they're currently facing out of that.
39:02So it's pretty sweet. I mean, also it opens up opportunities, right? Especially in the GPU-controlled world. You can get them from wherever you want. And then I think once you have data sets as well, then fine-tuning just becomes obvious, which is like, okay, now can I fine-tune this? Or maybe it's even like, hey, hook up your OpenAI endpoint using Base 10 so we can collect that data set. And then we can create that fine-tuned misceral or llama to to keep you the more you want so i think there's a lot of interesting things along that whole stuff and as i said to you guys earlier there's so much opportunity here for people building in the tooling layer brand ai and ml it's very exciting yeah well we appreciate uh you taking time out of doing great work in that layer to talk to us and share with our listeners This is a great conversation and hopefully we have you on the show in less than three years from now.
39:55But if not, at least three years from now, hopefully sooner. So thank you for joining us again and giving us an update and some insights around this. And really appreciate what you all are doing and appreciate you taking time. Of course. Thank you. Thank you so much for spending time on the show. I thought it was fun to be on it.
40:22Thank you for listening to Practical AI. Your next step is to subscribe now, if you haven't already. And if you're a longtime listener of the show, help us reach more people by sharing Practical AI with your friends and colleagues. Thanks once again to Fastly and Fly for partnering with us to bring you all Change Talk podcasts. Check out what they're up to at Fastly.com and Fly.io. and to our Beat Freakin' Residence, Breakmaster Cylinder, for continuously cranking out the best beats in the biz. That's all for now. We'll talk to you again next time.
From the publisher
We’re excited to have Tuhin join us on the show once again to talk about self-hosting open access models. Tuhin’s company Baseten specializes in model deployment and monitoring at any scale, and it was a privilege to talk with him about the trends he is seeing in both tooling and usage of open access models. We were able to touch on the common use cases for integrating self-hosted models and how the boom in generative AI has influenced that ecosystem.
Changelog++ members save 1 minute on this episode because they made the ads disappear. Join today!
Sponsors:
- Fastly – Our bandwidth partner. Fastly powers fast, secure, and scalable digital experiences. Move beyond your content delivery network to their powerful edge cloud platform. Learn more at fastly.com
- Fly.io – The home of Changelog.com — Deploy your apps and databases close to your users. In minutes you can run your Ruby, Go, Node, Deno, Python, or Elixir app (and databases!) all over the world. No ops required. Learn more at fly.io/changelog and check out the speedrun in their docs.
- Typesense – Lightning fast, globally distributed Search-as-a-Service that runs in memory. You literally can’t get any faster!
Featuring:
- Tuhin Srivastava – GitHub, X
- Chris Benson – Website, GitHub, LinkedIn, X
- Daniel Whitenack – Website, GitHub, X
Show Notes:
Something missing or broken? PRs welcome!




