In short
Practical AI Podcast Episode Notes
Episode Title
Generative Models: Exploration to Deployment
Episode Description In this episode, Chris Benson and Daniel Whitenack discuss the lifecycle of generative AI models, from experimentation to deployment. They explore how this lifecycle differs from traditional data science practices, model optimization, and serving.
---
Key Contributors
- Chris Benson: Tech strategist at Lockheed Martin
- Daniel Whitenack: Founder of Prediction Guard
Episode Highlights
Introduction
- The hosts introduce the episode, mentioning their recent experiences at various tech conferences, including the Intel Innovation Conference and GopherCon.
AI Hardware Trends
- AI-Enabled Applications: Discussion on building desktop applications running AI models locally, reducing the need for network calls.
- Confidential Computing: Importance of trusted execution environments (TEEs) in safeguarding AI workloads and ensuring data integrity.
Cloud and Infrastructure Developments
- Cloudflare's Workers AI: Introduction of serverless GPU solutions, exemplifying the growing trend of serverless computing in AI.
- Intel's Cloud Innovations: Introduction of the Intel developer cloud, enabling users to spin up virtual machines or bare metal instances with the latest processors.
---
Generative Model Lifecycle
Experimentation
- The lifecycle for generative models is distinct from previous data science practices; these models are often not trained from scratch.
- The hosts discuss the importance of understanding model behavior before deploying.
Model Selection
- Where to Find Models: The best platform for discovering models is Hugging Face, which hosts a substantial library of models (around 345,000).
- Evaluation Criteria: Check download counts, user engagement (likes), and model licensing (commercial vs. non-commercial).
Running Inferences
- The hosts emphasize the importance of testing models in a controlled environment, such as Google Colab, to evaluate resource consumption and performance.
- They suggest starting with smaller models and incrementally testing larger models based on application needs.
Model Optimization
- Why Optimize?: Reduce hardware requirements or increase performance; tools like Bits and Bytes, OpenVINO, and BigDL can assist.
- Operational Concerns: Automating deployment processes, potentially using DevOps practices.
---
Deployment Considerations
Deployment Strategies
- Serverless Deployment: Using platforms like Cloudflare and Base10 for flexible, on-demand GPU resources.
- Containerized Model Servers: Running models in Docker containers on VMs or bare metal with dedicated GPU resources.
- API-Based Interaction: Treating model serving as a separate API, allowing flexibility between application and model infrastructure.
Tools and Frameworks
- Hugging Face Transformers: Comprehensive library for various AI models and data handling.
- Model Optimization Libraries: Optimum for Hugging Face models, Bits and Bytes for quantization, and others for specific architectures.
- Deployment Frameworks: Base10's Truss for easy model packaging, Hugging Face's TGI for text generation inference.
---
Conclusion
- Hosts conclude with a reminder to approach model experimentation thoughtfully, leveraging the myriad of resources available online.
- They encourage listeners to share the episode and engage with the Practical AI community.
---
Sponsors
- Neo4j: Promoting the NODES 2023 conference.
- Fastly: Bandwidth partner for the podcast.
- Fly.io: Infrastructure provider for deploying applications.
Additional Resources
- [BigDL GitHub Repository](https://github.com/intel-analytics/BigDL)
- [Hugging Face Blog on Bits and Bytes](https://huggingface.co/blog/4bit-transformers-bitsandbytes)
- [Previous Episode: Running Large Models on CPUs](https://changelog.com/practicalai/221)
---
This episode provides valuable insights into the practical aspects of working with generative AI models, offering listeners a roadmap from exploration through to effective deployment.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:28Welcome to Practical AI. and database close to your users. No ops required. Learn more at fly.io.
0:42Welcome to another fully connected episode of the Practical AI podcast. In these episodes, Chris and I keep you fully connected with a bunch of different things that are happening in the AI and machine learning community. And we talk through some things to help you level up your machine learning game. My name is Daniel Whitenack. I am a founder at Prediction Guard, and I'm joined as always by my co-host, Chris Benson, who is a tech strategist at Lockheed Martin. How are you doing, Chris? I'm doing great today, Daniel. How's it going? It's going good. You know, this week I was, well, it's been an interesting couple of weeks for me in that I was at the Intel Innovation Conference out in San Jose the week before last.
1:30And then this week I was at the Go Programming Language Conference called GopherCon and taught a workshop there. And so that was really enjoyable. So two weeks in sunny California or mostly sunny California, I guess. That was really cool. So maybe even just highlighting a couple of cool things that are happening in those communities at Intel, there were a couple of things that were highlighted that might be of interest. One is, it seems like Intel is really diving into the idea of AI-enabled applications on your local machine, which I know is something we might talk about a little bit in this show in particular.
2:15That is like, hey, if I want to build a desktop application that people actually run on their laptop. And I want that to run stable diffusion as part of the application and not, you know, reach out over the network to some API. How would I build that? And what would those sort of like AI PCs is, I think, what they're calling them, what would those have to look like? And they're thinking about that with some of their processors, which is interesting. And then on the data center side, they had a bunch of things, including announcing the Intel developer cloud, which is cool because you can go on there similar to other cloud environments and spin up either a VM or actually connect to a bare metal instance that has their latest generation of processors, including these Gaudi 2 processors, which are from Habana Labs.
3:13They were acquired by Intel, I forget when, but they have, so they would be sort of on the data center side. You're running accelerated workloads on these. And we're actually running some of our prediction guard stuff on these Gaudi processors and seeing really great performance. So those are a couple of things highlighted from there. And yeah, I don't know. Have you heard those themes in your conversation as well in terms of either new processors, advances in data center technology or this kind of local inference side of things? I have quite a bit actually. And I'm certainly not an expert on microelectronics by any stretch, but I have friends who are and listen to them closely when they talk.
3:57There's a bit of an ongoing revolution on the microprocessor side. And so, you know, many of us that have been in the AI world for a long time. There have been, for instance, GPUs from NVIDIA have been kind of a core to that. But there's a lot of chip types that have been coming out by a number of different vendors to compete with that. Famously, Google was probably the first one well-known with their TPUs, Tensor Processing Units. But there's all sorts of specialized chips and chiplets that are coming out that are enabling these types of things. So I think Intel is definitely one of the global leaders in that and looking forward to having it'll be nice when everyone's laptops and phones and everything are all completely equipped with everything they need.
4:44Yeah, yeah. It's super interesting, especially for use cases where it's like your personal assistant, AI enabled personal assistant that really is tied to you personally. Applications like that, I think you'd want to run a lot of those things locally and not be sending a lot of that data all around. So that's kind of interesting. They also talked a lot about confidential computing, which is an interesting topic that I think maybe some of our audience, at least, wouldn't be familiar with as much from what we talk on this show about. But it is very connected to the AI world in the sense that if you are running kind of secure workloads through AI models, whether you're doing that on NVIDIA chips or other chips like we've talked about, there are ways and toolkits to enable you to actually secure the environments that you are running those models in and actually provide attestation to know that nothing has been tampered with inside of those kind of secure environments.
5:56So I'm going to surprise you. I actually know quite a lot about that. Those are trusted execution environments. Let's just say I've touched on those quite a lot. I think Intel's version is like TDX trusted. Yeah, it's something. They have a couple of different versions. That's the one that's out in the marketplace right now. But yeah, it's the idea of ensuring that when you normally, if you're running a program for audience, if you're running a program and it has to transit, obviously, from system to system, every system has a processor it's processing on. And even if you're running encryption at the application layer, you have to unwrap that encryption for the processing to happen in the chip.
6:33An adversary, you know, if it's on the order of a major nation state, has the ability to steal unencrypted information that had been encrypted in transit straight out of the processor memory. And Intel and other vendors are starting to push trusted execution environments and products and services around that which protects and guarantees the safety of that data inside the processor something i've spent some time on actually yeah it's super interesting and i think even the cto and his talk had like a t-shirt that sort of had a venn diagram kind of thing between like security and ai and at the intersection of that is a lot of you know what he talked about this sort of idea that, hey, whatever hardware you're running on, if you can combine AI workloads with these sort of trusted or confidential computing ideas, that can be very powerful and take care of at least some of the security and privacy concerns that people have with AI workloads in general, which is cool.
7:41So yeah, the two are converging in a big way because while trusted execution environments, which are referred to as TEEs, have been around for years in processors, now that we are having large federated workflows, which is really classic on cloud-based AI jobs, where you're distributing an AI inference or training across many, many systems with very, very important data that you would not want to get into an adversary's hands, that federation is really kind of pushing AI and chip providers together in that way to guarantee that. We didn't see lots of workloads that would be falling in that category until we hit the AI space, and it's chock full of them.
8:22So I keep remembering things that happened over the past couple of weeks while I've been traveling and people have mentioned, But one maybe other noteworthy thing for people to be aware of on the more of the infrastructure side, which I think we will talk a little bit more about in this episode, is that Cloudflare announced their workers AI. And I think this is the latest in this sort of series of serverless GPU solutions. So these worker AIs are Cloudflare's version of the serverless GPU type environment that we've talked about with things like modal or base 10 or banana. And there's a lot of these coming out, but I think it's worth noting that a very large player like Cloudflare is now kind of dipping into this serverless GPU space, which I think also signals that we'll be kind of seeing in the cloud side more and more push towards serverless GPU workloads and environments that support that.
9:29Interesting. Very interesting. Well, that's a bunch of infrastructure and confidential infrastructure and computing and security stuff that has crossed our paths in the past couple of weeks. But one of the questions that you asked me leading up to this recording was about things are moving so fast. And I think deploying and managing an AI workload may look different now than it even looked six months ago. and it's been a while since we talked through the kind of developer or technical team perspective on how you might if you want to use one of these models that's coming out all the time so mistral ai's model just came out the ones that received huge amazing amount of funding just earlier in june and now they have their first model out it's released apache too so you can download it so the question is, let's say you want to use one of these great models that's coming out these days and you want to host it in your company's infrastructure or even just play around with it as a developer.
10:41What does that look like currently? Because there's also, along with these models that are coming out, new tooling that's coming out all the time. So what does that look like these days and what are the various options and things to consider as you're interacting with these models and considering even hosting them yourself or integrating them in your own infrastructure? That's a fair question because it's been a while since we talked through some of the infrastructure, I think, Chris. It has. And for what it's worth, I'm going to brag on you for a second since I know that you would not do that to yourself.
11:16With Daniel being the founder of Prediction Guard, this is a topic that he is a global expert in, really, really knows what he's doing. And as we were talking about, I've had so many people asking me these questions that Daniel was just talking about lately. And I was like, well, you know, one of my best friends is a real pro at this. So thank you. If you can kind of start walking us through, and this is a moving topic, as you just pointed out, it's changed in the last few months and, and we'll continue to evolve over time. But yeah, if you can start walking us through what that looks like today.
11:50You know, we're in the beginning of the fall of 2023, something that might help the rest of us for at least the next few months. And maybe one note on this is I'm also getting these questions all the time. And like you say, I'm deploying models all the time with PredictionGuard. I think a lot of people, if you're a developer or a infrastructure person, you just have that natural desire. Even if you end up using a model that's behind some API that's hosted by someone else, it can be useful and instructive in building your own intuition even to just try deploying one of these models, see what's involved, see how they run, that sort of thing.
12:31It's also kind of worthwhile from my perspective to experiment with different models before you say, you know, lock yourself into a certain model family or something. it's relatively easy now with the tooling to get somewhat of a sense of how these different models perform and build up that intuition for yourself even if you end up using a model that's behind an API. I mentioned I was at GopherCon this week and that was some of the questions that came up to taught a workshop on generative AI and that was a good long discussion in there that people had a lot of questions about was hey let's say I didn't want to use one of these APIs how do I pull down a model and use it.
13:14So yeah, let's jump in. Let's first maybe talk about something that I know that we've touched on before, but just to emphasize here, where can you get models? And let's say that we're putting aside for a second, the kind of closed proprietary chunk of models. These would be ones from like OpenAI, Anthropic, Cohere, et cetera. they have their own APIs. They host those models. Let's say that we're interested in either an open access model, but it could be either an open and somewhat restricted model or an open and somewhat permissively licensed model. And we've talked about that on the show too.
13:58For example, there's models that come out that are licensed for commercial use or non-commercial use or research purposes only, but let's say you want to use one of these open access models. The first question that might come up is where do I find these models? The best place that you can find these models is on Hugging Face. So if you go to the Hugging Face website, just huggingface.co and you click on models, you'll see that there's at the time of this recording around 345 ,000 models on Hugging Face. A few to choose from. Yeah, yeah, a lot to choose from. And think about this, those of you that are familiar with GitHub, right?
14:40How many GitHub repositories are there? There's a lot of GitHub repositories that are someone like tried something in an afternoon and uploaded something to their GitHub repo, right? It doesn't mean that's the most useful thing for you to use in your workflows, although you could kind of learn from it, maybe. It's similar on Hugging Face there's a lot of people that might like oh I tried fine-tuning this model and now I uploaded it to my repo on Hugging Face and similar to GitHub one of the things that you want to look at just as a practitioner is look at how many people are downloading the model look at how many people are hearting the model or you know liking the model and you can filter by those things.
15:28So if I click on model, I can then click on a filter like the task that I'm interested in, a computer vision task or an NLP task or an audio task. And then I can look at both the trending models and how many models were downloaded, filtered by things like licenses and languages. So yeah, I think the first thing to be aware of is just the landscape of models and where you find them and the best place for that currently although there are other repositories is by and far hugging face and go there and treat it similarly to github and that there's going to be a lot of there that might not be of interest to you but there's going to be some really great things there as well
16:22what's up friends there's so much going on in the data and machine learning space It's just hard to keep up. Did you know that graph technology lets you connect the dots across your data and ground your LLM in actual knowledge? To learn about this new approach, don't miss Nodes on October 26th at this free online conference. Developers and data scientists from around the world will share how they use graph technology for everything from building intelligent apps and APIs to enhancing machine learning and improving data visualizations. There are 90 inspiring talks over 24 hours. So no matter where you're at in the world, you can attend live sessions to register for this free conference.
16:58Visit Neo4j.com slash nodes. That's N-E-O, the number 4, J.com slash nodes.
17:10Okay, Chris, I'm on Hugging Face and I see a bunch of different models that are potentially available to me. and I can click on, for example, object detection and see that the trending model that I'm looking at is from Facebook, DETR ResNet 50. Seems like people have used ResNet quite a bit. 603 ,000 downloads. And so maybe that's a good place I want to start if I'm looking at object detection. If I go to, let's say, automatic speech recognition, up at the top would be OpenAI's Whisper model, which is a great choice and released openly that you can use for speech transcription. If I go to, for example, text generation, which a lot of people care about these days, the trending one right now is this new Mistral 7 billion model that we mentioned earlier was just released.
18:13So let's take those as our kind of example. Let's say I want to run something like OpenAI Whisper, or I want to run text generation with Mistral, 7 billion. Or there's even a range of sizes of models, right? The 7 billion model from Mistral, Falcon, 180 billion was released recently. So one question that I think people have is, how do I know which model might serve my task well? And one thing I'd like to recommend to people is even before you try to download the model yourself and run it, you can go in and click on these models. Like if I click on Mistral 7 billion version 0.1. if you notice on the right hand side of the hugging face model card for that model a lot of these models already have a hosted interactive interface that you can just click the compute button and see the output of the output of the model so it's kind of like a playground that you can see a bit of the output of you can do the same thing with a lot of you know computer vision models or audio models and then below that you'll see a little thing called spaces using mistral 7 billion or if you're on whisper spaces using whisper these are little demo apps that are actually hosted within hugging faces infrastructure where people have actually integrated mistral 7 billion.
19:49And a lot of these are kind of just a simple input output interface. And so even without downloading the model, if you're just trying to get a sense for what these models do, you can click through some of these spaces that are using them or just look at that kind of interactive playground feature and just try, you know, upload some of your own prompts or upload some of your own audio or whatever that is to see how the model operates. I think a lot of people might miss this if they're just scrolling through. Let me ask you a quick question. When, if you're looking and you're trying to narrow down, you know, which model you want to pick, we've talked on previous episodes about some of the concerns that go with different sizes and such.
20:32So are there some models that, unless I have a very large infrastructure available to me, many, many GPUs, for instance, that I should probably disregard? Is there like a minimum and maximum practical threshold that would say that I have some hardware, but not everything that I would dream about that I might want to go for? So there's kind of an answer to this and then a follow-up. One is for this sort of transformer language models, oftentimes if you go much beyond 7 billion parameters, maybe pushing it up to kind of 13 to 15 billion parameters, you're not going to be able to run it very well.
21:14just by default by downloading it and running it with the kind of standard tooling on anything but a single accelerated processor like a GPU. And even then, most of the time, not on a consumer GPU. However, the follow-up to that is that a lot of people have created open source tooling around model optimization that may allow you to run these models on consumer hardware or even on CPUs. And I'd like to talk about that here in a bit, that a lot of times you may want to consider this sort of model optimization piece of your pipeline when you're considering how to run the model because sometimes the sort of default size and default precision of the model might not be best for you both in terms of your needs in terms of performance or in terms of the hardware that's available to you but i would say in this phase of like what model is going to be good for me go ahead and put that sort of hardware concern although it's important put it a little bit to the side and focus on which model is giving me the output behavior that I want, right?
22:29Because you have a certain task in mind, right? And if you could figure out, hey, this model kind of does what I want, and it seems like it's giving pretty reasonable output. And then you find out, oh, well, I can't run it on the GPU that I have, or I need to figure out how to run this on a CPU, then And then that kind of narrows down the type of tooling that you're going to have to use for optimization or you might not need to optimize at all. So kind of start with the smaller models and build up to something that fulfills the behavior requirements that you have by just using some of these demos, using some of these spaces.
23:07And then think about, OK, I've now figured out I need Falcon 180 billion. So what does that look like for me to run that in my own infrastructure? Then there's kind of a follow-up series of things that we can talk about related to that. Gotcha. Thanks. So I was kind of getting ahead of myself then a little bit in terms of worrying too much about hardware first. Yeah, yeah. I think the question, well, maybe it's because I come from a data science background, right? My data science experience always tells me start with the smaller models and work your way up to the bigger ones until you find something that behaves in a way that will work for you and then figure out the kind of infrastructure requirements around that.
23:54Because if you start smaller and work to bigger, it's going to be easier to work with that smaller model infrastructure wise and latency wise and all of that. But some people do have really complicated sets of problems where they need a really big, like, let's say, you know, I want to produce really, really, really, really good synthesized speech or really, really good transcriptions from audio. I'm going to need maybe a bigger model than a really, really small open AI whisper model. So it has to do with the requirements of your use case as well, I would say. okay so let's say you identify a model and you're you've kind of picked what you want to do where do you go from there yeah so let's say that you've picked a model and let's take the first case where it's a model that could reasonably or you think it could reasonably fit on a single processor a single accelerator or by your own sort of infrastructure constraints you need it to operate on a single accelerator and even if you don't have those infrastructure constraints i think one recommendation i often give is it's just way easier to run something on a single accelerator or a single cpu so i personally recommend to people even if it's a bit larger of a model convince yourself that you can't run it on a single accelerator or a single cpu before you make the jump to spin up a gpu cluster or something like that it's just a lot harder to deal with even with good tooling some good tooling around that side which we can talk about so yeah let's say that you found a model i don't know let's say it's our mistral 7 billion model You should be able to run that on a single instance with an accelerator or GPU.
25:53I would then look at that model and depending on the type of the model, oftentimes in the model card on Hugging Face, hopefully if it's a nicely maintained model in Hugging Face, then it will likely just like a readme and GitHub, it will likely have a little code snippet that says, hey, here's an example of how to run this. What I usually do in that case is I just spin up a Google Colab notebook because I want to see how this thing runs and how many resources it's going to consume. So I'll spin up a Google Colab notebook. If people aren't familiar, Google Colab is just a hosted version of Jupyter notebooks with a few extra features like you can have certain free access to GPU resources.
26:44There's similar things from like Kaggle and PaperSpace and DeepNote and a bunch of others. So spin up one of these hosted notebooks and just copy paste that example code in that notebook and try a single inference. And oftentimes what you can do in these environments is if you look up at the top right corner of Google CoLab, there's a little resources thing and once you load your model in you can actually look at oh how much gpu memory am i taking up right how much cpu memory am i taking up and that gives you a good sense of hey i loaded this model in i performed an inference if i just do nothing else like the the most naive thing i can do then i'm consuming 12 gigabytes of gpu memory or something like that And that kind of tells you if you don't do any optimization, then you're going to need a GPU card that at least has 12 gigabytes of memory.
27:43And so maybe you use like an A10G or you could use an A100. That might be a little bit overkill in this case. But one of these with maybe 24 gigabytes of memory, you have a little bit of headroom there. And you can say now you've narrowed down not only the model, but potentially the hardware. where assuming you don't do any optimization, potentially the hardware that you could use to deploy it. So as of yet, I haven't spun up really any infrastructure. This is kind of my standard thing where I'm like, hey, what's the deal with this model? How do I perform a single inference? And what kind of resources am I going to need?
28:20It's a nice little cheat code equivalent of finding out what you're getting into, it sounds like. Yeah, yeah, for sure. And if you happen to have, the other way I've done this in the past is if you happen to have a VM or maybe it's just your own personal workstation and you have a consumer GPU card, if you have Docker running on that system, you could pull down a pre-built transformers, hugging face transformers Docker image and just run it interactively, open a bash shell into that Docker container and run an inference, just like I said, or spin up the model loaded into memory in Python. And then in another tab or another terminal, just run Docker stats and it'll tell you, you know, how much memory you're consuming and that sort of thing.
29:11Or run NVIDIA SMI or the similar for other systems or other processors that would tell you how much GPU memory you're running. So this is kind of a next phase that I do. The first is like, maybe what kind of model do I want? The second is how do I run an inference with this model? Then kind of is a whole branching series of funness, which is either you go down the path of saying, I want to optimize my model in some way to run it either faster or on fewer resources. Or I want to go down the path of saying, nope, this is fine. I can run it with the resources that I figured out it needs. And then you kind of move on to the deployment side of things.
30:07so
30:19okay chris let's say that we want to follow the path on our choose your own adventure that you want to do some model optimization on your model. Okay. The reason you would want to do this is one of two reasons. One is, hey, it turns out I crashed my Google CoLab trying to run Falcon 180 billion because I ran out of GPU memory. And turns out you need more GPU memory for that or multiple GPUs. And I don't either have access to that or I don't want to pay a bunch of money to spin up a GPU cluster and run the model in a distributed way. Or it's maybe even a smaller model and you want to run it either faster or on standard non-accelerated hardware.
31:09Like I heard a talk at GopherCon about a workflow where people were running a model at the edge in a lab to process imagery coming off of a microscope. And it was all disconnected from the public internet. So in that case, you just have a CPU, maybe you need to optimize on the CPU. So there's gradually more and more options that are out there to do this. Some people might have seen things like Llama CPP, which is sort of a implementation of the Llama architecture that's very efficient and allows you to run Llama language models on like your laptop or on like an, I think a lot of people were running them on MacBooks with M1 or M2 processors.
31:55If you want to kind of scroll through this set of optimization stuff, if you go to the Intel Analytics Big DL repo, that's Big DL, like big deep learning. First of all, the Big DL library does a lot of this sort of optimization or helps you run these sorts of models in an optimized way. But they also have this little note at the top which is actually a very i found it to be a very helpful little index as well they say this is built on top of the excellent work of llama cpp gptq ggml llama cpp python bits and bytes q laura etc etc etc these are all things that people have done to run big models in a smaller way i guess would be the right way to put it so bits and bytes is a good example of this Hugging Face has a bunch of blog posts about this where they've run the big Bloom model in a Google Colab notebook by loading it not in the full precision, but in a quantized way.
33:05But there's a lot of different ways to do this. And that's a kind of a good reference to see a bunch of those different ways. At some point for a future show, we should come back and revisit that. That sounds really cool. Yeah, yeah. And I think it probably deserves a show in and of itself. People might refer back to an episode that we had with Neural Magic on the podcast where they talked about the various strategies for optimizing a model to run on commodity hardware like CPUs. But there's a ton of different projects in this space, both from companies and open source projects like OpenVINO and Optimum and Bits and Bytes and all of these.
Read the full transcript
33:49So if you are needing to take this big model and either make it smaller or run it more optimized on certain hardware, then you might want to go through this model optimization phase. Assuming you did that or you didn't need to optimize your model, then we get to deployment. Now, Chris, what's in your mind when you think of these days, where might people want to deploy models? Yeah, I think so. It's one of those situations where a lot of people I'm talking to are trying to decide between cloud environments. And we're seeing some people that had dived into cloud pulling back and investing in their own, as well as starting to explore some of the other chip offerings.
34:36So people are kind of reconsidering that go cloud when it's too big for you now and looking at these open models in their own hardware and trying to figure out, okay, I don't really know how to do that at this point. So that's where I'm really curious is let's say that we go ahead and buy a reasonable GPU capability in-house, but it's not too big. What can I make of that? If I'm willing to do a little bit of investment, but we're not talking millions and millions of dollars kind of thing. Yeah, yeah. Yeah, so it might be good for people to kind of categorize the ways that you might want to deploy an AI model for your own application.
35:16And even before I give those categories, I think I also normally recommend to people that I think still the best way to think about deploying one of these models, if you're deploying it to support some type of application in your business or for your own personal project or whatever it is, any type of scale, I think you're going to save yourself a lot of time by thinking about the deployment of the model as a REST API. and then your application code connecting to that model or a REST API or a GRPC API or whatever type of API you want. But the purpose of the model server is to serve the model. And then you have your application code that connects to that.
36:00Now, that could be running on the same machine or the same VM as your application code, or it could be running on a different one. But as soon as you make that separation a little bit, I don't really promote people microservice everything, But I think in terms of model serving, it's useful because you can take care of the concerns of that model, maybe the specialized hardware it's running on, and then take care of the concerns of your application separately. And if your application is a front end web app or is something written, an API written in Go or, you know, Rust or whatever it is, then you don't have to worry about like, oh, how do I run this in a different language or that sort of thing?
36:42You just handle that through the API contract. So that's maybe one kind of classical separation of concerns, you know, that any developer would be doing. Yep. Yep, exactly. And then you can test each separately, all of that good stuff. Sure. But if we think about the categories of how you might deploy these things, there's the case where you would want to run this in a serverless way. like we already talked about what cloudflare just released but there's a whole bunch of these options like cloudflare and banana and base 10 and modal and a bunch of different places where you can spin up a gpu when you need it and then it shuts down our scales to zero afterwards and there are so depending on the size of your model and how you implement it the sort of cold start time or the time it takes to spin up that model and have it ready for you to use might be somewhat annoying for you.
37:42But the advantage is you're not going to pay a lot. So you could at least try that first. There's kind of this more and more offerings in that space. But a lot of them have like, you know, base 10, the Cloudflare thing, whatever it is, you're going to be running it in someone else's infrastructure. So if you have like your own on-prem thing or something like that, maybe a little bit harder to deploy that sort of serverless infrastructure because they've optimized those systems for what they are. So likely in that scenario, you're signing up for an account on one of these platforms and you're deploying your model there and then you can interact with it when you want.
38:22A second kind of way you could do this is like a containerized model server that's running either on a VM or a bare metal server that has an accelerator on it, one or more accelerators on it. Right. And so you could spin up an EC2 instance with a with a GPU or, you know, you could even run this as part of an auto scaling cluster that's like a Kubernetes cluster or something like that. But these would be VMs that have a GPU attached or something like that. And they would be probably up either all the time or they would have uptime that's different from the serverless offerings. Sure. And so you'd just be paying for that all the time.
39:06And in those cases, maybe you could use a model packaging system like Base10's Trust is one that I use. But there's other ones as well, Selden and others, that will actually create a model package in a Dockerized way that allows you to deploy your system. Is there any standardization yet in that space or does each vendor have its own approach? I think each vendor has its own approach. Like if you look at Hugging Face, they have the TGI or text generation inference project, which I think is what they use a lot to serve some of their models. And that kind of is set up differently than Base 10's trust, which is set up differently than Selden's system.
39:52There are some standardization in that if you have a general ONNX model or something like that, there's various servers that take in that format. But the way in which you set up your REST API might be different in different frameworks. So this is a very framework dependent thing, I would say. Gotcha. And there's also an additional layer of choice here, not only in terms of what framework you use, but also in terms of optimizations around that. So there's certain optimizations like VLLM, which is an open source project. So this doesn't modify the model, but it modifies the inference code that allows the model to run more efficiently for inference.
40:40So this is not the sort of model optimization that we talked about earlier, which is actually changing the model in terms of precision or in other ways. But this is actually a layer of optimization of how the model is called that helps it run faster. So, yeah, there's a lot of choices there as well. And I think once you get to that point and you've chosen, let's say you're using Base10's trust system and you've deployed your model either on a VM or in a serverless environment or you're using whatever system you're using, I think then kind of gets to these additional operational concerns about how do I plug all this together in an automated way?
41:30So if I push my model to hugging face or if I update my inference code, how does that trigger a rebuild of my server and then redeploy that on my infrastructure? And that gets closer then into what is more traditionally DevOps-y infrastructure automation type of things, which is its own whole land of frameworks and options and that sort of thing. But it's more of a standardized thing that software engineers are familiar with. Right. That's kind of, from my perspective, that's if we were to just summarize, you kind of go from model selection and experimentation, which I would say, don't spin up your own infrastructure necessarily for that.
42:16And once you figure out a behavior of a model that works well for you, then decide if you need to optimize it to run it in the environment you need to. if so optimize it and then once you're ready to deploy it think about a model server which is geared to specifically inferencing of your model and that's the separation of concerns and either you you can use a framework like one of these we've talked about or you could build your own you know fast api service around it or whatever api service you like and deploy it in a way that is ideally automated so that you can do all the nice DevOps-y things around it.
43:00That sounds really good. So you've done a fantastic job of laying everything out. I think I've talked to you hoarse at the moment, trying to cover everything. Be careful what you ask for, Chris. So as we are winding up for this episode, what are some of the kind of open source go-to tools that pop top of mind for you that you tend to find yourself going to over and over again, you know, for folks to explore? Yeah, I think on the pulling a model down and running it for inference, just that sort of series of things, there's really nothing, in my opinion, that beats the hugging face transformers library.
43:45And this is not for people that aren't familiar. This is not just for language models and that sort of transformers, but this is general purpose functionality that you can use also for speech models and computer vision models and all sorts of models, both in terms of data sets and pulling down models and extra convenience on top of that. There's not really anything I think that is more comprehensive than that. And Hugging Face has a great Hugging Face course where you can online, if you just search for Hugging Face course, it'll walk you through some of that. In terms of the model optimization side of things, I would recommend checking on a few different packages.
44:31One of those is called Optimum. It's collaboration between a bunch of different parties, but it allows you to load models with the Hugging Face API. So similar to like how you would load them with Hugging Face, but then optimize them on the fly for various architectures like CPUs or Gaudi processors or special processors. In terms of like quantization and model optimization of the actual model, like the model parameters, you could look up bits and bytes by Hugging Face, OpenVINO by Intel, this big DL library from Intel, which I mentioned that readme and that GitHub also links to other things that people have done.
45:17So it's nice that you can kind of explore that as well. And there are other projects like Apache TVM and others that have been around for some time and do model optimization. Yep. And we've talked about that one before. Yeah. And then on the deployment side, there's an increasing number. The one that I've used quite a bit is called Truss from Base 10, T-R-U-S-S, like Bridge Truss. And that allows kind of packaging and deployment of models. You don't have to use their cloud environment. You can deploy to their cloud environment if you want, or you could just run it as a Docker container. But it's really this packaging.
45:57But there's other ones I mentioned too, like the TGI from Hugging Face or VLLM if you're interested in LLMs. So yeah, there's kind of a range there. And of course, each cloud provider has their option to deploy models as well, like SageMaker in AWS, which a lot of people use also. So I think you've given us plenty of homework to go out there and explore a bit. Yeah, yeah. There's no shortage of things to try. It can be a little bit overwhelming to navigate the landscape, but I would just encourage people, you know, that first step of figuring out what model you need to use doesn't require you to deploy a bunch of stuff.
46:38Just try it in a notebook. And once you figure that out, then find a way, even just search for like, oh, you found out you want to use Llama to 7 billion. Just search for the great thing now is you can search and say like, running llama 7 billion on a cpu and there'll be a few different blog posts that you can follow to figure out how people have done that and so just follow that path and kind of follow some of the examples that are out there it's not like any of us that are doing this day-to-day don't do the exact same thing like when we deployed recently on the gaudi processors and intel developer cloud I just went to the Habana Labs repo where they talk about Gaudi and they have like, you know, text generation dot pi example or whatever it was called.
47:28And, you know, there's a lot of copy and pasting that happens. So that's OK. And that's how development works. So fantastic. Well, thank you for letting me pick your brain on this topic for a while. Sure. And like I said, I think you're almost hoarse after this one. But that was a really, really good instructional episode. So I'll actually personally be going back over it. Cool. Well, it's fun, Chris. Thanks for letting me ramble on. And I'm sure we'll have some follow-ups on similar topics as well. Absolutely. All right. Well, that'll be it for this episode. Thank you very much, Daniel, for filling both the host and the guest seat this week.
48:06Another fully connected episode. I'll talk to you next week. All right. Talk to you soon.
48:17Thank you for listening to Practical AI. Your next step is to subscribe now, if you haven't already. And if you're a longtime listener of the show, help us reach more people by sharing Practical AI with your friends and colleagues. Thanks once again to Fastly and Fly for partnering with us to bring you all Change Talk podcasts. Check out what they're up to at Fastly.com and Fly.io. And to our Beat Freakin' Residence, Breakmaster Cylinder, for continuously cranking out the best beats in the biz. that's all for now we'll talk to you again next time
From the publisher
What is the model lifecycle like for experimenting with and then deploying generative AI models? Although there are some similarities, this lifecycle differs somewhat from previous data science practices in that models are typically not trained from scratch (or even fine-tuned). Chris and Daniel give a high level overview in this effort and discuss model optimization and serving.
Changelog++ members save 2 minutes on this episode because they made the ads disappear. Join today!
Sponsors:
- Neo4j – NODES 2023 is coming in October!
- Fastly – Our bandwidth partner. Fastly powers fast, secure, and scalable digital experiences. Move beyond your content delivery network to their powerful edge cloud platform. Learn more at fastly.com
- Fly.io – The home of Changelog.com — Deploy your apps and databases close to your users. In minutes you can run your Ruby, Go, Node, Deno, Python, or Elixir app (and databases!) all over the world. No ops required. Learn more at fly.io/changelog and check out the speedrun in their docs.
Featuring:
Show Notes:
- BigDL
- Article: Making LLMs even more accessible with bitsandbytes, 4-bit quantization and QLoRA
- Previous episode: Running large models on CPUs
- Baseten’s Truss
- Seldon
- Hugging Face’s TGI
- Intel Gaudi 2
- Intel TDX
Something missing or broken? PRs welcome!




