In short
Practical AI Podcast Episode Notes
Episode Title
Large Models on CPUs
Description In this episode, host Daniel Whitenack speaks with Mark Kurtz, Director of Machine Learning at Neural Magic, about optimizing large AI models for CPU inference. The discussion covers the challenges of model size, the inefficiency of unused parameters, and the advancements in model optimization techniques, particularly in the realm of CPU-based implementations.
---
Key Concepts and Discussions
- State of Model Optimization
- Definition: Model optimization involves techniques to make AI models smaller and faster while maintaining their performance.
- Techniques:
- Pruning: Removing unnecessary connections in a neural network.
- Quantization: Reducing the precision of model parameters from FP32 to INT8.
- Distillation: Training a smaller model to mimic a larger one.
- Importance of Model Efficiency
- Deployment Cost: For enterprises, costs in deploying models can far exceed training costs, with inefficiencies resulting in significant financial loss.
- Use Cases:
- Edge Computing: Real-time applications where latency is critical (e.g., security cameras).
- Server-Side Applications: Higher throughput and cost savings through optimized resource usage.
- Running Models on CPUs vs. GPUs
- Performance: Recent advancements allow large models traditionally run on GPUs to be efficiently executed on CPUs, significantly reducing costs.
- Efficiency: Dynamic cache hierarchies in CPUs can outperform GPUs by reusing data efficiently.
- Sparsity in Large Models
- Parameter Usage: 90-95% of model parameters may be ineffective during inference.
- Example: ResNet50 can maintain accuracy with only 5% of its weights used.
- Sparsity Techniques: SparseGPT and other methods help retain model performance while reducing size.
- Model Optimization Techniques
- Training Aware vs. Post-Training:
- Training Aware: Continuously training the model while optimizing.
- Post-Training (One-Shot): Utilizing a smaller calibration dataset to optimize without retraining.
- Open Source Community and Tools
- Sparsity Tools: Neural Magic's SparseML and SparseZoo provide frameworks and models for practitioners to optimize their own models easily.
- Community Contributions: Collaboration from researchers and developers to share sparse models and optimization techniques.
- Future Trends in Research
- Post-Training Optimization: Research focuses on reducing the need for retraining while enhancing sparsity.
- Lower Bit Quantization: Efforts to push quantization down to INT4 or lower for efficiency.
- Sparse Training: Seeking methodologies that allow for sparsity from the beginning of the training process.
---
Key Takeaways
- Model optimization is critical for reducing costs and improving performance in AI applications.
- The disparity between GPU and CPU performance is narrowing, allowing for broader accessibility of large models.
- Techniques such as pruning and quantization are essential for efficient model deployment.
- Open-source tools like SparseML and SparseZoo are making it easier for practitioners to optimize models without extensive expertise.
- The landscape of AI is rapidly evolving, particularly with the increasing relevance of generative AI, and effective optimization will play a key role in its deployment.
---
Sponsors
- Fastly: Bandwidth partner providing fast, secure, scalable digital experiences.
- Fly.io: Platform for deploying applications and databases closer to users with minimal operational overhead.
Additional Resources
- [Neural Magic](https://neuralmagic.com/)
- [SparseML](https://neuralmagic.com/sparseml/)
- [SparseZoo](https://sparsezoo.neuralmagic.com/)
---
Conclusion The episode highlights the advancements and practicalities of running large AI models efficiently on CPUs, dispelling myths about GPU necessity and emphasizing the importance of model optimization in real-world scenarios. As the AI landscape continues to evolve, these techniques will be vital for practitioners looking to leverage AI effectively.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:06Welcome to Practical AI. If you work in artificial intelligence, aspire to, or are curious how AI-related technologies are changing the world, this is the show for you. Thank you to our partners at Fastly for shipping all of our pods super fast to wherever you listen. Check them out at Fastly.com. And to our friends at Fly, deploy your app servers and database close to your users. No ops required. Learn more at fly.io.
0:43Welcome to another episode of Practical AI. This is Daniel Whitenack. I'm a data scientist building a tool called PredictionGuard. And I am not joined today by my co-host Chris, but I am joined by an amazing guest who is an expert in all things model optimization and efficiency and running on CPUs, which is super exciting. I've got Mark Kurtz, who's director of machine learning at Neural Magic. Welcome, Mark. Thank you, Daniel. Thanks for having me on. Yeah, yeah, of course. So let's maybe just start out with a kind of like state of model optimization right now. So first off, could you kind of describe like when you're talking about model optimization or that set of tooling.
1:36What do you mean by that? And how does that fit within maybe the things that a data scientist or an AI person would want to do? So whenever we're looking at model optimization, we usually focus on a few different techniques. But the ultimate goal is to make the overall model smaller and faster, right? Neural networks known to be very large models, especially compared to more traditional machine learning. And it turns out the size of those models is the important part in terms of exploring a large dimensionality of a space, but actually doesn't use all of those pathways at inference time. So what we specialize in is specifically pruning, where we're going to remove connections within that network, quantization, where we're going to reduce precision of those connections in the network.
2:24So going from, you know, the typical FP32 down to Ent8, and then additionally distillation, where we're taking larger models and trying to teach a smaller model to mimic the capability and the functionality of that larger model. So it's kind of, you know, overall high level. And yeah, it's a very exciting space right now. It's kind of exponential in terms of the number of research papers that are constantly coming out on a topic. Everybody's very excited about sparsity specifically, mainly because you can turn these large models and get rid of up to 95, even 97 percent of the weights are actually useless in these.
3:02Obviously, you can use that for a lot of efficiencies around performance and energy. And that's specifically where we've been focusing in at Neural Magic and what I've been focusing in on my work. Awesome. Yeah, that's I definitely have felt this problem. So I'm sort of asking this question maybe for others out there that maybe haven't felt this problem as much. Why is it important to make models smaller or make them more efficient? How does that fit within what enterprises or even users running smaller applications? Why is that important for people, I guess is the question. Generally, there's going to be two cases that we're looking at in terms of deployment.
3:46And one would be an embedded space where we're running on the edge and trying to work there. So generally you want real-time latency and optimizing the accuracy as best as possible. So if you're using an object detection model, you want to make sure that, you know, for example, you're on a security camera, trying to draw object detection and make sure that you know when a person walks in a frame and whether or not that's alarming or not versus a dog or something like that. So in general, on that edge application, what you can do is use a larger model, remove a lot of the pieces from that larger model so you can keep the accuracy of the larger model, but take up the space of the smaller model.
4:23So significant improvement in terms of accuracy on that edge device while still maintaining, you know, the constraints that were set for you in terms of memory and latency that you need to request back with. And then the second one would be on the server side. And that's generally where we're looking at, you know, more throughput based applications and potentially also latency if they're shipping the data up to some server to be processed either on NLP or computer vision. But overall there, what we're looking at is, especially whenever people get into larger deployments on ML and neural networks, the cost significantly shifts, not from training, but to deployment.
5:08you're for a lot of larger enterprises that are actively deploying 80 90 percent of their costs is purely in deployment on these machines so what you can do is take the exact same model that you have reduce again the amount of a compute that you need to run it so that one it'll run faster but two ultimately what that means is that it's going to run significantly cheaper right and we have a cost savings on the order of you know 10x 20x even larger if you're really trying to specialize and optimize. There can be a significant reduction once you're at that scale. I would say definitely if you don't have anything deployed yet, don't worry about optimizing the model.
5:48Worry about getting a use case that works and something that you can prove out. As soon as you go into deployment, model optimization is a great thing to start because it's essentially just free performance that's left on the table that can significantly affect your bottom line. we've mostly been talking about kind of like model size and optimizations and i do want to get sort of down and get into the nerdy stuff around like how some of this works but before we do that i'm also curious about this element of deployment on gpus versus cpus it seems like some of what's indicated at least in like the tooling that you're building is like the potential to take a large model which might require a GPU at inference time and potentially run that on cheaper commodity hardware that only has a CPU maybe doesn't have a GPU like what is the state of that now and like how far can you push that or maybe also like how could people best think about that in terms of like when and when that might not be possible I guess as you said we specialize almost entirely on CPU performance.
7:03And in that, actually our latest ML inference results on MLperf have come out. So in that we show that we're running faster than T4s and A40s and things like that on this commodity CPU. So server-based CPUs, stuff that you have in your laptop, desktop, things like that. And it's very surprising that what's thought of as these little CPUs can outperform the GPU. And we see this generally across every domain that we've tackled. And that's been across image classification, object detection, excuse me, segmentation. And now we're working in the NLP and NLG space and actively coming out with that. But overall, we're seeing the same use case where these models are overparameterized.
7:50We can take away a lot of that compute. And what that means is that you can actually get the CPU and the GPU about equivalent in terms of compute throughput because with the sparsity and the dynamic setup of CPUs, we can run and skip all those zero multiplications, right? So significant reduction compute, they're about even, but then the CPU has a unique cache hierarchy, which means that we can reuse that cache more often than what you can get on a GPU. L1 and L2 being extremely quick, faster than GPUs main memory, and L3 being about equivalent. So overall, what we do on our performance optimization is skip all the compute to get even and then use that cache hierarchy as efficiently as possible on the CPUs so we can get faster memory access than even you can get on a GPU.
8:36And we pay a little bit more by doing a little bit more compute by doing that. But overall, it works out that you can actually beat the GPUs and that setup with just pure software. I anticipate this is maybe a question that you get sometimes, but it's like you hear so much about the necessity of GPUs for running these large models. Do you find generally practitioners are just unaware of this like possibility of running these large models on on CPUs? And like, how has that been for you and those that you work with that are actually doing this amazing work and have these awesome tools? Like, how has that been in terms of overcoming that barrier of perception that people have?
9:21It was a barrier that we hit, especially early on, a couple years back, where it took a lot of convincing to even get in the door to talk to anyone because they just didn't believe what we were saying. Now it's gotten to be quite a bit better, especially with the newer software that's been pushed out of the newer chipsets on CPUs. They're getting a little bit more even in terms of GPUs, so it's, you know, within a stone's throw to try and match GPUs. So people are a little bit more accepting of it. but yeah whenever we show them the numbers generally the first reaction is well let me try that on my hardware because you guys got to be doing something weird let me let me try to do that let me replicate that and as soon as I do then the next question is okay well how do I do this to my model and that's usually where the tricky part comes in is model optimization hasn't always been the easiest thing to do it can take a lot of research to enable new architectures and things like that.
10:18But that's what we've been also specializing on at Neural Magic is making all the research that we're doing, being able to put that into open source, and also building out a SaaS platform on top of it so everyone can easily play with hyperparameters and get something that is consumable. But I would say that's probably been the biggest gap in terms of trying to get people off of GPUs onto CPUs is the model optimization that needs to take place first to be able to run faster than the GPUs. You talked a little bit about sparsity, which I want to dive into over time. And I want to get to those tools and open source stuff.
10:53You mentioned like these large models and people are probably used to hearing these numbers, you know, 3 billion, 7 billion, 13 billion, like however, and up from there, like these models have in terms of numbers of parameters. Could you describe a little bit what you mean when you say 90 to 95 % of these connections, or maybe less than that, but a high percentage in some models have no impact on the actual forward pass or inference in the model. Could you describe a little bit more by what you mean by that? Yeah, definitely. And I'll take two steps to doing that. One is just covering kind of the 90, 95 % class, at least where we've been able to get to on those.
11:35And the second is looking specifically at large language models. So for the first one, whenever we're looking at getting rid of 95 % of the weights, let's take ResNet50 as an example. This is our toy benchmark model. This is essentially what we prove out all of our technology on because it's a common feature in MLPerf and for most performance tests. So what we can do coming in is looking at those convolutional layers. It has, I forget how many million parameters within it, but it's definitely not the 3 billion, 7 billion or up on top of that. But within that, we can actually zero out. So what we're doing is taking all, imagine taking all those parameters, dumping them into a giant array, and we're just going to zero out the ones that are not important.
12:18And figuring out the ones that are not important is part of the research. The easiest assumption is just saying that the weights that are the largest are the ones that you want to keep. So the ones that are furthest from zero are the ones that you want to keep. Generally, you can think of this in two ways. One is that as the model's training and being regularized, the weights that don't matter are going to move towards zero. And then the other thing is during that forwards pass, the weights that are higher magnitude have more an effect on the output, right? And everything else is going to be noise in between.
12:49So we're able to essentially get rid of just our, whenever I say get rid of, I mean setting those parameters to zero within 95 % of them. So you're left with 5 % of your weights that are non-zero. And that's actually all that you need to preserve the accuracy on ImageNet for ResNet 50, for example. And some quick kind of intuition in terms of how I've been able to think about this and why it works and things like that. You can see as we increase the size of our dimensionality in our optimization space, what we're doing is, and there's a few research papers out on it, that we're able to connect more of the local mins.
13:26right so the optimization process will slowly converge further and further down because more of the local men's are connected generally though there's only a few of those pathways that you actually need to connect those local men's so all that we're doing is we're following down that most optimized pathway and removing everything else around us in terms of that dimensionality so it's kind of one of those things that as you're training it's slowly selecting the weights that matter that gets you down to that local men and there's very few so the important part was that large dimensionality of the optimization space but not every direction mattered right so then we can get rid of it and then diving in on the llm side and large language models we actually have a recent paper that came out from one of our principal research scientists dan alistar called sparse gpt and that's where we're looking at taking opt and bloom models all the way up to 175 billion parameters and be able to optimize those and remove as many weights as possible all in this case in one shot so just using the model without any retraining we're able to get rid of around 60 percent of the weights without doing anything and there's a new paper out of cerebris actually that was looking at the llm story and they're able now to get to 80 sparsity on these llms with retraining so that's kind of the research direction that we're headed down now is proving out how optimized we can make these models.
14:52Because there's also a lot of interesting stuff that happens with the large language models, specifically because it's generating one token at a time, very latency bound. And that means that it's a lot of memory access to load those weights. So if you can quantize those and then get rid of half of them, you're already at anywhere from a four to six X speed up just on your inference times. And that's generally where we're focused and looking at currently to try and get those LLMs to run faster. The other thing to call out for those two is, you know, 7 billion parameters and 175 billion parameters.
15:26Those don't fit in a single GPU. So now you have, you know, clusters of GPUs to serve one model. And a lot of that compute is just completely wasted because all that it's going to is trying to maximize the memory on the GPUs. For CPUs, you can throw a few terabytes on there and it works out fine. So that's the other thing to call out with the LLMs in terms of GPU versus CPU.
16:01This is really interesting, Mark. I want to follow up on what you were just talking about, which I think is a really, it's sort of a subtle point, but it's really interesting in that I think if I understood you right in what you're saying like let's say that I have one of these large models 175 billion parameters or whatever and even for inference I have the necessity to have multiple GPUs just to load that model into the memory of the of the cards whereas on a CPU you can have terabytes of memory what I'm assuming is like you could load that in as long as you're able to execute it quickly, which I guess is the other piece.
16:41So am I right? You sort of have to have both like the ability to load it into memory and you have a bit more space in that on the CPU side, but then you also have to be able to execute it very quickly, which I guess is like why you would think about both space and sparsity. Is that an accurate way to put it? As you said, you know, you have a total space you need to take up and then a minimum latency that you want to respond to the user out, right? And that's going to set the constraints for your hardware. And for CPUs currently, at least for the smaller models, you can get to a usable speed on those.
17:18If you've seen like llama.cpp, they're doing int4 and things like that on smaller models and they're usable, but they're less accurate. So what we're trying to do and what we're actually working on right now is making sure that we can get that GPU class speed while maintaining the large memory advantage of CPUs. So you can deploy this 175 billion parameter model on something local and you don't have to worry about data privacy, anything like that. It's just there working and available and highly accurate for you. I definitely heard people that I've talked to who have tried various optimization techniques and have maybe been dissatisfied with the performance hit that they're getting, not in terms of compute, but in terms of actual model performance or accuracy or whatever.
18:04Does that performance hit often come about because of the quantization that's maybe part of the optimization techniques? Or are there multiple sources of that performance hit? How should people think about that? I think the biggest thing there is honestly just the amount of choices that people have to apply and not knowing when to apply them. because generally for quantization, for example, you can apply quantization to pretty much anything at end eight for both activations and weights and have it recover. But there are definitely cases where, for example, we've been quantizing efficient, that's quite a bit on our image classification side.
18:48There's one or two layers in some of these that are extremely sensitive for whatever reason to quantization that you can't quantize those. So removing those, then you get 100 % recovery, right? So it's a lot of these kind of little things that researchers know intuitively in terms of having used this constantly. What will work? What won't? But that's not really coded into software anywhere to make it easy for people to use. So generally, they'll go through, try and quantize, and there's no feedback loop. There's no methodology. It's just, hey, I was able to apply it in one shot, but it lost 5 % accuracy.
19:22Talking about quantization. but if you do you know a quantization or a training scheme generally you'll recover all of that back and it generally works completely for that same thing on pruning pruning you'll definitely see more of a drop and it's much more of a requirement to do training aware on the pruning side at least to get to really high sparsities but you definitely will see this kind of the choices that are made in the hypergrammeters that are chosen those can significantly affect the recovery and the quality So generally, I'd say, you know, if they were seeing drops in performance, it's primarily because of those choices and those issues and just the wide breadth that's available right now and not knowing how to narrow it down.
20:02And that's what we're actually working on. That makes sense. Yeah, yeah. That's good for people to myself. I want to develop a little bit more intuition around these things, because like you say, sometimes you're just like, oh, here's the command that I run on the command line. and like I get this file out that is smaller, right? But I don't have a great intuition about like, similar to hyperparameter tuning, right? Like that takes time to figure out like, okay, how should I think about changing my learning rate if this happens or if that happens, that sort of thing. You mentioned one thing which I think would also be good to kind of clarify and help people understand is like training aware optimization optimization versus just non-training, I guess, optimization, or I don't know what the counter to that is.
20:53Could you talk about those, like how they're differentiated? Some people might guess what that means, but like, how are they differentiated and how does that work out in practice in terms of how you would optimize a model? Technically, we have three categories generally available. One, and the two you're going through is one training aware. And then we have post-training or one-shot, which are kind of interchangeable. And then we additionally have on sparse transfer, which is something that we've been pushing a lot because the research has worked out quite a bit for it. So I'll cover all three of those in a little bit more depth.
21:24So for training aware, what we're doing is taking the exact same model that you're wanting to deploy in the exact same data set it was trained on and continuing the training process further. So while we're continuing that training process, we could be continuing it. This is where the hyper parameters come in but generally it'll be about half the time that it originally took to train it we'll train it for that much longer and as we're doing that training we're iteratively pruning away or we're applying quantization or both and the reason we're continuing that training is because as we're iteratively applying these optimizations we're slowly moving the model away from its local men right and it has to adapt and adjust back so by slowly doing that and training over some time we allow small jumps that the optimizer can recover from and adjust the remaining weights for rather than doing it all at once.
22:15And the all at once piece is where we get into post-training in one shot, where we're not going to try and retrain the model at all. We're going to take a small calibration data set, and then we're going to use some heuristics or algorithms to figure out, using that calibration data set, how to optimize that model and remove weights or quantize. So the most common case would be static quantization. We're using a calibration data set to figure out the activation ranges for each layer, right? And once you have the activation ranges for each layer, you can set up a simple quantization scheme to say, given that it's going from, this layer is going from negative six to six.
22:57Now I need to fit that range into an intake scale of zero to 255, right, and be able to map that in. So that would be a simple post-training or one-shot application. And then the final one is sparse transfer, which works exactly the same as transfer learning or fine-tuning. It's just we're starting with a sparse model to run through, and that's a lot of what we've pushed up in neural magic into our sparse zoo are these open source sparse models that we have you know sparse birds and res f50s and yolo v5s things like that and you can just take those plug in your data set and transfer over to it so the sparsity mass stays in place and it just adjusts the remaining weights to fit your data set we have a few papers out on that as well that shows that sparse transfer works just as well as regular transfer yeah that's really interesting.
23:51I guess this would somewhat depend to like, if I'm just thinking of like the average practitioner out there, right. Probably in a lot of the space that I work in, in like the larger language model area, then like, I'm not going to be able to retrain one of these large models on the original data set, right. Or even half of that or for half of the epochs or whatever. But these other things are certainly things that I do all the time, right? Like fine tuning, transfer learning. So it's cool to understand that there's options out there. Am I correct in assuming that like you mentioned the zoo, the sparse zoo that people can find on your website and we'll link in here too, that like researchers, your team, practitioners, whoever are the people out there are also putting in work to actually release some of the things that are going to be able to do.
24:48of these sparse models publicly to the community so that I can take those and then maybe do fine tuning on that. Or maybe it's just good enough from what's released in the community. Could you tell us a little bit about that community and what's being released? So on the open source side, pretty much everything that we have on the sparse suit currently has either been from our lab. We have a few that are from Intel's lab up as well. And some hugging face examples and things like that, primarily because a lot of the sparsification research is all built around a few models like ResNet 50, BERT, and things like that.
25:27They don't expand it out past those models to prove out their algorithms. So we have the best of those. And then, yeah, that's exactly what our team is working on. And what we get some community contributions in every once in a while for sparse models that people have generated or transferred, right? So our goal is to able to push these up so that as you said anyone can come in from the community and be able to pull those down and get value out of those models so you can think of them as sparse foundational models rather than the dense foundational models
26:10mark in addition to the sparse zoo which is really cool i know that neural magic is producing some pretty uh interesting and useful tooling otherwise as well in terms of actually doing some of this optimization themselves. And I see certain things, DeepSparse, SparseML. Could you describe a little bit about if I'm a practitioner, I have a model and I want to do optimization, what does that look like for me right now with the tooling that's available? So we have SparseML, which is our open source model optimization framework. It's built primarily on top of PyTorch. And then we have integrations with Torch Vision, Hugging Face, YOLO, Ultralytics, YOLO v5, pretty much all the common repos that most people are using.
27:00We've already integrated with, so you can use our integrations and just plug in your model and go along with that. And then the other part is that we have recipes, and I'll go through recipes in a second. But we have those kind of pre-coded integrations, or you can create your own integration. Usually it only takes a few lines of code. we've done all the hard work in terms of making sure that whenever we want to optimize a model that you just have to wrap the optimizer in PyTorch. Essentially wrap the model and the optimizer, which is what our code handles, and then it's going to go through and handle the optimization past that.
27:29So there's no coding really from your side or from the practitioner's side on implementation. And then the other part is then coming up with an optimization recipe that people want to use. And what that's going to lay out is saying that I want to prune from this epoch to this epoch, for example, and then apply quantization and target these layers and at this sparsity level, things like that. We have automated ways to generate those as well as examples on the sparsity and in other places. That's generally what it would look like is, you know, get up and running. We definitely recommend checking out the sparsity first, seeing if there's anything that you can do to just transfer the model onto your data set, because those are the quickest and fastest ways.
28:09and then otherwise if you do have a specific model architecture that you're looking at then you can start going down this integration pathway and generating your own recipes and that's the current state that we're at the other thing that i wanted to call out is that we're working on a saft platform right now to make all of this more intuitive so you have a ui to be able to predict where that model is going to end up before you start optimizing it and then additionally actively benchmark it across your different deployment scenarios so this is called sparsify we've had an old kind of alpha state for a while that was a downloadable and we're actually going through alpha testing right now so anyone that is interested in trying that out definitely reach out we're going through alpha currently looking to go to beta and probably next month two months and then ga following up after that so definitely check that out as well and yeah yeah that's great um yeah i i love how you're thinking towards uh usability as well around these things because i i I do see this as something that kind of blocks people on optimization a lot because they get stuck.
29:13Like you say, it's like, well, what recipe do I use here? It seems like there's a million options. Like, where do I start? So that's really awesome to hear. Just so people can understand, like, there's options available in the Spar Zoo. I'm kind of scrolling through there now. There's a lot, even ones that I know that I use, like Distilbert, which is fine-tuned on squad. I use that all the time for question answer. And apparently I should use the sparse one because, yeah, that would help out a lot, both in terms of compute and speed. Let's say that a model is not there. And in particular, people, of course, are interested in all of these things being rapidly released all the time.
Read the full transcript
29:55So one of the things I know that I've seen in optimization platforms over time is like, it's hard to maybe support new architectures as they come out. So how are you all approaching that? And like, what is the state of like being able to be flexible with these optimization schemes for a variety of architectures? looking at the base framework that we have pushed out in sparse ml for example everything's implemented such that it's supposed to be able to run with any model architecture so there's no real big assumptions on there other than you have convolutions and you have linear layers inside of your model somewhere that we can target for optimizations so everything's set up very generically which means that whenever there is a new architecture that comes out or new weights things like that it's going to take some time for you know us to tackle that and that's if it gets onto the list of you know top used models things like that but that is the nice thing about having an open source community is that people are welcome to come in the tooling and the framework there should work out of the box with everything and they're more than welcome to be able to commit or push up you know whatever they're working on and we have an active slack community and github community for people to come in.
31:15Our engineers are actively on that. They can come in, look at it and easily get support on any issues they're running into. Awesome. Yeah. And we'll make sure for the listeners who are wanting to get plugged into this, we'll make sure and include the links to the slot group and the GitHub in our show notes. So make sure you visit there and get plugged in and start optimizing your models. As we kind of get a little bit closer to the end here, I'm wondering about a couple of things. One is, I know that, like you're saying, you're actively involved in research in this space. What trends are you seeing in research around optimization?
31:56And in particular, what are the directions that your team is sort of excited to go into in the near future in terms of the research side of this? So I would say there's two big trends right now. One is around the focus on post-training. Our principal research scientist, Dan Alistar, he came, him along with his lab, came up with this algorithm called OBC-OBQ, which actually we have a webinar which will have already aired by the time that this comes out. It's airing tomorrow. but that algorithm specifically is where a lot of efforts going around which is requiring as little data as possible and no retraining and trying to increase sparsity as much as possible so that's one of the key things that people are trending down is trying to do that the second one i would say is a large push around quantization in terms of getting to lower bits and that's something that's been around for a while but it's something that's becoming more and more prevalent as we're looking at the larger models, mainly because their execution time is mainly dominated by trying to pull in these large matrices of weights.
33:04So there's been a big push there now around trying to get down to, you know, pass N8 down to int4, int3, int2 quantization on these active representations. I'd say that's the other trend. The final kind of bonus trend that I'd throw in there, because I said too, the third one would be more research around sparse training. So specifically trying to figure out how to start with an unoptimized and untrained model and be able to make it sparse from the start and then keep sparsy throughout training. Because generally what we do to guarantee accuracy, we'll start from a dense converge model and then iteratively print on top of that, which adds training time.
33:47So now there's a lot of research going into trying to figure out how quickly the model can be pruned and then be able to carry that over as training. So that's where the other big, big pieces. And all three of these are definitely active areas that we've been investing heavily in, especially looking at the generative AI space now going through that. Yeah, yeah, I know it must be a crazy time for you all, just like it is a crazy time for everyone. But yeah, I think this is a really important piece of it. I know one of the trends we've even talked here on the show about, it seems like a lot of people are talking about like serverless deployments of machine learning, deep learning models.
34:31And I know a lot of the issues related to that and things that people are dealing with is cold start time and loading models into memory. I don't know if that's impacted you all at all, but it seems definitely relevant. Like if you're going to run your model serverless, you probably want it as small as possible, I would imagine. Yeah, absolutely. Absolutely. Absolutely. So you mentioned where people can find out about neural magic on Slack, on GitHub. I would really encourage people to do this. As we close out here, what are you kind of personally excited about during this? I mean, like I say, it's a crazy time for everyone right now with generative AI and the way things are trending.
35:12What's exciting to you right now about the AI community and certain things you're seeing? What do you see as kind of positive trends, I guess? The part that I'm most excited about is that generative AI space specifically in being able to augment humans. Obviously, there are a lot of privacy concerns and data concerns and bias issues and things like that in this, which I don't want to see LLMs deployed everywhere become a default response for like Google search or something like that. but it is really exciting to see even in my day-to-day starting to use these actively to augment what I'm doing around content generation and framing and things like that so it's one piece that I'm really excited for and with the work that we're doing in neural magic we're especially looking at these because one we want to see that continue to grow to open source and I think that's been the other push that's been really big and really exciting to see is that whenever GPT-4 came out, it was completely privatized.
36:13They put out a little white paper on it that had no details about it at all. A lot of data concerns and things like that within that. But the open source community has already released, I mean, I can name probably 10 models so far that have been released since then that are chat GPT-like or GPT-4-like. So it's really exciting to see that. I think the next stage from those open source models is going to be making them runnable anywhere, right? So you don't need this big GPU cluster farm to get something that is usable. And that's where we're really looking at going. We're actively working on the LLM deployment issue right now and hope to have something out in the next few weeks, the next few months that people can start actively using, download it, they can run it anywhere they want on any CPU and it'll be just as fast as GPUs.
36:59Cool. Yeah, well, keep us posted. I know I'm personally interested in that one. So yeah, thank you so much for joining us, Mark. This is a really fun conversation. And I love getting into the weeds of these practicalities because this is a topic where people get stuck a lot is on the deployment side and the optimization side. So, yeah, thank you for all that you and your team are doing at Neural Magic in this area. And, yeah, keep up the good work. We're excited to see it. So thanks for joining. Thanks, Daniel. It's great talking with you.
37:42Thank you for listening to Practical AI. Your next step is to subscribe now, if you haven't already. And if you're a longtime listener of the show, help us reach more people by sharing Practical AI with your friends and colleagues. Thanks once again to Fastly and Fly for partnering with us to bring you all Change Talk podcasts. Check out what they're up to at Fastly.com and Fly.io. and to our Beat Freakin' Residence, Breakmaster Cylinder, for continuously cranking out the best beats in the biz. That's all for now. We'll talk to you again next time.
From the publisher
Model sizes are crazy these days with billions and billions of parameters. As Mark Kurtz explains in this episode, this makes inference slow and expensive despite the fact that up to 90%+ of the parameters don’t influence the outputs at all.
Mark helps us understand all of the practicalities and progress that is being made in model optimization and CPU inference, including the increasing opportunities to run LLMs and other Generative AI models on commodity hardware.
Changelog++ members save 1 minute on this episode because they made the ads disappear. Join today!
Sponsors:
- Fastly – Our bandwidth partner. Fastly powers fast, secure, and scalable digital experiences. Move beyond your content delivery network to their powerful edge cloud platform. Learn more at fastly.com
- Fly.io – The home of Changelog.com — Deploy your apps and databases close to your users. In minutes you can run your Ruby, Go, Node, Deno, Python, or Elixir app (and databases!) all over the world. No ops required. Learn more at fly.io/changelog and check out the speedrun in their docs.
Featuring:
Show Notes:
- Neural Magic
- SparseML
- SparseZoo
- Neural Magic Scales up MLPerf™ Inference v3.0 Performance With Demonstrated Power Efficiency; No GPUs Needed
- Deploy Optimized Hugging Face Models With DeepSparse and SparseZoo
- SparseGPT: Remove 100 Billion Parameters for Free
Something missing or broken? PRs welcome!




