In short
Podcast Summary: Speculative Decoding and Efficient LLM Inference with Chris Lott - #717
Podcast Overview Podcast Title: The TWIML AI Podcast Episode Title: Speculative Decoding and Efficient LLM Inference Host: Sam Charrington Guest: Chris Lott, Senior Director of Engineering at Qualcomm AI Research Release Date: Not specified Episode Link: [TWIML AI Podcast Episode #717](https://twimlai.com/go/717)
Episode Description In this episode, Chris Lott discusses the challenges and solutions for accelerating large language model (LLM) inference on edge devices. He delves into the encoding and decoding processes of LLMs, the hardware constraints affecting inference metrics, and various techniques for improving efficiency, such as KV compression, quantization, pruning, speculative decoding, and the use of small language models (SLMs). The conversation also touches on future directions for on-device agentic experiences.
---
Key Concepts and Discussions
The Nature of LLMs
- Two Distinct Computations:
- Encoding: Involves processing the input query in one go, highly compute-intensive.
- Decoding (Generation): Involves generating tokens sequentially, with every new token requiring the model to process the entire context. This presents bandwidth limitations.
Hardware Constraints
- FLOPS (Floating Point Operations Per Second): Essential for understanding the compute capabilities of devices.
- Memory Footprint: Limits the size of models that can be run on edge devices due to DRAM constraints.
- Bandwidth: The key limiting factor for generating tokens in real-time applications.
Metrics for LLM Inference
- Time-to-First-Token: Measures the time from input submission to the generation of the first token.
- Tokens per Second: The rate at which tokens are generated after the first token is produced.
- Tokens per Joule: An energy efficiency metric, crucial for mobile devices.
Techniques for Accelerating LLM Inference
- KV Compression: Reducing the size of key-value pairs to save memory.
- Quantization: Lowering precision of model weights to fit within memory constraints.
- Pruning: Removing unnecessary parameters from the model to make it more efficient.
- Speculative Decoding: Using smaller draft models to predict multiple tokens and then validating them with the main model. This process helps utilize excess compute capabilities more effectively.
Future Directions and Innovations
- On-Device Agentic Experiences: Enhancing user interactions through personalized AI assistants on mobile devices.
- Hybrid AI Models: Efficiently integrating local processing with cloud capabilities for improved performance.
- Recursive Speculative Decoding: A new method for speculative decoding that leverages tree-based structures to efficiently explore token generation.
Conclusion The discussion highlights the ongoing evolution of LLM technology, the challenges of efficient inference in constrained environments, and the innovative solutions being developed to improve performance. Chris Lott emphasizes the importance of hardware-software integration and the need for continuous research and adaptation to meet the growing demands of AI applications.
---
Key Takeaways
- Efficient LLM inference is crucial for enhancing user experiences on edge devices.
- Techniques such as speculative decoding and KV compression can significantly improve performance metrics.
- The field is evolving rapidly, and ongoing research is focused on enhancing model efficiency and integration with existing systems.
---
Additional Resources
- [TWIML AI Podcast Show Notes for Episode #717](https://twimlai.com/go/717) - Complete show notes and links to relevant research papers discussed in the episode.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00A language model has two very distinct kinds of compute that you do. One is you're going to have this big input that just represents a query that you're putting into the input field. And in one gulp of the model, you're going to try to compute all of that. So we call that the encode problem. And that really ends up being very compute limiting. The real problem is the generation or the decode problem. The key point is every single one of those tokens that you generate, you've had to pass through the entire model once. So immediately you would say, well, we have to then have parity with bandwidth and compute power.
0:31But the reality is the physics of all of this is such that no system generally is like that.
0:50All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. And today I'm joined by Chris Lott. Chris is a Senior Director of Engineering with Qualcomm AI Research. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Chris, welcome to the podcast. Thanks, Sam. Very happy to be here. I'm excited for our conversation. We are going to be doing what I think will be a bit of a deep dive into efficient large language models on the edge. before we dive into the topic i'd love to have you share a little bit about your background and how you came to work on all this stuff okay um yeah i started out actually working on test instrumentation and then early days of gps so my background communication signal processing but then i got into control theory and and then went and did a phd in control um and specifically stochastic control, which has a lot of commonality with what we call AI and ML now.
1:54So then I joined Qualcomm, and in the early part, I did system design and standardization of all the wireless systems, and that led to all these voice systems and high-speed cellular data that did it up through LTE. And then as Qualcomm developed, we went beyond modem and we became an SOC company, which means we started adding more and more integration of all the different functionalities that you might want on your phone, on your tablet, into one chip. And then that brings in a lot of compute aspects, GPU, CPU. And then we started adding accelerators for AI. And then this led to, I guess, getting into the whole AI field for those of us who had the right kind of background and the right interest.
2:37And so that's where I've been in, say, the last probably eight years. And your focus is on the engineering side within research. So is that primarily working with productization of innovations on the research side? So our group is called AI Research, Qualcomm AI Research. But we are organized, we want our research to be, basically there's two poles and then a lot in between. So one pole is contributing to the research community focused on papers and the frontiers at that level. And then there's what we call system engineering, which is integrating, basically doing design and being the main system design and support group for AI in the company.
3:21But we actually are organized in a way we want our AI researchers to also be part of system engineering at different levels. So everyone has some aspects of both, but my focus has been more on the system engineering, partly because of my system design background, but also that's where a lot of our action is right now. We're going through this period where AI models, we're getting to the rubber hits the road aspect of where are they useful? How do we make them efficient and how do we actually integrate them into full design systems? systems. That's maybe a good segue to talking about LLMs on the edge and why you see that as being so important, you know, relative to using them via APIs on the cloud.
4:08No, exactly. And that's a big discussion and this topic that'll be going on for a while. But there's a number of reasons we think supporting LLMs on the edge is the future. And it's partly because we can We have the compute and bandwidth and just storage aspects that we can do a lot on the Edge. We have that capability, and it just makes sense, economic sense, to use it where it's feasible. But also, the Edge is also a place where there's a lot of unique information and personal data, etc., that can be used to improve how the LLM is used. So obvious examples are camera, audio, sensors, and just awareness of, you know, the phone indicates something about your situation.
4:58And you can use that information to help when you say, oh, what is this? Where am I? Or any question you might pose, you can use the sort of contextual information that is only, you know, growing over time. So primarily situation and personalization aspects? Yeah. Yeah. So some of it, so I was going to say it also, it can extend over time, something that it would learn from you as you are using the device. Okay. And then, you know, it's certainly feasible to store all this on device in a very private way in your own database that can then be accessed while you're, you know, trying to achieve different tasks.
5:36So yeah, it can learn your preferences, your habits, and just the typical things that have happened in the past that often are a best indication of what you're looking for now. You mentioned storing that type of preference information in local databases. What's the relationship between databases and models with regards to personalization and preference and intent type of information? So you might view this as a form of RAG, which is just instead of having the model have to store everything, you're actually storing information that it can then retrieve when it needs to. And so, I mean, there's definitely one major trend, which is just the standard database, non-SQL, SQLite kind of things that are very readily supported on a mobile device.
6:30And then the system simply is deciding what aspects of your daily life to store. And this could be, you know, your recent emails, your recent chats, discussions, and, or just other aspects of your behavior, your calendar, of course. And all of these things can be made available for an agent that's on the device that's trying to anticipate, you can try to anticipate what you need, or simply retrieve things that may be hard for you to remember. But, you know, as simple as, oh, you have this thing coming up in your calendar, but also far more, obviously far more. What new LM technology lets us do is do all of that far more deeply.
7:10And so you mentioned kind of this personalization and situation opportunity, but I'm also imagining, and you also mentioned kind of the availability of the local resource, you know, compute and the like. Like I've got to imagine latency is a big part of the attraction of using the device. I'm thinking of my own experiences using like OpenAI, advanced voice mode. Like it's amazing, but it also, you know, there are some glitches that at least I experience frequently that seem like they have to do with latency. you know like you know the prematurely stopping the a generation or listening and it seems like if we could tighten that loop we'd be the performance and the experience would be better it depends on what you're trying to do and what the model is because there is a limit on what we can support and how long it takes to run on the local device and and whether we can do the task that you're needing right now.
8:21But our goal definitely, so this brings up this whole question, and this is one of our great questions now. How much can we pack into the smaller language models? And how efficiently can we have them run with these databases and things where the model doesn't have to store everything itself, it just has to be an effective agent and communicator, right?
8:42And how are the resources on the device being used to efficiently be able to respond? So we size everything. So yes, to minimize that latency. And that is what we are working on also is new developments in hardware, new developments in software, and also algorithmic things that I can also describe on how we do improve that efficiency. But yes, the goal is literally you have this device in front of you that is capable of just immediately understanding you and responding with minimal latency. And via voice, not just text. Oh, yeah. And this, in fact, if you know what one of our main, I'll say pitches right now, is that this is all going to lead to voice UI, which is literally, and if you want the very high level simple scheme, anything that you are currently doing on your phone right now, you actually will be able to do instead of clicking or touching or controlling it with your hand.
9:41And for at each, say you're going to an app or a webpage and having to click through and figure things out, it will learn how to do things. And also that you'll be able to give high-level commands in voice and say, oh, please just go do blank. And something that it has learned to do or already knows how to do. And that literally running apps on the phone, this becomes how you do it in the long run. Yeah, I was at the Snapdragon Summit when Cristiano was talking about this idea. vision. And, you know, while there are a lot of, you know, gaps that need to be filled in order to get there, like the idea that, you know, part of what a quote unquote agent on your device is doing is masking or abstracting away a lot of complexity, you know, both, you know, hardware and software and settings and all these things that exist on our devices and making them easily accessible via, you know voice uh natural language voice commands i think is a compelling one because these devices can do so much yeah and they clutter our mind having to remember how to do this and that whereas we know we just want to do blank and we don't want to have to remember all this stuff yeah christiano our ceo his i remember his statement is something like we want to break the paradigm of the app construct or something like that but but it is this exactly um but you know it really just plays into an age-old dream of AI agents, AI just being able to have enough understanding of everything that it really just becomes this useful little assistant to you.
11:14So, of course, we're here to dig into the details. And I think maybe a place to start there is to really narrow in on the device constraints and where the way we need to use the devices to work with LLMs bump up against those constraints. Yes, let me mention, so one of the first things we did as this whole generative AI field has been effectively exploding is how far can we push things on our devices? We've been working on AI for a while, but large LLMs are a new kind of, they have very specific properties and they stress the hardware and software in unique ways. And so just I do think a good example is when we said, okay, we're going to put the LLAMA 7 billion model on a smartphone.
12:08And a little while ago now, but that was our first say, okay, what does it take to do this? And what are the limiting factors? So I'll just mention a language model has two very distinct kinds of compute that you do. One is you're going to have this big input. It might be hundreds, thousands of tokens. It just represents a query. It could also represent data in context learning that you're putting into the input field. And in one gulp of the model, you're going to try to compute all of that. So we call that the encode problem. And that really ends up being very compute limiting for almost any kind of hardware.
12:44and um and so but you know we've we've put enough compute power in our uh in our uh edge edge devices that they are able to do this pretty efficiently and i could give numbers at some point but like yeah you can do you know we we look at like 4 000 input size fields and we do it in a couple seconds something like that and that's already quite powerful actually um but then we realize the real problem is the generation or the decode problem of LLMs, which is they're autoregressive in the sense that what they do is they take all those inputs that have been encoded and then they say, okay, what's the, then it just starts generating one token.
13:24And here token really can just think of as a word, right? One word at a time. And it goes, okay, generate the next word. Okay. Now that's the context. That's, that's the history. Now generate the next word. And you do that one word at a time. And then the key point is every single one of those tokens that you generate, you've gone through the entire model. You've had to pass through the entire model once. Okay. And in fact, in our computer architecture circles, they have a parameter they call the intensity, which really is every time you read something from the memory, how much compute do you do with it?
13:59And it gives you some sense of the balance of how the bandwidth is supporting the compute. Maybe I'll just contrast with in the old convolutional network days, every weight that you read from the model was used, you know, dozens, hundreds of times something. And that just made it, that's usually what leads to being compute bound, the way things are sized. But with transformer decoders, they literally have an intensity of one. And it just, which just means you do one MAC, one multiply, accumulate for every weight that you read in over the bandwidth. with. Now, and so, so immediately you would say, well, okay, we have to, we have to then have parity with bandwidth and compute power to make that balance.
14:43But the reality is the physics of all of this is such that you don't, no system generally is like that. Because moving data takes energy and, and takes, you know, it takes a certain amount of time. And basically bandwidths are some orders of magnitude lower than the compute power on many devices, most devices, I would say. So we're talking about DRAM bandwidth here. Those are things like DDR3, 6500, that kind of thing. These models are so big that you don't put them on like the local memory that might be local to the processor itself. You will have like a local, we call it the TCM, the tightly coupled memory to the processor itself.
15:29But that memory tends to be measured in the megabytes, maybe even single digit or double digit megabytes. But we're talking about models that take gigabytes. And so it's only cost effective at all to do that in DRAM. So then, yes. So then, and that is the whole fundamental crux of token generation, which is we're going to generate one token at a time. We have to read gigabytes of memory into the processing for that purpose. And that's going to happen over the DDR bandwidth. Yeah. So, and then you're absolutely right. There's these, there's these standards for the memory controllers that, that support various bandwidths.
16:05And, and just, I think maybe we're saying we're currently in what we call LP5. Okay. And sorry. Yeah. Yeah. We're currently in LP5 and in the future we'll be moving to LP6. And, but the kind of deltas you get in bandwidth currently are, our, our processors are supporting 50, 60, 70 gigabytes per second. Okay. And in the future, that'll go to like 113 is what I understand. That's what LP6 will give. That's quite, you know, it's almost factor of two. That's nice. But it's still, we're still heavily bandwidth limited. It's nowhere near what the processors could support just from a compute point of view.
16:42You talked about the difference between encoding and decoding and encoding is compute intensive and decoding is bandwidth limited. how do those limitations translate into llm metrics that people care about so when talking about encode i'll just bring up what is the most standard one which is called time to first token and it literally is the time as i was mentioning in the encode problem you do maybe you'll do 4 000 tokens all at once you can do that in one pass of the model if you have the right structure so it all comes down to how fast you can compute for for for like a 4k input and and then when you're done then you then as soon as you finish that you're going to generate your first token of output so if you're the user you may have asked a question and that it may have pulled in some external information put that into the input to the lm and and the lm will process all that use that and then start rendering the output so that the time it took from you to to ask your question to when that first token appears, we call time to first token.
17:47But once the first token appears and then things start being generated, your output response generated, then it's that time of, we call that tokens per second, of course. It's just the amount of time it took. So the output might take a few hundred tokens. That would be typical for some kind of response, or it might be something even longer. And then tokens per second is that metric. There are other very important metrics for the edge. In fact, they might be somewhat unique to edge devices because these devices are mobile often and they have their own power source at least some of the time. And you may realize in smartphone design, it's like battery life is like one of the most fundamental things that people care about, just power, how much power you're using.
18:32And so we also have - So imagining you're going to say something like tokens per watt or something. Right. Actually, tokens per joule, if you know, because the point is, watts is joules per second. So you're using up some power while you're outputting. So it's really the relative rate of those that tell you, like, really the point is how much joule is an unit of energy. How much energy did you use to output 100 tokens is how much you sucked from the battery. You can think of the battery as storing joules, and then each token takes up some joules from the battery. you could do that in milliamp hours by the way if you want a different unit but that we actually write we actually do things in tokens per joule or tokens per picojoule or however you want to however you want to do it given the constraints like a big part of what you're doing from a research perspective is trying to optimize around those constraints what what are some of the kind of foundational ways that you think about delivering lms on the edge in light of those constraints Yeah, I'll make one more thing that was not even on my radar at the beginning, but is very much what we all think about now.
19:39The other big economic issue with smartphones, tablets, etc. is how much DRAM you can put into them. And it's pretty much purely a cost question, but it's a very non-trivial one. And the typical phone these days, I believe, four gigabytes, many have eight gigabytes. That's about the typical range. And then what's interesting, like consider LLMs that we want to support. You need a certain size LLM to just start having some of the real, you know, impressive power that they can show. And, you know, at some point we thought maybe... Power in terms of capability, not in terms of... Capability, I'm sorry.
20:17Yeah, just impressive performance they're giving on applications, right? Right, right. and in fact that's why we said look we have to at least try lama 7 billion and see if we can do that okay you know on a little smartphone and um and you can uh but if you want to know what the primary issue the primary issue more than anything else it turns out it's not even token rate or any of those things the primary issue is the footprint in the DRAM so just being able to fit the model in the RAM so you can even think about running it that's right that's right and and interestingly uh And I'll mention all of these devices, pretty much probably every computer will have a concept of this sort of fast access DRAM and then a kind of drive, like a flash drive solid state, something where it's storing things more permanently.
21:00We do have more things you can store on. You can store larger things on the flash drive. Okay. But to run a model, as I'm saying, every time you generate a token, you have to use the entire model during generation. And that means you have, pretty much that means you have to have it in DRAM. and okay so it may not be in DRAM constantly but it has to be able to be in DRAM and the way phones work the way all these systems work they'll be some of the DRAM is always filled with the OS like could be Android or iOS and then part of the DRAM is running your persistent apps which most people will have something that's just some sort of persistent app okay and and so it leaves you whatever is left and say eight gigabytes to run your LLM okay all right and then um so then consider like a llama, literally that model has 7 billion parameters.
21:49So even if you encoded it one model parameter and one byte, so there's a whole thing here I'll just mention briefly. All these models, when they're first developed, they're developed as floating point numbers for simplicity and for training, et cetera. It's really one of the great questions. And I really think there's a very interesting just whole set of things that we think about here. Basically, what are the limits of computation to achieve comparable performance to what, say, the floating point version of a model is. Okay. And now this has been studied a long time in various guises, and then DLN context is bringing a lot of that in.
22:26But the main thing, it comes down to, can we pare it down? Can we prune it? The model may be 7 billion parameters. Does it really need? It was trained that way, but does it need all those parameters? Can you prune it down? Or can you quantize those floating point numbers into smaller and smaller, you know, numbers, how far can you go? And, okay, so, and, you know, and still achieve good performance. So let me go back to saying, if a 7 billion model is in floating point, and even that would typically be FP16, if you know that term, that's two bytes per parameter, that's, you know, 14 gigabytes, right there, okay.
23:10All right, and even if it were, you know, if you could go down to one byte, well, that's eight gigabytes. All that is still way too large, seven or eight gigabytes. Okay. So then, so the first thing we do, and I'm sure this is something well known, but these LLMs, interestingly, generally you can take them down to four bits if you do quantization the right way. You can take the weights down to four bits. And that's an interesting thing. This is something very conceptually interesting about why and how the model is working properly, even under those conditions. But that is how it's done. And that's like the most standard thing now is to take these models, take them to four bits.
23:42And so a 7 billion parameter model can be taken. The model parameter is down to about three and a half gigabytes. That's at least feasible on a smartphone. And that's what our first demos and that's all the things that we were first doing. And I can just mention on the site, quantization is its own long topic. We have whole teams that do a lot. And it's really something that we actually, our research group is very known for. But actually... And I can include some links to conversations with some of the researchers about various quantization efforts in the show notes. Okay. Yeah, perfect. And then, you know, like one of the trends now in quantization is, well, different parts of the model may quantize better than other parts.
24:24And we call this multi-precision. And, you know, some of it you can even take down to two bits, interestingly. But other parts actually do better at eight bits, the overall performance. So just those are all the things that, yeah, that we're currently doing. And then one more thing I think is definitely a trend. We have efforts in this. Is how can we just, as I was mentioning, power, tokens per joule is very important. All that's fundamentally how much energy you're burning reading the model, reading the model to where it can be computed. And so it is, we are, there is rethinking of architectures for that purpose.
25:02which is minimize data movement and uh yeah you get down to some really fundamental the physics of of moving data computing and doing the computations and storing them and what's an example of that so uh i'll just say what it really comes down to is can you locate can you minimize can you can you really make the the compute as coupled with the memory as you can I mean, one example is beef up as you can the tightly coupled memory and see how far you can push that. But that's one example. So you mentioned SLM, small language models, which is that fundamentally trying to address the limitation in DRAM size?
25:50Or is it, you know, equally trying to address performance? like how do you think about where SLMs come in and why there's so much interest in those nowadays so it's definitely about the footprint the DRAM size um but it is also about those other things about the the time to pursue but but I whether those are I those may not be as driving as the DRAM size which is an interesting fact okay maybe you can accept a little slower token rate under some conditions if you're going to get a better answer but um but you can't have a bigger model under, you know, if you don't have the DRAM. Okay. And so, yeah, so there's this, I think a lot of great effort now on what we have come to call SLMs.
26:29And for me that the range of that is, you know, maybe two to four, maybe up to seven gigabytes, sorry, billion parameter, billion parameter models. But what's turning out these days, what we see is there's this sweet spot that seems to be three to four billion parameters. Is it a meaningful distinction to, you talked a little bit about the effort to get the 7 billion parameter LLAMA model running on a device. Presumably you started with a native 7 billion parameter model, and then there's a optimized smaller version of that. You can correct me in a second, but the real question I'm trying to get at is like, is when you talk about SLMs, is it a meaningful distinction to talk about like a smaller version of a larger, natively larger model or like a architected to be small language model?
27:33You know, it's interesting. It is. I am speaking about the ones that were targeted for being small language models. Okay. Mostly. Okay. By the way, they might use larger models to distill some of the capability of the large model during their training. So in that sense, you can view them. That's just not that interesting. Like, it's a confusing distinction, maybe. Well, okay, but most of them, okay, if you know, like, the PHY models from Microsoft, just pick those. And by the way, even these days, Lama has a three billion, right? Those, however they trained it, so it doesn't even really matter how they got there in some sense, right?
28:06And they will all, they'll train from scratch in the pre-training part, those giant, you know, Terra token training sets that we use. and then whether they can improve performance by distilling from a large model or they often do, what Microsoft does is these curriculum learning things and synthetically generated data of very particular types and they have all these methods to try to see if you can pack more capability into these smaller models. So whether you do it that way or distilling or whatever combination, that's not fundamental to the point that it comes down to what can we do with a model that is 3 billion parameters because that's something that very readily can go on our devices right now.
28:47And then we try to figure out what we can do with, you know, how far we can take that. And then the other, I think, really fascinating topic is I sometimes view what's happening now is there's a kind of system engineering phase for all of this in our terms, okay? And that's really, to me, what like compound AI is or any of these things where, okay, this model is not going to be standalone. It's in a system design. And how do we do that? And, um, and then, and then, then the trade-offs are not even, well, I can't fit it. So I'm going to use the smaller model. Then the trade-off is a model of this size is sufficient to perform this function in this larger system.
29:23And, and it runs faster and it's more efficient and it has less energy. And in fact, it's better for, you know, this question answering that we're doing using RAG where we're grabbing the data from outside anyway. Um, and, and there may be, you know, lots more you know we know a lot of a lot of examples now of uh scanning a lot of the things going on on the on the phone processing it coming up with summaries or or or using that information from the past where the smaller models are perfectly adequate and and much more efficient when you think about small small language models as a direction like what are the areas that you're specifically working creating new small language models creating slms based on other people's llms um you know figuring out what you know are there complementary system components that can make slms more viable like how does that translate into practical research and engineering directions for for you so maybe there's two main aspects we do try we do train our own slms and we're studying the architectures, seeing what architectures are performing well, and also how they are mapping to our hardware.
Read the full transcript
30:37We do want the flexibility to try out things in that space. But we also are users of the public SLMs, or there are those from, we do have sets of customers that bring in their own unique ones that we are supporting. but then we study this problem in our orchestrator which means how do we support an SLM in this sort of larger context of a system design and the software that's using it and you know the RAG system and the databases and all the things that might be associated and how is that all running on our devices. In that context we you know what we also will find we'll fine-tune existing models we have our own models we find existing models we have customer models all of those things are happening in in some ways it's kind of a traditional qualcomm play in the sense of like you've got this model type that's gaining popularity how do we add a bunch of value by creating components around it that will let it do more things yeah i actually that's that's a good way to put it um we we have this philosophy and this is the system engineering side we'll do end-to-end design even if we are not the provider of all components in the end design And we use this analogy in the wireless system, which I had a lot of experience, which is we would do end-to-end mobile base station, drive around and take all the data, support the standards, write the standards off.
32:01And in the end, our high volume is the mobile side chip, right? But we are also fully aware in developing the full end-to-end system design. And I think that is a good, you know, if not analogy, just exact example of what we're doing here, which is the end-end design here is down to what is actually happening on the device, on the smartphone, on a daily basis, and all the components that we provide, how are they playing into it and doing it all efficiently. And the only way to really even, even to even plan our hardware software developments, we have to know what our system, you know, what our future systems are going to look like.
32:39You've emphasized thus far kind of the memory constraints and the bandwidth for generation. Are there things happening on the compute side as well? I've got to imagine like as you're trying to use longer and longer context windows, you know, it's still important that you have efficient compute structures, right? Oh, yeah, that's where it really happens. Let me go back to the distinction between the encode and the generation decode part. um as i was saying the encode part is very compute bound which means the time it takes to first token is limited by how what your compute power is and that just becomes more and more so as the as the context window that you're basically the number of tokens you're computing increases and there's this trend um i always refer to google for their their one million context length that they have been claiming for gemini um and so the compute itself during encode isn't n squared uh n squared in the max if you like um because every token attends to every other token so you um so number of tokens you're attending to and then how many tokens that are attending and that's an n squared so um so you can get to and then you know people want to support i was mentioning 4k 4k is like the sweet spot of what we currently do on mobile devices but um but that there are many that are saying we want to do 16k we want to go to 128k you can imagine how that grows.
34:04Because then you can start doing significant things all in one process of an LLM, right? Summarize this entire novel, if people use that example, right? So, and then the N squared problem, you know, I'll just say right now on our mobile device, that'll just kill you at 128K. You know, it could be a background task, I could even say, if you're willing to wait, it's possible. But, okay, so what we do instead, though, is, and then the flip side, too, is once you have that long context and now that you're generating tokens that also slows down because every new token is based on all the previous history if that previous history is 128k that's a lot of processing you have to do just to generate the next token okay um and then i should go back to the footprint dram question so i'll introduce this term kv which people may may know but the way transformers work is uh for every token in the past that you've processed like the encoding problem is entirely take all the input tokens and create their KVs, which stands for key and value.
35:03Okay, and that may be enough here. So, but the thing is you have to allocate that. That has to be stored also in memory as you're generating going forward. So that actually takes up DRAM. And right around, you know, that starts mattering right around when the context window size is on the same order as the model size, if you want an interesting little rule of thumb. So like a Lama 7 model, Lama 8 model, the model size is 4K. That's the embedding size of the model, the token embedding size. And it's right around on that order, that context length, where the compute aspects and these DRAM aspects really start taking off and start dominating.
35:42And is that related to that intensity calculation that you mentioned? Actually, for, well, the intensity is still the same from a compute point of view, but this is the number of, it really, actually during generation, it's the bandwidth of reading in all those KVs is what's slowing you down, usually. You're still not compute bound, you're still bandwidth bound, but you're having to read more and more. So you have to, I should maybe say it that way, you have to read the entire model and you have to read the whole history of the KVs, okay, for all the previous tokens. every time you generate a new token.
36:19And if those KVs are relatively small, when you start out, those can be relatively small. But if you have a longer, longer history, it's really interesting, it flips, and suddenly the KVs are everything. And for the truly long context, yeah. Okay, now that's in kind of the raw form. And that is what puts these limits on what are practical context windows that we're currently support. But we do a very active effort, and we will have a paper coming out pretty soon on this topic, which is, we call it KV compression. and it because it's a very you know how much of those kv historical kvs do you actually have to store and now the first thing is as usual there's a quantization question uh and we and everybody realized you can basically go to eight bits for kvs but going to four bits doesn't always work so if you don't know where these boundaries are okay um but you can also interestingly you can drop some kvs or they're not all important not all the previous historical you know input words effectively have the same importance you can imagine how you could compress that just doing it as a straight compression problem without any semantic context doesn't work okay but if you have you but you can compress them based on their relative importance and we have ways this is what the whole thing is how do you estimate the relative importance of historic the historical inputs and that that even depends on what the task you're trying to achieve is so um and people do other ways like question answering yes a different question answering well you can't lose the fact of what you have to answer the question that's in one but but you know some of the tokens are the yeah right yeah and who knows you know what what really is needed so but actually and so there's ways you can do it that's very semantically aware or even context of i should say uh domain aware of what you're trying to do but another um another approach one of the things we've been doing is It's more generic, but just using the statistics of them and how they relate to each other to decide what you really need to store.
38:11And so what we'll typically, if you want to know what a very practical like on-phone question is, as I say, you have this fixed DRAM, you have a fixed allocation you're going to put for the model itself. And then you have a fixed allocation for how much KV you're going to support, okay? Because you have to have it available. And then you say, okay, but so how far our context? So simple thing is, well, that KV can support a certain, we'll pick eight bits for KVs and we'll do a context window that fits that okay that's fine but um but you can also go beyond that and then you have this compression problem you're going to decide what you're going to fit into this fixed size that you have and that's the interesting question uh that we've been working on for a while and this is very practical like this will be this is something very um customers and people are already using these kinds of things and you mentioned that the using 8-bit kvs like is there any work to look at like variable variable uh length and codings for kvs is that interesting area yeah absolutely uh we definitely have done that also ourselves um there's always this trend up in complexity versus how much it actually is worth it you know we're totally into system engineering stuff but it absolutely in the same way that it's true for weights yes it also can apply to kvs actually um that is something that we even we even we even saw instead of you You had this question of dropping an L token or making it two bits so you know there was something there.
39:34And exactly, those all trade offs exist. So KV compression is an area, you mentioned a little bit about compute architectures to adapt the compute to the model. You know, the converse of that is model architectures to adapt the model to the compute. I'm imagining that's an area of research as well. Oh, yeah. And that leads into something we track and consider even contributing more ourselves, which are the new architectures that are potentially replacing transformers or at least replacing layers and transformers, such as state space models. Yeah, maybe the best example. And I think as often understood these days, hybrids of transformer layers and state space layers, you know, like the Mamba 2s, et cetera, may end up being a real contender because you can remove these KV growth problems and storage problems.
40:33You're effectively saying, we're just going to, you know, it's funny, I'll make this comparison. And when we say we're going to fix the KV size and we have to practically, we just have to do that. And yet the KVs in the vanilla transformer will keep growing. Well, we're going to compress. Well, that's something that's quite similar in a sense to what a transfer, I'm sorry, what a state space layer would do, which is they just fix the size. They say, here's the size no matter what. We're going to aprorate this side. This is the size. And we're going to pack all the previous information that has happened at this layer into that one vector.
41:06and um which fits perfectly the kind of constraints we actually do see is what i'm trying to imply and but the but the problem is um that doesn't always work well it depends uh how far you can take that pure design and it's probably because you know as things if you're running longer and longer you know more and more things happen it's very hard to know exactly what you should have stored and again that's domain dependent it's related to the kv compression problem of depending on what you're doing whether that really works or not is what we find. Okay. But if you have some amount, the idea that you need a full KV for every layer in the model, which is how the basic transformer works, that's definitely something that could be questioned.
41:50And I didn't mention it before, but you could also, even this KV, when we say we're going to minimize, we're going to have a fixed KV size and we're going to have to compress it, that can also be a per layer decision. Or you can, I mean, you can, some layers may need less than other layers. And that actually ties directly then into what these state space models are also trying to do and their method of compression you know in many cases seems to work quite well and but standalone maybe not so much you need you need that some layers need the long memory and some don't so to us it's a natural then development beyond this our fixed kb size compression stuff into bringing in other layers that just are inherently limited in their and how much memory you need and let's see what accommodations of that work best for a given problem to me.
42:36So I would say it's not like we're going to converge on a single architecture. It's going to depend on what use case is and how much other resources you have. Again, trying to map this onto what you would do versus what the customer would do. Presumably within some bounds, if I want to write a state-space model and you know run it on the hardware it's you know it's gonna work and what you're trying to do is make it work better so like you asked a question that we are definitely we're always engaged with okay what is our best role and um so i mentioned we always do we'll do the end-to-end system design because that's how we understand everything and how it all fits okay and and so very much first thing is we can then work with we have some major customers they're doing all sorts of things, some of the major phone vendors, et cetera.
43:27We help them effectively. We give them, we say, look, we did this, this is how well it worked, and that'll influence how they do their designs. But you're right, they'll own the model in the end. But for example, we'll just give them this example of here's a state space model running on our hardware, here's a hybrid model, more likely. And then there's things that we think about that only certain others would start at least thinking about, which is, do state space models quantize well? and when you're comparing states based and transformer of these different architectures not every they generally won't start with that but it's one of the primary questions for us and and the way they the way they work where the state is limited it's a limited state size and um those and and it's it's iterative autoregressive um inherently you know it could have more quantization limitations than a transformer than just storing kbs at a certain quantization level.
44:19That's just an example of what we will do. And so the truth is, in the end, we span this range of, we give reference designs, we show, we do our, and then we provide the pieces that are fundamental to running things on our hardware. And we will give example of, you know, how we fine tune this model in this way, and this is how that fits, you know, things like that. Yeah, what's kind of coming to mind for me or becoming apparent as we kind of talk through these different ways of supporting LLMs and the edges, like almost a maturity model, like quantization and pruning. Like you guys have been doing that for years.
45:02I've been talking to you, you know, your researchers about how that space has been evolving for years. And so, you know, it's mature and there are primitives baked into the hardware that, you know, do some of this stuff. SLMs, you know, you've been working on a little bit longer as, you know, LLMs have been popular and like there are modules and the hardware that do that. KV compression is probably newer and, you know, maybe that's a software thing, not baked into the hardware yet, but maybe some point in the future it's baked into the hardware. you know the states-based models that's brand new you're feeling it out building the nm systems learning and you know maybe that gets you to a point where you say oh we can put this in our software toolkit or oh we can put this in the hardware and it kind of wrote that's kind of how it rolls out is that fair yeah yeah that's right um although it's an it's always this ongoing what is worth we call that hardening right when you um you know put things as in hardware support and you're kind of, it's a, it's, you know, it's a faith-based thing where you're saying, we believe this is so important that we're going to, we're actually going to, it's expensive to put things into hardware, right?
46:10So we're going to, we're going to, we're going to need to support this, or we believe, and, you know, and things can change so fast. Right. So it's like a, yes, we bet on a transformer that's going to be around for a while. Do we make that same bet on a state space model? And if so, like, you know, which one and what's the right timing for that and that kind of thing. I can use an example, which is part of, uh, part of what takes time doing, um, encoding for LLMs is the soft max, just the simple soft max function. And, um, and, you know, you could, you could go so far as to say, well, we believe soft max is so important.
46:47We're just going to absolutely accelerate, accelerate it in hardware. Um, but there are alternatives to soft max. And, and in fact, you know, the trend, uh, state space are different and there may be other things coming. So we're faced with this decision of we want to support the existing models as efficiently as we can, but in a flexible enough way that we can adapt, right? Yeah, so even how quantization works has been evolving and we've been adapting how we design things. So we've been talking primarily about accommodating the LLMs exclusively on the device. Are there approaches that leverage the cloud and the device together?
47:28Yeah. In fact, the way I like to talk about this is, you know, most apps are already some, you know, apps that are running on your phone are already some combination of some cloud component and what's running on your phone. And it's just the natural thing. You want to utilize the local cloud processing as you can in storage. But there's some things that may be better persistent in the cloud or compute heavy in the cloud or some large database in the cloud. So this definitely will happen as we go forward with these AI models. And everything that we're trying to enable, models of a certain power and certain functionality that can exist on the edge, but it will work in conjunction with whatever you're trying to achieve in some cloud side as well.
48:13And one of the very interesting problems is how to know what to do where. And so we call this hybrid AI. That's one of the terms. And it can be anything as simple as, here's my app. It uses the LM for these functions. You can just pre-decide. It uses the local one for these because we know that works. It'll run out the cloud for some of the more maybe planning aspects or just anything that's more difficult. But we also do, we are developing methods where that decision can be made on the edge device based on its own determination of where it needs, you know, effectively the most efficient thing is use the edge device until it needs to go to the cloud and for the more complex things.
49:03So, yeah, that's definitely the future for all of this. That sounds distinctly like where this orchestrator thing that you mentioned earlier might come in. Exactly. So back to this whole point that we do, we do this edge dead design. And as we've been doing this for a while, the clear, what becomes clear is we need to have this capability on the edge device that, you know, orchestrates all the different, all the different components. and then that can include cloud components. It definitely, it always will include going to cloud for information, going to APIs, going to different information sources that exists out in the world, but also it can go to external device, external large models, AI models in the cloud for when the local model is not sufficient.
49:58And it handles the local database, as I was describing, being data collection storage on the local device, all the context, basically all the smarts of what exists on the agent that actually exists on the edge device. So as part of doing that system design, what we, it's just another good example that of what comes out of it for us, which is this is actually integrated into our chip design and with the OS, this orchestrator, which is a layer that is gonna provide support to all of our customers that are doing these kinds of things. It's like a software product. I've heard it come up on several occasions, but I wasn't very clear on timing.
50:42You said it's going to, so it sounds like it's something that's being developed now as opposed to in people's hands. You know, some versions of it are in the hands of various customers. Yeah. I mean, I know the exact state. I should be careful. So one thing that we haven't talked about that's come up in other research interviews is this whole area of speculative decoding and speculative methods to address some of the bandwidth constraint in particular. Can you give us a recap on that and how that comes into play and some of the various research efforts that have been supporting that? Yeah, sure.
51:23So as I mentioned, very early on, we all realized, wow, we are so bandwidth limited when we're generating tokens, that whole previous topic. And well, and given the hardware as is, it's just like, what do you, okay, what do we do? And when we jumped in, there had already been a couple of early papers, very recent at that time on speculative decoding. Because this, you know, that's not just for mobile size. Let me add any, even, even big NVIDIA GPUs have this problem. Yeah. So, but the idea is, well, we have this compute power that we can, you know, as we, instead of going through and computing one token at a time for generation, as I was describing, we have the compute sitting there that we could compute all we need for multiple tokens at the same time.
52:09So how do we, how could we use that? And basically it's the balance of bandwidth and compute. We're just so way, way out of whack balance. Can we, can we, basically, can we use compute to reduce the bandwidth requirements? Okay. And that's what speculative decoding does. And it has various forms, but in its most basic form, what you do is you have a draft model. It's a small model that you can run very fast. It's basically trying to mimic what your main, say it's an 8 billion target model in our case. The draft model might be 100 million or something. something so and and and you know literally token generation is basically proportional to that size you know the the bigger it is that that's all it takes on a given bandwidth so um so you know it can run literally you know might be a hundred times faster uh than the than the than the target model okay um i should be currently 10 to 10 to 50 is what we would normally do i should say okay in terms of how much more okay um anyway so it can it can generate um Um, so it's going to have the same input.
53:11It's still an LLM. It's just a smaller one and it can generate tokens. So what you do is you say, okay, why don't you generate a few tokens into the future? And that might be four, it might be eight, you know, it might even be more in some cases. And, and then we can actually say, um, now those are not final, those are not final generation. You take those and then you're going to run the target model with those as input. And, and what you're doing is that's kind of an as if it's like, what if this first token is the next one? What if the, what if the first token is the next one? and then what would the next token after that be?
53:40And then do that for every one of those tokens. Assume the previous ones are the actual ones and then compute what you would do for the next one. Okay. What these models do is they create a distribution on the next token selection, probably distribution. Okay. And then there's a nice statistical argument, rejection sampling that can, what you can then do is after the target model does that process is it can one by one go through and decide which of those are statistically valid. And I won't get into what that means exactly. Okay. But it just means there's a way to do this such that it will accept some number of those tokens and say, okay, I'm going to give the stamp of approval from the target model on those tokens.
54:19And I'm going to throw the others out, by the way. But the idea is that the computation that the target model needs to do on these draft tokens is not the full decode. It's a subset of the decode. It actually is the full decode. I'll say this. So the target model has this ability to say, given the input so far, here's a probably distribution over the vocabulary of what the next token is. That's what it does. Okay. And the point here is you can do that exact thing for all eight input tokens, let's say, eight draft tokens, all in one parallel computation, but it's the same thing you would have done.
55:01Let's say you pick the fifth token. It's the distribution on that fifth token, given the first four were the input. Okay, even though they're speculative, these are speculative tokens, but we're saying this is the distribution on the fifth one as though the first four were actually correct. Or accepted or correct, yeah. Got it. So that once you do accept it. So there's the idea that if the fifth, the distribution in the fifth one is high enough, you can assume that the first four are good enough? It's a little more involved than that, and that's the thing. You know, there is a version of that's deterministic, by the way that just one at a time like you don't have to get into the whole rejection sampling statistics i'm trying to really understand why you know how you're able to skip steps if you know ultimately you're doing full decodes right oh here's the point the whole the whole point hinges on can i compute these distributions for all eight tokens in the same amount of time it would have taken me to do it for one token and when i say compute the compute is much higher for the eight it literally is about eight times more compute for the eight tokens than for one token but the point really is you can do that ax compute in one without the bandwidth without because the bandwidth is the same going back and forth yeah yeah okay so you're using the excess compute that you have that was just dormant if you didn't do this because it's sitting there waiting for the for the stupid DRAM to give me the next thing that's really what happens okay and then there's a whole statistics so these methods can guarantee that the distribution like the thing that you're generating has the same distribution as the original model that's that's the key statement okay so there's no approximation going on it really is generating exactly what the target model would have done meaning there's a way to do the draft and a way to do the rejection sampling that yields a system with this guarantee and better performance relative to tokens per second.
57:06That's right. But even more, let me just disentangle two things. The draft model, how the draft model works does not affect whether the distribution, whether what you're generating is correct or not. Oh, really? I would have thought you would need to ensure a certain degree of goodness of the draft model in order for the whole thing to work. There is a degree of goodness, but it only impacts the token rate, which means... Oh, meaning if it's not good enough, the target model is going to have to generate, you know, do full decodes more. Yes, it's going to reject this. Yeah, yeah, yeah. These are bad like that.
57:42Okay. You see the point, yeah. It is interesting that that's the thing. You get this guarantee that in the worst case, you're just going to go... In fact, in the worst case, you can always accept one token. Okay. And so in the worst case, the only slowdown is how long it took to run the draft model to generate some number of tokens. But that was wasted compute anyway. Yeah, but that does cut it because the target model is waiting. You're burning bandwidth. It all comes down to bandwidth in some sense. You're burning bandwidth for the draft model to run, to generate. And that does waste. So the worst case token rate is a little bit lower than what it would have been without this method at all.
58:22Okay. But not much lower. But generally, of course, as soon as you start getting some hits, then it gets better. And so typical gains, we say 2x, 2 point something x is typical for these methods, which, you know, it's like doubling the bandwidth of your system. It's really quite, you know, it became important very quickly. Yeah, it's interesting. A couple of things interesting. One is that, you know, from the time I first heard about it, which was a conversation with one of your colleagues in research, I think it was Fatih, a year and change, maybe two ago, or a year and a half or so. Like, from that time, I've heard it come up, you know, a lot in the intervening time, especially recently.
59:09And also interesting is that, and I think you alluded to this in referencing like NVIDIA GPUs having the same problem, like it comes up in the context of not mobile, but just high performance LLM inference hardware. Yeah. Any imbalance between bandwidth and compute, as I was describing earlier, this can benefit because you have, yeah, you can have this excess compute that you try to use in some way. Right. In fact, I might want to just say at that point, we have a lot of excess compute in our chips, okay, in general, compared to bandwidth. And this method I just described where you're, okay, I'm going to predict some number of tokens into the future.
59:57The value of going further and further, well, you're burning this time to run the draft model to generate more tokens. and then for those later tokens to be accepted, all the previous ones have to be accepted to even have a chance. Then you basically check them sequentially, the third one, the fourth one, and you only go to the next one if the previous ones have all been accepted. So the point is, and I'm sure it's understandable, that there's diminishing returns on how far you want to predict into the future. Okay. So, and what we found very quickly is we're still with that sort of baseline speculative, we were not filling our compute.
1:00:34And I'll tell you, you can look at our, we have this like TerraOps, our chips can support a certain peak rate, a certain TerraOps peak rate. And then look at the peak bandwidth I was saying is in the tens of gigabytes per second. And there's quite a disconnect there. We have a lot more compute. Okay. And we don't generally use that peak rate. There's other things, but I'm just saying we can support 32. We look at the 64, even in the hundreds of tokens at the same, you know, without and still be basically bandwidth limited. Okay. And so the sort of predicting four tokens or eight tokens into the future, it is not filling what we can do, what our compute can do.
1:01:16Okay. So it is one of one of, it was one of our innovations is once we see that, okay, what can we do? so then we created a version that we were calling the tree-based uh and i will add that by the way this has become something very widespread but it is something that we did early on in all of this development we do have a paper a paper on it okay um which is instead of just having one path into the future you you know very naturally you say well i'm going to pick more you know just pick two times yeah uh well okay actually there's many there's many possible flavors of this um And one of them was just do multiple.
1:01:53Yeah, you could do it multiple times. That's what you just said. Okay. But we found actually even better than that. You can do, I guess from tree base, I'm inferring pick a first token and then branch off of the first token and then branch off of those and branch off of those. Yes. That kind of thing? Okay. Yes. And then what's interesting is there's actually a few different ways to do the rejection sampling. We actually experimented, or this is literally now just writing out the equations for proving that the distribution is maintained in various ways. And I'll say we had a few ways, but the one we ended up liking we called recursive speculative.
1:02:30Okay, it's recursive rejection sampling. And that is the same or different than the tree-based? Well, I would say the tree-based is, think of it as generating the tokens in a certain way. and then how we interpret we use the distributions that were generated and the target distributions to decide which ones are acceptable which tokens end up being acceptable and there's actually an interesting there's a few different ways to do that and we had to as i say um so even within the tree-based generation methods of doing the rejection sampling there's a few different ways to do that yeah i'm envisioning like the you know something analogous to like a binary surge or something like you have all these leaf nodes and you need to figure out the most efficient way to qualify or disqualify, you know, all of these leaf nodes based on their, you know, their root, their parents.
1:03:21Something along those lines? Yeah. So the target, so now let me say the target model now, it actually inputs this tree rather than just this, you know, linear set of tokens. And it can process in the tree, our typical tree these days is 32 or 64 tokens that we are processing on one pass of the target model. and and so it creates it creates now there's a lot of what-ifs now on each of those paths and there's distributions created and and given how it was generated you have to do the statistics the rejection sampling like it just uses the the draft model statistics the target model statistics together to make that to decide which ones to accept to give maximal throughput while still maintaining the original distribution.
1:04:07Yeah. Very cool. And maybe one more quick thing to say. Then what we realized is, at first we were doing the simple thing that you said, which is user branching factor of two or three or something. Okay. Just as sort of some fixed thing. But then we realized, actually, you can do likelihood of paths, not just on a per node basis. What you're really doing is generating, like the tree can be a like it doesn't have to be a you know like a fixed two oh a symmetrical yeah yeah it can be effectively think of it as pruned in some sense you're saying all this pruned during creation like yes self-pruned or something yeah okay that's yeah that and then how do you handle the how do you handle that at the target level um yeah a lot of things to say there yeah i'll just say something we end up uh be going back more to the practical side is our hardware does symmetric things nicely like why gpus are like this too and because when you parallelize well it's easier if there's some kind of nice symmetry so there's some interesting trade-offs there from a practical point of view but yeah but we so we so we called that recursive speculative decoding and i will add um whatever the cause and effect is here there's actually a lot of this tree-based speculative decoding out there now and um and and we continue we continue our developments and and and the gains we see what i said maybe typically in a real realistic scenario we might see 2x from speculative we're trying to approach 3x with these methods and see a nice like 50 % gain okay but but it's very much diminishing returns um this uh conversion of compute into bandwidth is how we think about it um it does not take away the need for more bandwidth in the long run um but it's just one of those yeah you know it's one of those things that's So, you know, like on a communications channel, no matter what you're getting from a raw bitrate, throw some coding on, you'll get some better bitrate.
1:05:57But, you know, there's a limit. You still need higher bandwidth eventually. Are these tree-based approaches to speculative decoding kind of the end of that line, at least so far? Or are there other techniques that are being explored? There's another whole direction that is still in the speculative family that I think I'll mention quickly. one of the downsides of speculative in general is we talk about draft models you have to have this draft model and you have to create it you have to train it however and that that's actually a pretty lengthy process and then but you know that that seems to be okay especially for deployments that have you know persisted but but a real issue is very common these days is you take a large foundation model and you're going to fine-tune it you may be with lower parameters or something for specific tasks, leverage that big model for all sorts of things all at once.
1:06:48And then the issue is, well, this draft model we trained was for, maybe it was for the foundation model, or maybe it was for one of the applications. But the problem is it really just may not work that well for other things you would fine tune the model for. And this becomes, now it can be handled by just, you know, you can fine tune the draft model or you can figure out, you know, other ways. But we actually have another thing that came out of our lab that we call self-speculative decoding, which is you actually are leveraging the target model rather than having a draft model. You're adding, you're just augmenting the target model itself in a relatively simple way.
1:07:25And then the target model in the same step that it's verifying previously predicted tokens, it itself is predicting new tokens. So it's sort of self-doing it. We call it self-speculative. And what's really nice about this, It does add computational complexity, of course, which we have to some extent. And we can get about a 2x gain from it as well. But its real benefit is once you've trained it for the target model, then if you fine-tune the target model, you don't have to change these parameters. It just will naturally track because of the way it's leveraging the target models to generate them in the first place.
1:08:04So whatever you do to the target model will affect the prediction. and we find that that's actually really robust to the fine tuning. You can add LoRa parameters to the target model and the pre-trained self-speculative decoding parameters don't have to change and it still maintains it like a 2x game. Yeah, and maybe the last thing I'll say on the speculative stuff is we definitely are also, we're engaged in the research community and there's quite a bit of effort still ongoing in speculative. And one of the new ideas recently is to leverage embeddings from the target model and still do a kind of draft model prediction.
1:08:42It's almost like a combination of these two. And then there's definitely some nice gains. There's still some advancement going on in speculative methods and how much token rate speed up we can get. Now, one thing that jumps out at me is we've talked about all of this in the context of traditional LLM inference, but now inference is getting even more attention because we're finding ways to use it more to get greater reasoning abilities out of LLMs. I'm thinking about the O1 class models and now DeepSeq R1 and others that apply quote-unquote inference scaling to increase model performance. I'm assuming that makes all of the things that we talked about even more important.
1:09:36Yeah, this is definitely something on our radar, of course. We've had some efforts in these directions just to study what inference scaling, you know, how it's another way to think about using our excess compute to improve the final results. And I'll say this one way. there's a question of whether how much of this is like it definitely it definitely puts a premium on fast token generation and so all these methods that we you know the bandwidth limits the speculative methods that all the things that go into how fast we're actually generating tokens um because when you do the these these uh o1 type methods end up being some kind of tree search with lm generation okay but it's um but it also raises so if you can just do that generation and then you have some kind of um verification model that'll tell you okay we that let's back up that wasn't a good direction we'll search in this other direction um then how fast you can generate all those you know i think of it as like a depth first tree search uh you're kind of going down this tree of possibilities you're you can generate faster and faster okay with that then so that wasn't go this backup, let's go this other direction.
1:10:50But it also can be done with parallel generation of the different branches in the tree itself. And by the way, those of us, we've had some work in code generation where they do a lot of sampling, beam search, whatever they do in their generation, because the model is capable of generating many things and you have ways of verifying and pruning very nicely in that space. And so here as well, we're going to have some kind of verifier as part of the search process. And so it begs this question of the parallel, like breadth first versus depth first search in a tree effectively. Okay. And what interests me, at least in one context here is, well, we have these methods that we're going to burn our compute in order to generate tokens faster.
1:11:37That's speculative. But we can also use that compute just to do true parallel you know like the draft model we're talking about was doing these tree generations right do you can just pick this tree and have as many paths as you can support and how that plays out what is the best way to do this o1 reasoning type stuff in a hardware with our kind of limitations that's the problem that's really engaging us right now and too early to say what that ends up looking like or are there ideas in terms of what the direction what the direction might... I mean, it sounds like, you know, I mean, I guess you kind of alluded to them, like one is like doing things speculatively, right?
1:12:20And another is doing things in parallel. Right? So maybe, did you already answer my question? But they trade off. That's the only next step. Okay. Is that, because we do have fixed, you know, if I burn all my compute power doing speculative, it'll let me go down the tree fast. and I can back up as I need. And that's certainly one mode. But I could have been instead taking some of that compute and generating search in a parallel sense across the tree rather than straight down the tree. Okay. That I would know. You realize it's not like these methods. First of all, like the Owen methods, we don't know exactly what they are, right?
1:12:59And now DeepSeq we do and there are others and then there's things that we experiment with. So maybe back to the theme here, which is this is what we do. We're going to do the system. We're going to understand the system design. However we get there, we're going to do some ourselves. We're going to pull what we can pull from others' work. But then we have some real questions we're looking at. And then we'll eventually come up with this is what we think is the best way on our hardware. And yeah, we'll propagate. And these inference scaling reasoning models, the inference scaling is one lens of viewing those models.
1:13:36but there's also kind of a recurring reinforcement learning aspect to the way that they're trained. Does that come up at inference time or is that training only? Does it change anything about the way tokens are generated? Yeah, it depends. I actually, I think you can, it depends how you're verifying and how you're making the decision making at inference time. And I think that how you're doing that will definitely play into these decisions of where best to put the compute. That's really the best, what I even, yeah, where we are is that kind of question. Got it, got it, got it. Any additional thoughts in terms of what's next in future directions in this whole space?
1:14:25There's just definitely going to be development. And there's a lot of thinking, as I was saying earlier, on how our, you know, one of our major questions, strategic questions at Qualcomm is how our hardware is going to evolve. And there's definitely a lot of thinking going on there. And trying to, you know, try and try, and as I was saying earlier, what it makes sense to sort of dedicate to these kind of use cases or just the AI problem in general. Yeah. So that definitely will be coming. But it's also more... it's not as concrete as some of these other things. Like that will take time. Well, we'll also be including links to the speculative decoding papers and some of the other ones you referenced in the conversation.
1:15:12But thanks so much for taking the time to run through all this. Like it had a bunch of conversations with folks there and elsewhere about a lot of these things. And I think kind of running through it from top to bottom, help fill some gaps for me. So hopefully that's the case for folks listening in as well. All right. Thank you. Yeah. Very interesting conversation. And I should say it's a very exciting time. Maybe we all know. Awesome. Thanks so much, Chris. All right. Thank you, Sam.
1:15:57you
From the publisher
Today, we're joined by Chris Lott, senior director of engineering at Qualcomm AI Research to discuss accelerating large language model inference. We explore the challenges presented by the LLM encoding and decoding (aka generation) and how these interact with various hardware constraints such as FLOPS, memory footprint and memory bandwidth to limit key inference metrics such as time-to-first-token, tokens per second, and tokens per joule. We then dig into a variety of techniques that can be used to accelerate inference such as KV compression, quantization, pruning, speculative decoding, and leveraging small language models (SLMs). We also discuss future directions for enabling on-device agentic experiences such as parallel generation and software tools like Qualcomm AI Orchestrator.
The complete show notes for this episode can be found at https://twimlai.com/go/717.




