In short
The TWIML AI Podcast - Episode #697 Summary: Simplifying On-Device AI for Developers with Siddhika Nevrekar
Episode Overview In this episode, hosted by Sam Charrington, Siddhika Nevrekar, AI Hub head at Qualcomm Technologies, discusses the evolution and implementation of on-device AI technology. The conversation covers the motivations for transitioning model inference from the cloud to local devices, the necessary hardware solutions, challenges faced by developers, and the introduction of Qualcomm’s AI Hub platform designed to simplify the development process.
Key Topics Discussed
- Background of Siddhika Nevrekar
- Siddhika's journey from computer engineering to obtaining an MBA.
- Nine years at Microsoft focused on machine learning and cloud services.
- Six years at Apple working on machine learning for mobile devices, contributing to features like Face ID.
- Co-founder of Tetra AI, later acquired by Qualcomm, leading to the development of AI Hub.
- Motivations for On-Device AI
- Cost Efficiency: Reduces cloud service costs related to hosting and inference.
- Privacy: Keeps sensitive user data on-device, minimizing data transfer to the cloud.
- Connectivity: Ensures functionality in areas with limited or no internet service.
- Challenges for Developers
- Transitioning models from cloud to device introduces complexity due to diverse hardware and software environments.
- Developers face difficulties in managing different operating systems, chip architectures, and AI runtimes (e.g., ONNX, TFLite).
- The need for extensive testing across various devices and models, requiring significant time and resources.
- The Role of Hardware
- Discussion on the importance of powerful system-on-chips (SoCs) and neural processors in executing AI workloads efficiently.
- Advances in chip technology enable running complex AI models on devices, but memory constraints can still impede performance.
- Developers need to understand the capabilities of the hardware they are targeting and optimize their models accordingly.
- AI Hub at Qualcomm
- Introduction of AI Hub as a platform to facilitate the testing and optimization of AI models across devices.
- Developers can upload models and receive instant feedback on compatibility and performance.
- AI Hub aims to demystify the process, enabling even less experienced developers to implement on-device AI.
Key Metrics for On-Device AI Evaluation
- Performance Metrics:
- Time to first token, inference speed, and overall accuracy of the model.
- Power Consumption:
- Understanding power usage is complex but crucial; developers often rely on CPU/GPU utilization as proxies for power efficiency.
- Model Compatibility:
- Ensuring that models run across different hardware generations and operating systems.
Future Outlook
- The evolution of AI technology is expected to accelerate, with a focus on creating seamless user experiences.
- Potential for innovative applications of AI across various industries, particularly in IoT and autonomous vehicles.
- Continuous advancements in hardware and software will lead to more accessible and powerful AI solutions.
Conclusion Siddhika Nevrekar’s insights into on-device AI highlight the transformative potential of this technology for developers and end-users alike. With platforms like Qualcomm’s AI Hub, the barriers to entry for implementing AI on devices are being lowered, paving the way for widespread adoption and innovative applications.
For more details, check the complete show notes at [TWIML AI Podcast Episode 697](https://twimlai.com/go/697).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00This is where ML is going to always surprise that it doesn't stop innovating. It is constantly coming up with new things. So here we are, we were like, oh, finally, vision is solved. You can now take a photo and edit it. And we are on to, can we do things like crazy LLMs, which has vision and audio and some language and editing. Tell me what more I can pass together.
0:35All right, everyone, welcome to another episode of the Twinmill AI podcast. I am your host, Sam Charrington. Today, I'm excited to be joined by Sirika Navrakar. Sirika is head of AI Hub at Qualcomm Technologies. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Sirika, welcome to the podcast. Thank you for having me here. Excited to talk to you. I'm excited to speak to you as well. We're going to be digging into your work to bring on-device AI to the world and the various hardware solutions that enable this, as well as a bit of what you've learned in speaking to hundreds of developers, AI developers in particular, over the past few months.
1:21To get us there, I'd love to have you share a little bit about your background. Sounds good. I'll start way back in the past. When, you know, going a little bit back, I got my computer engineering, like a lot of folks in the field. And then I went on to do the MBA to just understand business aspects of how do you build a product, build a company, and so on. And then over the course of time, I went into Microsoft, where I spent about nine years, and that was all on ML and cloud. So my memories and learnings of AI is all based on what I learned building a search engine. And it was all about relevance and ranking and large and enormous amounts of data, ginormous models, hundreds and thousands and millions of inference queries responding to so many questions that people ask on a search engine.
2:13So it was all about that, going onto maps and navigated a lot of that. Was this on the Bing, one of the Bing teams? Yes, it was Bing Maps, Bing Shopping. I was very lucky, very fortunate to work on a bunch of verticals within Bing and overall ranking relevance as well. Then I took a deliberate decision of moving on to device because it was all cloud. It was all a world of services and servers. And I thought, okay, let me learn something new now and make myself extremely uncomfortable. And that's when I went to Apple, which is, this is the device world. And I got a chance to work on ML there. Very, very fortunate again.
2:52And to be honest, Sam, I did not believe ML on device existed. Like, that's just, that's not true. You know, that was my first instinct, sort of first thought that came to mind. Because again, these are large models. You've got to have so much data to train these models, be able to host them, be able to, and the model has to understand so much more. And clearly that cannot happen on the device. The six years I spent at Apple, that thought completely changed because my first year was when Apple Neural Engine chip was launching. And there was this big push on what is the key feature to showcase this?
3:29You know, it's the AI chip that goes into the device. How can we make the most of it? And that was the face ID moment where can we have phones unlocking with face? And that's a big deal because you have to recognize the face no matter what happens. You change your hairstyle, you change your facial hair, so many things that get older and the phone still has to recognize you. And working through that, that was the first experience. Working through that, building a platform for that, that made me realize that, oh, this is just the beginning. So spent six years working on a variety of use cases, going from detection on the watch to foil detection on the watch to cameras, which have 25 models and so on.
4:12And then left Apple, started my own company along with Krishna Shreder. I was the founder of Tetra AI. And that's where this beginning of, can we bring any model to any device started? I cut short, zoom in, got acquired by Qualcomm, moved into Qualcomm, and now built this product that we'll talk about, which is AI Hub. Yeah, I think your six years at Apple and your realization that, hey, this getting AI running on devices kind of actually calls them to focus a question like, why is this still a big deal? Why do you have teams there working on this? Apple's been doing it for six years. We've been using AI on our devices for several years.
5:01What are the outstanding challenges in that world? That's a really good question. The biggest challenge for an app developer is when an app developer decides, hey, I want to put an AI experience within my app on a device. They are essentially bringing the model which was on the cloud, where all the queries could go from all of their users, from all the apps, whether it's on an Android phone, on a Windows laptop. That's easy. You have one system to manage, one system to host a model which answers all the queries. But now this app developer is saying, okay, I have to take this model and put it into application instance of every user on the device.
5:37All these devices are very, very different. Taking a step back, why are they doing that? If they've got an application that's working in the cloud, why are they taking that next step of making them native to the device? The purpose is very, very clear. And these developers said they want to save money. That's number one. So most of the times people will start with privacy, but I'm starting a little bit different. I'm saying saving money because every time you send a query to the cloud, you're paying for that cloud entry, paying for hosting, you're paying for that scale. They can save all that money because really they're using the free cloud in everybody's pocket, which are the devices.
6:16Everybody has a device in their pocket and why not? That's the user's device. And if they can run the model there, they can save on a tremendous amount of cost. That's number one. Then you go on to things like equally important privacy. Sometimes you don't want the data all the way to travel to the cloud. We all know that, especially as we are getting into the world of more and more personalized experiences. I don't want to share the kind of queries that I do or maybe my private data with anybody else. So you get to give that assurance to your user base that your data is going to be secure and safe.
6:48The third big one is connectivity. Oftentimes, you have experiences, you know, you want to hike and there's an experience. You don't want, you know, what happened suddenly, the experience went away. You want your users to have that experience, no matter whether they have internet connectivity or not. technologies have advanced in all different spaces, including how powerful the chips are, how fast the connectivity is. Yet we see delays in our video conferencing sometimes, you know, you see the image being stuck. Oh, for whatever reason, the nomenclature is the signal is not strong and I can't do this because it's going to...
7:23Right, right, right. Four bars, three bars. Correct. And that is not a problem if the model is on the device. It's too close to you. You're doing things directly with the device, which is in your hands, and hence very fast. So some experiences require that blazing fast experience. And that's really what has brought this purpose of doing AI on the device. And we are realizing and finding more and more applications every day now. It sounds very much analogous to something that I see a lot in our community and other communities, which is folks wanting to run models on their laptops, like using Olama and similar types of tools to let you run locally.
8:07But I've got Olama running on this laptop here. I know how to pull a model down. I'm not sure I would even know where to start if I wanted to try to get something running on my iPhone, for example, or if I wanted to get something running on my local device. Yeah, you're absolutely right. I'll segment into, you know, AI Hub and slowly start talking about that. The main thing is now there are AI experiences happening all around us. And we see that because we see a Roomba in our house who's figured out how to navigate to all the spaces. You know, we see it every day and it might not cross our minds, but that's running AI.
8:42That's running some powerful AI. There are, you know, delivery services with little delivery robots. That's running AI. So it's become commonplace, yet we are not at a point where every developer, let alone developer, anybody can do AI on the device. How can we get there? Because that to us is going to be a moment where we will explore the kind of experiences that we can imagine. Right now, we are restricted by a few semi-advanced developers who are doing this. But how can we get this in the hands of the masses so we can build experiences which we can't even think about? That's the idea behind AI Hub, that the core challenge that we're trying to solve is what you mentioned just now.
9:22I don't know where to begin. And this is, your statement is very true. It even comes from the most experienced developers. I hear it all the time. Where do I even start? So we've got these developers, they're seeing cost, privacy, and other drivers as reasons to move their models to the device. Like me, perhaps they have no idea where to begin. And one of the things that you mentioned in there in talking about the motivation to run on devices that is blazing fast. When I think about these gigantic foundation models and running on device, it doesn't jump out of me that that would be a blazing fast experience.
10:02Close that loop for us before we transition to some of the ways you're enabling developers. The correlation of how fast something can be is, you know, directly related to how powerful the SOC is. That's one. And memory. And you said SOC, system on chip, the chip that's running the models. That's right. So what's happening now, interestingly, is we are starting to figure out how to get these large models, compress them, so the size is smaller, so that they can run on the devices. The chips on the devices are extremely powerful. They can actually execute really, really fast. if you get the model to them.
10:40But the memory is where we are getting stuck. So as the world has advanced, as technologies have advanced, we have made advancement in a couple of spaces. Now we hit memory. So when you run a Lama model, time to first token becomes very, very important. And you'll see that once the first token is generated, it's in the record of 10s and 20s tokens, depending on Snapdragon, where we are running, Snapdragon X Elite, where we are running this model. You can see that on a laptop, this experience of running a Lama is much, much more powerful and seamless. Same as Copilot. Copilot runs on the device and it's really beautiful.
11:16But taking that and going to a mobile brings constraint. You can't do the same. Mobile immediately becomes memory bound and you cannot have that same seamless experience. You get throttled, you get stopped in building an experience to run that token, produce the next token and run it fast. That's where we are hitting issues. However, there are techniques that have been used with quantization, with compression, with how we are kind of stitching various parts. You take this model, large llama model, you actually break it into four parts. You run one part after the other consecutively and you stitch them together.
11:54And there are lots of techniques that are happening now of how to seamlessly stitch them together and bring them together with powerful SDKs where software can take the burden and do some smart things on the devices to make this experience faster. Like you said, that's one of the pain points that everybody seems to be hitting. There's work going on. This is fairly new to us. ChatGPT has just about happened. So now we are getting into a mode where how can we do more? How can we do it faster? How can we bring large models? We are kind of entering that phase now. And I'll refer listeners to some of the conversations I've had with your colleagues on the research side about the various quantization methods and compression and, you know, quite a few conversations that I've had about some of the things that folks are doing.
12:48to enable these models to kind of fit and operate well on devices from a research perspective. But kind of getting back to, you know, the issues that a developer needs to think through, you mentioned SOC, you know, we're used to hearing about CPUs and GPUs. Now you're talking about SOCs and neural processors. Like what's the lay of the LAN from an on-device hardware perspective? A really good question. I think this is where most developers hit the biggest pain points, which is understanding what are the processes that are available on the hardware and how to best utilize them to run my model. What's happened even more is we've introduced AI chips.
13:34And these AI chips, they do some powerful math really, really fast. Matrix multiplication, processing of tensors. They are specialized to do just that. Their whole job is to take an AI workload and execute that workload extremely fast. But as a developer, when I'm trying to run my model on a variety of devices, I have no idea which device has a neural engine, this special chip, like you said, because it also was introduced a couple of years ago. It's not been there forever, like a CPU and GPU that we are more familiar with. So as a developer, I don't know which devices have it, which devices have the first generation of this special chip, which devices have the second generation.
14:15And with each generation, the chip has had more features. It has become more powerful. And I'm talking about not just about Qualcomm. Generically, around the globe, in all devices, this is true. Because we started with the first version where we understood what kind of ML we should execute on the first version of chip. And then the algorithms evolved. So we have two plates that are constantly evolving. One is the algorithm plate and the other is the hardware plate. Both are constantly moving and evolving. And the developers are stuck in trying to understand what works on what. And that's the part where if you look at the entire stack of how a developer goes from taking a model, which they build in PyTorch, to getting it running on the device, there are a variety of steps that they have to get to go through, which are not necessarily solved for or streamlined universally.
15:07So they have to learn a few things to actually take a model and get it running on the device. And that's the biggest pain point for anybody who's wanting to do this. What are some of those things that they have to learn? So let's take an example of an app developer who's trying to do this on mobile. And like I said, we spoke to 100 developers. So let's take, you know, there's a photo editing. Let's say somebody is building a photo editing app and decides to put a model in a photo editing app. For this developer, the model has to go within the app which runs on phones, which runs on tablets, which runs on laptops.
15:40And let's say there's some other proprietary device for namesake. They have to run this on not just these devices, but generations of these devices, because some of their users will have older laptops, some will have newer laptops, some will have newer phones, older tablets. Then they have to care about what OS is running on these devices. So the phones will run Android, most likely. The laptops will run Windows. Proprietary devices might run Linux. So now their model has to run with their application. Then the question comes up, okay, cool, if I want to run a model on Android, what is the framework that I should use on the device, which actually takes this input of the model and then executes it on the AI processor?
16:23So now they run into, should I use Onyx? Should I use TF Lite? Should I use something proprietary from the manufacturer of the chip? Let's say they finalize and they say, okay, for Android, I'll use TF Lite. For Windows, I'll use Onyx. And for Linux, I'll use something proprietary. Great. They're finalizers. They then run into a problem of, I have to take my PyTorch model and get it into a format, convert it, compile it to a format which this runtime on the device, the TF Lite or the Onyx consumes. How do I convert it? What do I use? As I'm saying this, you can imagine that their testing surface is increasing rapidly with every word I introduce.
17:05No combinations, right? And then in the end, they have to make sure that everything gets compiled, optimized to the point that the model runs on that AI chip on the device. That's the other challenge. All of these steps that I mentioned are learning for the developers. They have to figure it out, each one at a time. And across, you can see the explosion. As you think of the testing, it's an explosion of combinations that they have to go through. When you talk about the neural processors or neural engines, you described it as a module that's really good at matrix math. That's kind of what I think of as a GPU.
17:41Does the developer need to know which calculations to put on a GPU, which calculations to run in a neural engine? Or, you know, I would think that, I don't know if we're there today, but at some point, we just want software and APIs. Maybe it's PyTorch or maybe it's another layer to figure all that out for us and optimize. Are we there yet? We are to some extent. So these runtimes that I talked about, whether it's Onyx or DML or TF Lite or even some proprietary runtime, they do a good job of it. However, as you get into advanced applications, gaming, for instance, or heavy, large video processing AI experience, if you're building a random thought, a film kind of editing processing you're doing along with AI, at that point, you actually have to think very carefully about where does the processing workload go and where does the AI workload go?
18:38And you have to get into a little bit of more advanced work where you say that, hey, you know what, I want you not to touch the CPU or the GPU at all. When you have an AI workload, run it on the AI chip. Don't touch my GPU, don't touch my CPU because there's heavy processing of a video that's happening or some sort of graphic workload is going on. So don't disturb that. Otherwise, my experience will be completely disturbed. So there is a little bit of advanced work there that exists, especially for certain use cases. For most use cases, we've gotten to a point where the runtimes do a pretty reasonable job, pretty good job.
19:15I'll caveat that it works well only when the chip manufacturers also partner very well with these runtimes. So that's one of the key things to keep in mind that what the AI chip can do is very well known to the manufacturer of the chip. They know exactly what kind of math can execute. And they have to build those APIs, those integrations with the runtimes to say, hey, hey, runtime, when you tell me execute this workload, I will 100 % guarantee and execute that workload. And here are the kinds of workloads to make it very, very layman. Here are the kinds of workloads you can send me. Here are the kind of operators you can send me and I'll take care of that.
19:54So it's a dialogue between these two, the community runtime. I call them community because, you know, generally contributed by the world. Onyx gets, you know, TF Lite, there's contribution from the world. The community runtimes and the manufacturer of the chip need to have a dialogue constantly and to keep pace with this technology. You get new and new algorithms coming in every day. So they have to kind of stay in sync. Now, in one of our prior conversations, you mentioned that one of the dynamics that you see playing out in this ecosystem and why we're not further along is that there is no like single organization that's really driving things forward for, you know, AI kind of pan device.
20:38Can you elaborate on that take? Absolutely, Sam. There are so many choices. As I mentioned this, you could, you know, any listener of this podcast will think, yeah, I heard about Onyx and TFLite and, you know, there is DirectML and there is all of these choices. Each of these choices have come from a different origin. A couple of years ago, there used to be Keras and Cafe and so many other things we heard about. And TensorFlow was also a big one. TensorFlow is still very, very popular in terms of training in a lot of spaces. But it was started converging towards PyTorch. And this is from my, at least, personal learning and observation and being in the field.
21:20PyTorch became more and more dominant because of the ease of use, maybe community gathered around it. And now you had PyTorch and when you went to Android, your runtime was TFLite, which was built by Google and still is contributed heavily by Google, a fantastic runtime. But you have a dialogue of Meta starting of PyTorch and the community taking over, Google starting TFLite and the community taking over. And these two have to have a dialogue. Similarly, on the Windows side, Onyx came about, which was Microsoft's community offering and people contributed. And as world changed, as things evolved, and each one of these owners or the ones who came up with this concept, as they evolved, they went on to producing devices which were focused on.
22:12Surface came about. Pixel came about. So now your attention got, how do I make surface experiences better or pixel experiences better? So it just diverged a little bit and the burden of it fell on the community. And every time when the burden falls on community, a lot of choices come about. And that's what has happened in the on-device AI world. Someone who really likes TF Lite is going to contribute and take that and take a dependency and contribute more and more. But someone else is going to go after Onyx and build more and more. That is what has happened. So there hasn't been a, what is the one thing?
22:47Why isn't there one solution on every device? There isn't one solution on every device. And we are going to constantly see that for the near future, at least for a while. You mentioned all these steps that a developer, not really steps, but you mentioned all of these things that a developer needs to think about as they kind of phase through the process of getting their models on the device. at some point they've got something that's like kind of working. I think broadly in the industry, maturing from a testing and evaluation perspective, how do folks need to think about testing and evaluation in this very fragmented world that you described?
23:30That's a very good question. It's also a very hard question because as the devices are exploding, as there are more and more generations of chips coming with powerful functionality, you'll always have to think about the old and the new until the old phases out slowly. So a developer has to care about how do I test across all of these variety of combination of devices, OSs, runtimes? Now, how do I make sure that my model that I had for the previous generation of phones works for this generation and the next that's coming out? And they'll have to constantly care about that. And we have to, from Qualcomm's point of view, we think about this very, very carefully and hard because that is one key point for developers to adopt on-device AI.
24:16The more we make it easier for them and the more we guarantee that the experience are going to work consistently no matter which hardware, the more we'll see people exploring and expanding their journey into on-device AI. Today, it's a burden on them. Really, people have to set up a rack of hardware devices in their office or something to test this. And it's a very tough problem. There's no way around it as well. To some extent, you can say, okay, I'll have virtual machines and I'll test on virtual machines, which are form factor kind of similar. But at some point, you do have to test it on real devices.
24:50And that's where most developers struggle today because it's costly. It takes a lot of time. There's a lot of work that needs to be put in to guarantee that the model works on the device. That's also where AI Hub comes in. And we met at an event recently and I had an opportunity to hear you talk about what you're doing with that. And one of the things that struck me as super interesting is this idea that through this AI hub, developers can get access to actual physical devices that they can test their models on. And it really made me think about, you know, back in the days of kind of early, I don't know how early web, but like, you know, when we were developing web pages because of all the incompatibilities and browsers, you would, you know, sign up for these browser test websites and you would give it a URL and it would like fire up VMs and give you screenshots of your, your webpage running on, you know, various different types of browsers.
25:53I don't think we're doing that as much anymore. I guess, you know, maybe if you're, you know, a huge commercial app, you're maybe doing that. We also did that with like email clients back in the day before, you know, and that was very unstable. I remember like testing an email to see what it would look like on different devices and haven't done that in a really long time. But, you know, what you described made me think about those kinds of experiences. And, you know, to your point, you know, when you're at the point in time where the world is like very fragmented and confused and you've got a lot of different things that you have to design and develop for, like you need that kind of functionality.
26:31And, you know, I don't know if this would be good for you, but hopefully as developers, we don't always need this and you'll provide value other ways. But at least right now, I can definitely see how someone trying to navigate this world might need that. So that all being a bit of a segue, can you tell us a little bit about AI Hub and what you're trying to do with that? The thought of AI Hub came about when we launched this little startup. We were thinking, why is it so hard? AI is everywhere. Like you said, we see more and more people building it. But we still see people being hesitant. We spoke to 100 developers and every single one of them was talking about, oh, I have to have a rack of devices.
27:12This is hard. The testing is hard. But even figuring out whether the model will run on the device or not is hard. And I have to take out time from my development and do this, which again is a burden. Cloud setup is fairly streamlined. I can go keep that. This needs to change. We need to make it more and more accessible. That is where this thought of AI Hub came about. What if, and it was a very simple problem statement, what if as a developer, I could come with a PyTorch model, just a trained model in PyTorch. And I say that this is my model. I know nothing about devices. The statement that you made, I have no idea what it means to run AI on the device, but I'm going to upload this model to your service.
Read the full transcript
27:53And I'm going to tell you which devices I would like to run this model in. Can you give me an answer whether this model runs? And if it runs, how well does it run on the device? That was the initial thought and idea behind foundational principle of AI Hub. And that's what AI Hub does. A developer comes in, They upload a PyTorch model. They say that, hey, I want to run it on Samsung 24, 23, 22 devices from laptops to phones to tablets to drones, you name it. And we will take that model. We will get it. The platform will, under the hood, get it compiled and converted to the format that is going to be accepted by the device.
28:31Use the appropriate runtime. If you don't know about a runtime, we'll make the recommendation. We'll pick the right one for you. And we'll pick the one which will guarantee you the best performance and accuracy on the device. And we'll give you an answer within five lines of code and five minutes whether, A, the model works on the device or not. So it'll give you an answer saying, hey, this model works. You can bring this experience on the device. And if it does, it'll tell you how fast it is, how accurate it is. So your journey can go much, much more smoother. Your decision of moving from, I'm thinking of doing on-device AI to I'm doing on-device AI is sort of in a second.
29:07That's what we are going towards. And we hope people are going to build many, many beautiful experiences and surprise us with AI on-device. I guess I'm wanting you to kind of characterize this experience that we talked about of, you know, things being very different from one to another. Are you seeing a convergence and are the things that someone needed to worry about six months ago, the same things that they're worrying about today and the same things you expect them to worry about in another six months? Or like, how are those sands shifting? Things are definitely converging. Today, if you think about vision models, these vision models, the operators within the vision models, the mathematics within the vision models is fairly understood by almost all devices today because we have seen a lot of use cases in that direction.
29:55And we have made the community runtimes, their integration with the chips has been really, really, from a software perspective, from an hardware perspective, it's very robust. So anybody starting with vision today is an easier journey for them. But this is where ML is going to always surprise that it doesn't stop innovating. It is constantly coming up with new things. So here we are, we were like, oh, finally, vision is solved. You can now take a photo and edit it. And we are on to, can we do things like crazy LLMs, which has vision and audio and some language and editing? Tell me what more I can pass together.
30:32So we are going to always be in an innovative phase with ML. And that's a good thing. I think that makes it super exciting for all the pieces involved, from the silicon manufacturers to the software builders along the way. It's a very, very exciting time for all of us. When you think about the types of models and model architectures that a developer might want to put on the device, I'm imagining a spectrum, as you alluded to just now, from tried and true to kind of bleeding edge, just read about it in a paper and wrote the code myself to implement it. Can you talk about both from the perspective of the underlying hardware, the neural engines and framework software and the AI hub or testing infrastructure?
31:22Do I need to wait for cutting edge to be supported or do all of these layers have the ability for me to kind of roll my own when I need it? I don't know if batteries included, not required or whatever the right analogy is here. works for this, but how much flexibility does a developer have to implement things on the cutting edge or are they kind of trapped behind some abstractions and need to wait for the lower levels to evolve for them? That's a very, very good question. I'm sure many developers have the same question. But to some extent, there's flexibility. So I'll give you a very simple example of somebody came up with phenomenal algorithm for languages or for audio to text translation or something like that.
32:09And there's an operator, you know, it's really a mathematical equation. They're factored in this one single operator, which is brand new. And they introduced it in PyTorch because they can't. They can come up with a brand new operator. Now, everybody along the way doesn't know about this operator. So when it comes down to the runtime, runtime goes, oops, what is this? This is new. And this happens today. This is not future thing. This is happening today. What you can, of course, the runtime has to introduce it. It then goes down to the chip, to the SDK and the chip. And the chip also, the SDK has to tell the chip, this maps to, you know, these equations on your side or kernels on your side.
32:47They both also need to have a dialogue. And for them too, this is a surprise. Whoa, what is this? This is a new operator, right? Today, what most developers do, they can deconstruct this operator into what is known to all of these pieces. The flip side of that is if you deconstruct, you pay some price there. Either the speed will be impacted or power will be impacted. There is a trade-off. There's a price to pay until it's natively supported all through. And this native support can be, hey, the silicon itself has to evolve or the software that is communicating with the silicon has to evolve. Something has to evolve and we have to wait a little bit sometimes.
33:28The good news is most of the advanced algorithms are doing exactly this today because they want to, they know that the software will evolve. If it's within the control of software, it's a cycle of release of two months, maybe even 10 days sometimes, but they know software can evolve. If the silicon has to evolve, they can actually point the silicon in that direction and say that, hey, this is the next magical experience coming down the pipe. That's where we have to go. So it's a good thing to have a little bit futuristic, but having the ability to do it today, like you said, can I still experience that this is possible?
34:02That throughout the AI stack is available today for most of the developers, and they do that often. You mentioned power in there, and that's something that's clearly important for folks that are running on device. You wouldn't want your app to drain your user's batteries too quickly. We talked a little bit about kind of testing and evaluation via AI Hub. What are the kinds of metrics that are testable or measurable when you're operating in the environment? And I'm imagining, you know, you mentioned like just working, like does the thing work or does it not? So accuracy and correctness, presumably some measures of performance, maybe latency, you know, powers.
34:46Is that one of them? Can you talk about the various metrics and the way you see developers trading these things off as they're trying to build experiences? What the developers see today are following metrics. You know, what's the first upload time? How long does it take for first token if it's LLM? How long does an inferencing for a certain number of samples take? So I can do some sort of a load test to some extent. So time to inference, accuracy, how accurate is it? If I give you a sample, can you give me what accuracy you see? Do you detect if I'm doing detection? Do you detect that? Those are important.
35:22Power is one metric we do not show today. It's very case by case. We work very closely. And the reason we are doing that is because we want to actually build that feature really, really thoughtful. Power is measured in a whole bunch of different ways where you need to have a certain setup for the device, a cold environment, a warm environment. You actually have to have a lab which does replicate real-life temperatures as well. There are external temperatures. You might find this fascinating, but there are external temperatures that we have to create as well. You need to have a stable setup to do that.
35:59And then you have to measure the power while five apps are running. What happens? So in other words, unlike some of the other metrics that you described, power has a bunch of externalities that don't necessarily lend themselves to just running your model in a device farm because of, it sounds like, noisy neighbor or other things running on the device, power, et cetera. That's right. You need to keep in mind some external factors as well as internal factors to the device as you do power measurement. Are there like proxies for power? Like is performance or like CPU utilization or NPU utilization a decent enough proxy for performance that it will get you to a model out the door?
36:45That is a really good question. And that is accurate as well, where in the absence of having true power metrics, that's what developers look for. When my app is running, when my model is running, how much is the GPU, CPU utilization? Where is it going? How free is the AI chip? How much is it? And that is why I'm moving the workload to AI processor is what everybody wants. Literally, everybody wants that. AI processors are designed to be extremely power efficient. So everyone is looking for, why is that one operator getting dispatched to the CPU? I need to do something about it. I need to move the full workload onto the AI chip.
37:20That is the proxy they use in the absence. Yeah, I'll say that I don't think I really got that until I saw some of the demos at the AI day that I referenced earlier. These were demos in particular of the laptop devices. I forget, are they co-pilot devices? or... Snapdragon XE, yeah. Copilot devices. The Snapdragon copilot laptop devices. There were a number of demos of like, you know, really heavy-duty inference. And you could see that on these devices, it was pinning the NPU, but your CPU and GPU were not affected and you can continue doing your work. That made it like viscerally real for me because I know when I'm doing like heavy-duty processing, saying, you know, you can kind of tank your device and you can't continue working.
38:10But with those workloads shifted to these dedicated units that don't get in the way of your everyday compute, it not only allows them to perform much better, but for, you know, heavy duty inference does not, you know, interfere with the usability of the device. It was kind of eye opening for me in that regard. It's fascinating, isn't it, Sam, that now when we use, I don't know if you've experienced this lately, probably never when you kind of use your mouse and it's laggy or slow or the window doesn't go it's just it's so unacceptable now but that what you describe is the reason why it's unacceptable because we have done such a good job with these processors leaving uh you know the processor free to do the right thing and not disappoint the user and that has become very important for all the developers that i i don't want to that's that's the death of any machine or any application that oh it's so slow it just doesn't open uh so our conversation thus far has focused on mobile devices you know smartphones laptops to some degree uh but you and qualcomm more broadly also do a lot of work in iot and autonomous vehicles uh and other kind of specific fit to purpose types of devices can you talk a little bit about the degree to which you know do all these principles apply equally?
39:36Are there specific things that IoT developers need to worry about, that AV developers need to worry about? These segments are so different, Sam, that I'm glad you asked this question. IoT is, you know, when you think of IoT, some people even ask me, what is IoT? Internet of Things, what is that? What does that mean? Truly, we have clubbed, you know, everything that's not mobile, not tablet, not laptops is IoT. It's kind of, you've done that in some sense. But IoT has some interesting challenges. The good news is most of the, when people are using IoT, building for IoT, they typically control the entire experience and the entire, right from the OS to what goes into the device, all through and through.
40:17And you think of a drone, right? The drone is going to have the OS you decide, all the bells and whistles that you decide, the software you decide, the SDK you decide, the runtime you decide. So it's in some sense, end-to-end, it offers end-to-end control. The flip side of this is some of the IoT devices might be really, really tiny. So you get into this world of one-time-use devices, a couple of time-use devices, and now you want to, on top of that, run AI on these devices. The biggest challenge there is how do we continue to expand the length, longevity of these devices, which are already as is tiny, small, cheap.
40:54How do we build chips for these that are actually long-lived and powerful to run AI. The overall principles of how I get a model and how I get it running remains the same. The challenges remain the same, but the form factor has suddenly changed and hence the kind of workloads you can run on these devices definitely change. In contrast, a car is a ginormous computer. So everybody is super happy about that because Finally, we can have a large computer sitting in here. Yay. Now you can talk about AI on device is sort of what comes to mind to many, many people. So then there you go into experiences in ADAS, voice detection, everything that's going on with autonomous.
41:41It's a huge machine that gives you a lot you can do. But at the same time, there is a lot that needs to run right there. Not just AI, but you have so many sensors and they're constantly collecting data and somebody has to be processing that data and then you will take decisions on that data. That itself is a ginormous task. Even if there is this large computer, you can imagine multiple computers, they're doing many, many things to keep you safe in the car for self-driving. So those are the two form factors. Intelligent, you know, very, very interesting and intelligent use cases are of the one I have come across is a drone that's sent into a mine before a human goes in so that it can map out a mine.
42:24And it does use a lot of AI. And there you have a constraint of the drone having no connectivity as well as limited power. So it has to do its entire job, come back to the docking station and charge itself, just like our Roomba. We get disappointed to see Roomba lying around, not charging. that drone would go through that. So that's, I think, mainly the form factors are different, the capacities are different, but the software principles remain exactly the same, which is good news for us because you can replicate from what we have learned in the mobile and the laptop worlds, the tablet worlds, and bring them into these worlds.
43:00So along those lines, maybe setting aside AV, the use cases that are most apparent for IoT tend to be vision-oriented. Have you come across interesting ones for LLMs? There are few, very, very few that are cropping up now. And the reason, like I said, I think the reason we call them IoT use cases, because everything that you don't put in these segments, you put in IoT, you don't put in mobile. It's kind of a slush bucket of... Correct. So now you have a kiosk, and on the kiosk, you have phase detection. because Sam goes to that kiosk every day and gets a power bar. And the kiosk has learned. When Sam comes in, 10 o 'clock, power bar, you know.
43:41What if Sam wants to talk to the kiosk and tell it what he wants? And so that's a use case. That's what it's evolving now. And you will see some of these use cases appear. Different form factor. But we are seeing some LLM use cases slowly appear in the IoT world as well. We're at such an interesting time. every space is evolving. That's why I think there's this general phenomenon where even when we from Qualcomm go on to talk to customers, the first question is, and how do I use AI in this? You know, like everyone's asking that. So it's every use case is evolving. Everywhere there is going to be more and more things done to save our time.
44:21It mainly comes from a place of how do we give back time to humans so they do something better, bigger, and they don't have to do this road thing. You don't have to put in a bill and punch the numbers and get that power bar out. You can just talk to it. Talk to the kiosk. Be behaving fast. So in a world or an environment where everything's changing so quickly, do you have a sense for what the future holds? To me, that's going to be, that's something I'm really, really looking forward to seeing. What is going to be the next wave? We have seen the wave of AI come and everybody's talking about AI is here.
44:57What we see next is, I think, feels like a sci-fi movie. That's what I imagine is going to happen. I still don't have a robot who can fold laundry. And boy, I want to. Really, that's why. So I think the world... I saw one at CES a few years ago. It didn't work very well. Yeah, a good working one, you know, which actually does the job. But I anticipate more advances in silicon is going to happen, more advances in software stack is going to happen. And we're going to reach a point where we see just so many use cases that are up and coming in the AI field. Well, Siddhika, thanks so much for taking the time to chat about on-device AI and the work you're doing to support that.
45:45Thank you. This is fun. Thanks so much.
From the publisher
Today, we're joined by Siddhika Nevrekar, AI Hub head at Qualcomm Technologies, to discuss on-device AI and how to make it easier for developers to take advantage of device capabilities. We unpack the motivations for AI engineers to move model inference from the cloud to local devices, and explore the challenges associated with on-device AI. We dig into the role of hardware solutions, from powerful system-on-chips (SoC) to neural processors, the importance of collaboration between community runtimes like ONNX and TFLite and chip manufacturers, the unique challenges of IoT and autonomous vehicles, and the key metrics developers should focus on to ensure optimal on-device performance. Finally, Siddhika introduces Qualcomm's AI Hub, a platform developed to simplify the process of testing and optimizing AI models across different devices.
The complete show notes for this episode can be found at https://twimlai.com/go/697.




