In short
Podcast Summary: AI at the Edge: Qualcomm AI Research at NeurIPS 2024 with Arash Behboodi - #711
Overview In this episode of The TWIML AI Podcast, host Sam Charrington interviews Arash Behboodi, Director of Engineering at Qualcomm AI Research. They discuss Qualcomm's innovative contributions to the NeurIPS 2024 conference, focusing on differentiable simulations, conformal prediction, and efficient machine learning models for mobile devices.
Key Topics
- Differentiable Simulation in Wireless Systems
- Definition: Differentiable simulations help optimize parameters in complex systems (like wireless antennas) using simulators.
- Forward vs. Inverse Problems:
- Forward Problem: Given a model, predict real-world behavior.
- Inverse Problem: Given data, infer the underlying model.
- Challenges: Traditional simulators can be either:
- Statistical: Fast but not physically accurate.
- Ray-Tracing: Accurate but slow and complex.
- Workshop on Data-Driven Differentiable Simulation Surrogates
- Goal: To bridge gaps in current simulation accuracy and efficiency across various fields.
- Topics Covered:
- Use of machine learning in drug discovery and molecular dynamics.
- Handling uncertainty in simulations.
- Conformal Prediction and Information Theory
- Conformal Prediction Framework: Offers a method to quantify uncertainty in model predictions beyond point estimates.
- Provides sets of labels that likely contain the true label based on calibration data.
- Research Focus: Connecting conformal prediction with information theory by analyzing how uncertainty can be understood through concepts such as entropy and conditional entropy.
- Recent Papers on Efficient Models for Mobile Devices
- LoRA (Low-Rank Adaptation): A method for fine-tuning large models with reduced parameter overhead.
- Key Innovations:
- Hollowed Net: Adapting LoRA for on-device personalization.
- ShiRA: Sparse High-Rank Adapters that allow efficient switching between models without high latency.
- FouRA: Fourier Low-Rank Adaptation to improve diversity in generated images on mobile devices.
- Robotics and Compositionality
- Clever Skills Benchmark: A dataset to evaluate how well models can perform tasks composed of basic operations.
- Live Fitness Coaching Dataset: Focuses on asynchronous interactions between AI and humans, enhancing user experiences in fitness applications.
- NeurIPS Expo Preview
- Arash previews the Qualcomm Expo talk, highlighting several demos:
- Mixture of Experts (MoE) Caching: Efficiently managing model loading on devices.
- On-Device Video Editing: Demonstrating the first video editing diffusion model capable of running on-device.
- 3D Scene Data Generation: Creating realistic driving scenes for augmented reality applications.
Future Directions
- Discussion on the scalability of AI technologies and the future of agentic AI systems.
- Potential exploration of hybrid approaches combining neural and symbolic methods to tackle complex reasoning tasks.
Conclusion The episode offers insights into Qualcomm's pioneering work in AI research, particularly as it relates to mobile and edge devices. With advancements in differentiable simulations and machine learning applications, Arash and Sam provide a thorough overview of the upcoming challenges and opportunities within the AI landscape.
For further details and complete show notes, visit the [TWIML AI Podcast website](https://twimlai.com/go/711).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Suppose you want to optimize your antenna within a limited space to get a certain propagation pattern. that's something that you need to go through a simulator, see how it performs, and then iterate around that. And experts usually do that kind of mentally. They know what to exclude and where to start, but still doesn't solve the problem that iteration is quite similar. Therefore, surrogate models for differential equations, they become important exactly for a similar reason.
0:40All right, everyone. Welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. Today, I'm joined by Arash Beboody. Arash is Director of Engineering at Qualcomm. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Arash, welcome back to the podcast. Hi, Sam. Glad to be here. It's great to have you on. It has been a little while. I think we were speaking about ICML 2022 last time we spoke. Wow, it seems like such a long time ago. So much has changed in the world of AI. Why don't we get started by having you share a little bit about what you've been focused on recently in terms of your research?
1:26Yeah, exactly. It's unbelievable how much things have changed during past years. Yeah, well, last time we talked, we talked about inverse problems and all sorts of techniques that we can use for solving them and particularly using generative models and generative priors. And for what we discussed back then, it still continues to be one of our main interests, But the forward process, which is like what we observe, was quite a nice tractable, let's say, linear transformation or something like that for most of the cases. But for a large class of inverse problems, basically, which includes a lot of things that we try to figure out in our day-to-day life, the forward process is a very, very complicated system.
2:18And we use simulators, for example, to simulate various physical phenomena, various activities, and so on and so forth. And having good simulators with certain desirable characteristics is very, very important as well. And if you really want to solve an inverse problem through a really complicated simulator, it's a challenging task. So, you know, we should probably level set on forward and inverse problems. So an inverse problem is when you've got some data about the world and you want to extract the model from that. And forward is when you have a model and you want to kind of project that into real world behavior.
2:59Is that the right way to think about those two? Yeah, I would say so. I would say so. So, you know, very simply, you can look at the, we talk about the function, right? So you have a function that the forward part of that function basically is your forward process. And the backward part is the back. You're applying that function to some data, right. Yeah, I have this kind of mathematical thinking always. So for me, it's just a function, I would say. Okay. Yeah. But yeah, so I've been focusing on that research quite a lot during past years. And, you know, of course, like wireless being one of the main, you know, focuses of Qualcomm.
3:47I've been like really looking into different wireless simulators and the way we simulate wave propagation and those things. So it's been like my focus during past years. What's the intersection between the, or remind us, I think we talked a little bit about this last time, of the intersection between the forward and reverse functions that we talked about and wireless systems. Very, very good point. So, you know, when we're building wireless or wireless systems, starting, let's say, from your antenna design, from the wireless system design, on all sort of parameter tuning, we need to be able to understand how will our algorithms perform, right?
4:27and at the end of the day we're working with in wireless communication we're working with electromagnetic waves and electromagnetic prepropagations so we want to be able to simulate those those phenomena effectively that is the that is the goal now in wireless we have like different class of simulators so if we focus on let's say the more what we call the link level simulation which is about like a point to point communication we have on the one hand On the one end of the spectrum, we have the statistical simulators, which kind of statistically, on average, tells you how the propagation looks like.
5:04And on the other end, you have physically consistent simulators that, based on some ray tracing principles or even solving electromagnetic Maxwell equations, try to give you the exact propagation in the environment. So that's the whole spectrum. And you use that to solve different problems or vet your algorithms. your positioning algorithms, right? Your throughputs, your latency, all sorts of communication metrics that we have. So that is roughly the class of simulators we have, but they have their own challenges, right? So if I list some of these challenges, I can say, of course, statistical simulators, usually they are not physically, spatially consistent, meaning they're just a statistical model.
5:54So it's not like they model some concrete trajectories necessarily or the impact of certain buildings or certain objects in the environment. But they're fast. So you sample from them fast so you can iterate around them. On the other end of the spectrum, ray tracing-based simulators that give you quite detailed propagation information, They generally, of course, first of all, they are based on some mathematical models of physical phenomena, of diffraction, scattering, reflection, all those things. But they are also quite complicated to set up. They're professional. Their company is actually selling those services, professional softwares for that.
6:44And you really have to go into that. and then at the end of the day, you end up with a complicated simulators. What is common between all these simulators is that if you, let's say you want to solve an inverse problem through these simulators, for example, you want to figure out where should I put my access point, my home, to get the best coverage. It's a common problem that many people will do. This is a typical inverse problem. You want to find the optimal location that, given the 3D map, provides the basically best coverage in whatever sense you define. And if you don't have a differentiable model, you just have to find heuristics of iterating or use some sort of Bayesian optimization or those type of methods.
7:26So our focus has been trying to find different operation points within this spectrum. Maybe we can sacrifice some of the accuracies that we have, but we get faster differentiable versions of it. Or maybe by building specialized simulators, you're able to get accurate, but since it's specialized, you already have low complexity for that. How can we inject inductive biases into the simulator? We're already processing 3D space. So how we can leverage that 3D space and simulators of that 3D space in building the simulator. So we had like a range of papers during last years on all these different topics and covering all this.
8:14I should have mentioned that we're going to be focusing our conversation today on the work that you and your colleagues are presenting at the upcoming NeurIPS conference. And one of the activities that you've got planned is a workshop that you helped organize on just this area that we're discussing. There's a lot of overlap there. Data-driven and differentiable simulation surrogates and solvers is the title of the workshop. But it's not just application to communications. It sounds like you're applying these techniques or you're inviting folks to come talk about the broad application of these techniques.
8:57Can you share a little bit about the idea behind this workshop? Yeah. Yeah, so what, you know, the D3S3 workshop, which is the kind of the abbreviation of all these names. So that workshop, the goal was actually to get together, you know, all different people in the machine learning community who are working on similar set of topics. So if I, you know, take a step back and, you know, keeping the discussion we have about wireless in mind, if I take a step back, there are certain problems we're dealing with. We're dealing with a problem of the interreal gap. Basically, we're not accurate to representing reality and how we can leverage data to close the interreal gap.
9:44We have the question of basic accuracy versus complexity, accuracy versus latency trade-off. And then we have the question of differentiability that enables solving all sorts of inverse problems. So a similar set of problems come up in many different areas. And many notable works coming up recently. We have heard of drug discovery and molecular simulations and works around that. We are aware of the work that came, I think, one or two years ago on ForecastNet, which is about weather prediction using neural operators. So that's fully neural operators. So that's one version of that. Then we have Unisim, which was the outstanding iClear24 paper by Sherry Young and colleagues who was actually on your podcast as well.
10:40It's about interactive simulators that can be used for robotics and all sorts of applications. It's a video diffusion-based simulator. And then you have very quite recently a foundation model on basically the earth system, which is about all sorts of climate related behaviors and predict those. So we have all sorts of things. And actually a lot of folks who are co-authors or lead authors in these papers, either are speakers or organizers of the workshop. So we want to bring together different people working on similar topics to be able to discuss about these challenges, how we can address them, what are the challenges we're facing, what are the lessons learned from previous experiences, and things like that.
11:36One thing that I think will be a good segue for a later discussion we'll have is uncertainty. You know, we know that if the simulators are not perfect, how we can characterize their uncertainty and use them in our design process, right? So that's a very, very good, important question. And like discussing about that can be interesting. One question that I wanted to ask is related to this idea of differentiability and differential simulators in the context of this workshop. There's two concepts of differentiability. One is like the differentiability of your machine learning models and equations that allow you to do your gradient updates.
12:22And another is the idea of differential equations and modeling those using machine learning. Um, are you talking about one or both of those, or, uh, you know, how does that, how do those two ideas, I feel like both of those ideas are coming into play and what we're talking about. Um, but maybe elaborate on that for a sec. Yeah, yeah, absolutely. Actually, it's an excellent question because differential equations, they are mathematical models that try to capture certain, you know, physical system if you want. Right. and they're an example of a complicated simulator. You know, if you want to solve, let's say, the exact differential equation that represents a certain dynamics, those are molecular dynamics, it can be quite tricky to solve it for a large set of simulation scenarios.
13:20Meaning you've got some differential equation that has a lot of parameters and you want to map that to some set of data. Exactly. Yeah. Exactly. And just assume that you have to iterate on that. Let's say you have to iterate on some configurations and figure out what is the impact of configuration on the ultimate KPI. Again, let me give you an example of something familiar in the wireless domain again. Suppose that you want to optimize your antenna within the limited space, and you have to figure out where you want to put your antenna to get certain propagation pattern, right? That's something that you need to make an iteration, go through a simulator, see how it performs, and then iterate around that.
14:08And experts usually do that kind of mentally. They know what to exclude and where to start, but still doesn't solve the problem that iteration is quite slow. So therefore, surrogate models for differential equations, it comes in different shape and form. they become important exactly for a similar reason. And you have this kind of a physics-informed neural network is an example of that. You have graph neural network-based physical surrogates with all these equivariant symmetries. Neural operators is another way of parameterizing functions and building surrogate for these differential equations.
14:52So all these topics are very much related to, you know, differentiable simulation there. So that is about differential equation part of it. One thing I want to add on the differentiable simulation, just implementing things on PyTorch or having a, you know, neural network there might not be sufficient for solving an inverse problem. Like very, very, one easy example is that if you just have a step function or a quantizer function, the gradient of that is going to be zero. So when you're back propping toward that, you will not get useful gradients. So you have to be cautious about just implementing your simulator in PyTorch doesn't solve differentiability problems.
15:32So you need to understand, you know, how you, you know, if your gradients are useful and take a look at that. And these are, these are actually quite, quite important problems as well. So two distinct concepts, but you're kind of talking about both of them in this workshop because of the importance and prevalence of differential equations in modeling real world phenomena. Exactly. And so your team or the various teams at Qualcomm had quite a number of papers at this year's NeurIPS. You are co-author on one that we'll dig into, and that one is related to a topic that you just brought up, the idea of uncertainty.
16:15And the paper is called An Information Theoretic Perspective on Conformal Prediction. A ton to unpack there. Let's maybe start with a broad brush overview. What are you looking at with this paper? What's the core question that you're attempting to ask here? I think the starting point for us was we have the framework of conformal prediction, which I explained shortly. We have framework of conformal prediction on the one hand, which provides a qualitative way of characterizing uncertainty. basically instead of a single point it gives you a set of labels that contain your true label, right? It is qualitative on the other hand you have the information theoretic way of characterizing uncertainty based on let's say conditional entropy and entropy and those concepts.
17:10So the question for us was how these two are connected to each other are we talking about two different notions of uncertainty or are they closely related and what we've done in this work, we connected these two concepts through the lens of information theory, and particularly list decoding. So that's roughly the question we had in this paper. And of course, it has consequences that I will talk about later, but that was the starting point. I'd love for you to elaborate a bit on uncertainty from an information theoretic perspective and broadly. Why is it important that we model uncertainty and what is the connection between entropy and uncertainty?
17:56How do you think about the broad space of uncertainty? Of course, we know that for a lot of use cases, it's always good to know that how much uncertainty the model has about the output it's providing. And Particularly, you can have a very common justification of safety, right? You want to make sure that the model is certain in what it's telling you so that you can use it based on that, right? That is really a rough idea. So we can think of this as like if an LLM could say, you know, I'm only 20 % confident about what I just told you, then we can use that to fight things like hallucinations. But we don't have that.
18:44It's very hard to do. And so, you know, the area that you're exploring is how do we get those kinds of uncertainty bounds for different types of models, not necessarily LLMs. Correct. Correct. Exactly. So that's one, that's a good example. So that's like why uncertainty matters. The other thing is like what information theory has to say about uncertainty. So in general, when you look at the original notion of entropy, there are different ways of deriving or introducing notion of entropy in information theory. There are certain just axiomatic version of defining the entropy. And there are other ways that just say, okay, entropy is just expected value of minus logarithm of probability, something like that.
19:48So that's how you find it. When you look at the axiomatic version of entropy, it actually is quite intuitively related to the notion of uncertainty. It basically, let's look at like a binary case. You have just a binary random variable. The axiomatic version, it has something, okay, the entropy is going to be positive. The function I'm looking for is going to be positive. It's going to be bigger than between, you know, it's going to be positive. And then if you have two independent random variables, their entropy is going to be sum of their entropy. And you'll write down those, which was quite intuitive from, you know, the common sense entropy notion.
20:34And if you solve that functional equation, already you can see that if you have the, you know, if you write down those axiom, it kind of gets leading you toward the logarithm function, right? So suddenly you recover the notion of entropy that is defined just like this on the paper, like expected value of minus logarithm. So that is roughly the idea. And then we continue using that as a notion for characterizing different things. For example, if the uncertainty you have about something, knowing another thing is characterized by conditional entropy. Now, interestingly, if you know something and you know that the other thing is just a function of this, you don't have any uncertainty about the new thing knowing the previous one.
21:34And in those cases, the conditional entropy is going to be zero. And so there are all these very nice intuitive behaviors and kind of nice because we can always use the probability outputs of our models to come up with different characterization of uncertainty, whether it's epistemic, aleatoric, and all those things. So that's roughly the connection. Of course, there are a lot of caveats here. Some of the examples I made, it's for discrete random variables, for continuous random variables. Things are more subtle. But yeah, I think we don't want to bore the audience with all these details here.
22:15How does the idea of entropy in information theory relate to the idea of entropy in physics? Like the whole chaos theory aspect of it? Yeah, I think there are connections, what I can say. And there has been very fruitful interaction between a statistical physicist and information theorist in the past. Even there are, I think there are these stories that Shannon, when he wanted to name this notion that was later called entropy, He got the feedback from, I think, von Neumann and some others that this is actually very much related to the national entropy. I might be mistaken, but I remember such anecdotes.
23:02Yeah, the connections are there, but at the end of the day, we can look at this. So there are connections between, let's say, the free energy, free entropy and all those things. So there are connections between that if you look at those concepts and the notion of entropy as we have in information theory. And so conformal prediction. So conformal prediction has an interesting beginning. Of course, the idea goes back to, you know, work of Alexander Wolfkes, a student of Alexander Kolmogorov and like that school. But the underlying idea is, I would like to present it like that, which seems quite intuitive to me.
23:50Suppose that, you know, you go to your market and you buy a classifier, right? You buy your classifier and then you look at the output of your classifier, right? You give images and you get the labels and it gives you some number between zero and one. But, you know, you just bought it. You don't know if you should trust it or you shouldn't trust it. So those numbers doesn't tell you anything. It's just you might want to pick the maximum one, which might be the case. But what you do then, you say, okay, let me actually gather some data, which you call calibration data, which has my samples and two labels there, like my images and my labels.
24:34And then I look at how much the score of my classifier deviates from, let's say, in this case, let's say one, right? You can combine, you can build any score that you want, how much it deviates from one. And you look at your calibration set and you notice that actually for 90 % of my samples in my calibration set, for 90 % of them, I only have, let's say, 0.4 gap between 0.6 and 1, basically. My score is between them. So what you do later on, you know that if you look at the output scores and the scores fall between, let's say, 0.6 and 1, then you call that as your true label. So you know that this is a true label.
25:25And now, if you just have so many samples, in this case, mathematically, you would probably have a 1-1-2 sample that lies between 0.6 and 1. But let's say if that threshold is 0.3, you can have multiple samples that can pass that threshold. In those cases, you will have a set as the output. That's the idea of conformal prediction. Collect calibration data, then try to define some conformity score. And there are a lot of research on how you define this conformity score. And then look that, you know, we'll get some quantile that you choose, right? Let's say 90%. And is the conformity score the threshold that you've set?
26:07So the conformity score is not the threshold you say. The conformity score just shows how conformal your classifier is with respect to the behavior you, you know. So you expect it. But now you want to pick a threshold. So what you do, you say, okay, I've picked 90 % quantile, 90 % of the set. So let's pick 90 % quantile. And then you look at the threshold according to that. You say, okay, for 90 % of my data, that is my conformity scores, right? And you use it like that. Okay. So the interesting thing about conformal prediction is that it says if you randomly pick your calibration data, IID samples, and you follow this procedure, 90%, pick the quantile and compute the threshold like that, actually that 90 % quantile, that 90 % is going to be the probability, approximately the probability that your true label belongs to the set.
27:15So it actually comes with a statistical guarantee. That is the nice thing about conformal prediction. It's an average, if you want, guarantee across all different sampling that you can do, but it's still a very, very nice guarantee. It's basically the more precise expression is that your probability is going to be lower bounded by the quantile and upper bounded by quantile plus one over the number of samples plus one, something like that. So if you have a very large calibration set, basically you get the exact guarantee that it's the 90%. And you can pick that 90%. So that's the design that you can have based on the use case you have.
27:54If you're looking for 99%, this is what you get. So that's roughly the conformal prediction, and it provides a set of outputs labels that you can use for your prediction.
28:10Now, when we looked at this, we said, okay, how this is connected, how we can connect this to information theory? Like, what is the connection between all these entropy notions and this? And then we say, look, look at the output labels of your classifier. as an output of a noisy channel, right? It's just a noisy version of your true label. What you want to do, you want to build a list that contains the correct label. And this is list decoding, basically. And there is like a lot of work already, like from in 50s, there are many works on list decoding and establishing different properties of list decoding.
28:52And what we're interested in is variable size list decoding because we just care for the guarantee that we have. So the list size might change. And using information theoretic inequalities like funnel inequality and data processing inequality and all that, we can connect these two together. On the one hand, we come up with the conditional entropy, which is the uncertainty of the model, if you want. And on the other side, we come up with an expression that contains something related to the expected set size, basically, which is somehow measure of inefficiency of your conformal prediction. And that is the connection.
29:32So it's interesting. So you can basically translate a lot of these inequalities to bounds that connect conformal prediction and entropy. And is the benefit that you're seeking that you can use these inequalities to determine, you know, principled bounds for your conformal prediction problem? Yeah, that's definitely one of the main interests we have and we're pursuing. We're interested to see, like, you know, how tight we can get in terms of bounds. So that's definitely one. But there are other things that caught our eye when we were working on that. One, for example, was what we call conformal training.
30:21The idea is that during training, instead of optimizing cross-entropy, let's try to basically optimize the conformal objective. Right? So there were works that already started. basically they put the whole conformal procedure, conformal prediction procedure in the training loop and optimize basically the set size there. So, and there are all sort of like smoothing operation that needs to happen because there are a lot of non-differentiability there, but they try to optimize those things and that's the conformal training objective. What we did is that, okay, look, actually the bond we obtained from funnels inequality and data processing inequality, in general, they tend to be tighter than just the expected set size.
Read the full transcript
31:07So why not using them as the training objective? And we kind of went through the procedure and used those as the training objective, promoting efficiency of conformal prediction. And we showed some results. So that's one case. And just the idea there, as I'm understanding it is, historically when we talk about, and we alluded to this earlier, uncertainty and machine learning, It's like we've got these models and we want them to characterize their predictions with some type of uncertainty that is often challenging, usually challenging. Um, what I'm hearing you say here is that as opposed to kind of post hoc trying to extract uncertainty, what we're going to do is as part of the training of the model itself, bake uncertainty in as part of the training objective, meaning, um, you know, I'm training this model in such a way that I wanted to guarantee that its responses are, you know, 90 % certain or X, you know, whatever your threshold is, like some percentage certainty or some value of certainty and kind of bake that in to the model itself.
32:26Is that the key idea here? Yeah, that's a great explanation. Because, you know, when we're using cross entropy, this is a kind of a surrogate loss that we use to train the model. With conformal training, what we do is say to the model, look, you know, This is the uncertainty you have right now. Try to make it smaller. Try to make the set size smaller. Of course, we don't use that objective. Again, we use just some surrogate losses to smooth the operations of set size and all that. But nonetheless, that's the same idea. Yes, we want to reduce the uncertainty. Yeah, so that was one of the applications of these bounds.
33:10and the other one was side information. One thing that is nice about the notion of entropy and related notions like conditional entropy and mutual information is that the notion of side information can be very naturally injected into these operations. There are these chain rules in information theory that enable you to decompose your entropies and mutual information into different components. and we kind of looked into the problem of having side information. If you have side information, how you can leverage that to improve your efficiency or both during training and during infer. What's an example of side information in this context?
33:55Side information can be, let's say you have a classifier and you kind of someone tells you, you're actually classifying animals. It's not any plant or buildings or anything else. This is one example of side information. Another one is, let's say in feather-to-learning, you know the device ID. And that can be just the side information that you can inject into the weight aggregation procedure and all the things that you do. With this work, how do you think about evaluation? What did you benchmark against? And did you apply specific use case evaluations or are they broad evaluations? I think we applied on the image classification type use cases, kind of range of models that you have there.
34:48So basically that. And what kind of results did you see? In general, of course, the bounds that we obtain from our information theoretic bounds are tighter than just the set size that folks use for conformal training. So that is, of course, expected to translate into better performance. And we see, especially with the data processing inequality bounds and the variations of that, we see better gain in terms of efficiency, which gets smaller set sizes and all those things. So kind of a consistent view. And on the side information, it's kind of also obvious. We have more information contributes to get, you know, smaller sets and reducing your uncertainty.
35:37And we also see that behavior. So basically showing that our approach translates to what we expect. Let's maybe spend some time kind of reviewing some of the other papers from your team at the conference. over the past few years one of the topics that has come up quite a bit has been the has been related to kind of generative ai and you know there are several papers here related to that i think one of the ones on the list is looking at on-device personalization for text-to-image diffusion models.
36:17Let's maybe talk a little bit about those papers and what some of the key themes are. Sure. So let me actually, like we have a bunch of papers and let me kind of cluster them topic-wise. So of course, one of our main belief and goal is we want to be able to provide these models on edge device. So we want to be able to all these image generation models or LLMs and all that. Of course, as much as possible, but we want to be able to provide Gen AI experience on edge device. That's what we want to do. And a lot of works that we see in this conference, we try to address different angles of this challenge, right?
37:12As well as expos and demos and talks that we will cover later. So one thing that has been quite popular, one approach that's been quite popular in recent years for fine-tuning different large models, especially large language models, has been LoRa. lower-rank adaptation. So what you do there is that instead of just fine-tuning a neural network end-to-end, which is quite big and probably not doable for many, many people. Yeah, expensive and all that. So what you do, you freeze the model, and you would say, okay, let me pick a particular weight matrix. Let me build this mirror weight, which is decomposition which is decomposition to multiplication of two matrices A and B.
38:11And there are low rank matrices. So there's a low rank bottleneck here. So this significantly reduces the number of parameters. So when you have these low rank matrices A and B and decompose them, then you only fine tune those A and B matrices for fine tuning your model. So in that sense, now you're able to fine tune big models and at the end of the, once the training is done, you have the possibility of fusing those learned models with the actual weight. So because the input is going to be the sum of these two operations, so you just can fuse them into a single weight and use them. Or you can keep a library of those and plug and unplug depending on the use case.
38:56So there are different challenges that we have to address here. And the first thing you mentioned, this hallowed net paper that is about this on-device personalization addresses on-device LoRa. So what if we want to be able to do that on-device, right? And that paper basically tries to address this problem by just, I just sketched a rough idea. The rough idea is that, let's say you have a kind of a unit type architecture and you You just take the bunch of these bottlenecks and throw them away. That's what you do. So you just hollow them. You mask them out. And then try to use the fine-tuning using the rest of the network.
39:43And this is roughly the idea, using LoRa on that. And, of course, there are a lot of subtleties to make sure the back and forth is working. And paper actually goes into quite detail. I think last year around NeurIPS, I spoke with Fatih, one of your colleagues, and we spent quite a bit of time talking about UNET and some of the things that you're doing with that. It sounds like this is an evolution of that interest in UNET and now applying LoRa to UNET. Absolutely. Is that the idea? Absolutely. Absolutely. Absolutely. So Fatih actually is going to have our expo talk that I will talk about, about degenerative AI on edge.
40:23So this is very much related and I think this will be covered there as well. So that's one of the LoRa papers we have. The other LoRa paper we had is SHIRA, which is Sparse High-Rank Adapters. And this paper tries to address the problem of fused versus non-fused adapters. So what you can do, you can try to, when you have an adapter, you put it on the network and you fuse it with the weights. you do your inference and then when you have a new adapter you have to unfuse them right and then put another one and then fuse them again so that's basically the the procedure and this comes what you're describing is operationally at kind of inference at scale like assuming um i guess i'm thinking about it in the context of like on device versus cloud if you're on device you probably have one adapter that you're using and you're not doing this, but if you're doing inference across a bunch of users, you're swapping out these adapters?
41:31Is that what you're describing? I think let's focus on just the on-device use case, right? Let's say, suppose you have a base LLM on-device, and you have different adapters for different applications, right? Got it. And that can happen, right? So you have fine-tuned for different use cases, but instead of changing LLM every time, you just change the adapter. That makes sense. I was thinking of the personalization use case of LoRa. Yeah, the personalization. I think the on-device, the hollow net paper is on-device. It's a personalization angle. So you basically want to personalize on-device. That is the personalization.
42:10But for the Shira work, the idea is that let's assume that we have done that. We have already trained our LoRa. Now, the problem is, you know, the latency price we pay that every time you have to unfuse and fuse back because you need to sum up things and do this. So, and if you avoid this, you just use a single adapter, of course, you lose expressivity because, you know, you don't get the gain and flexibility that you get with multiple adapters. So what they try to do, they say, okay, instead of just using these low-rank adapters, let's try to use mask-based, sparsity-based adapters. Basically, you look at your network, you freeze still the whole network, except like a few weight parameters.
42:58Sparsely select those weight parameters and only update those. Now the situation changes because fusing operation becomes really easy, right? Because it's just a matter of replacing a few weights. So this paper basically tries to look at these sparse high-rank adapters because that's in contrast to low-rank. So that's another challenge. The other paper we have is FURA, which is like a Fourier low-rank adaptation. And that's because if you look at all these text-to-image generators, they tend to have typical lower fine tuning they tend to have not much diversity in terms of image they generate on device so that kind of becomes one of the challenges and the way we address this in this paper we go to Fourier domain so we try to do this lower rank adaptation in Fourier domain get a Fourier transform do lower rank adaptation go back but also we will do an input-dependent rank selection.
44:04So the bottleneck is going to have dependency on the input. And with this adaptivity, we managed to address some of these challenges. So that's another work which focuses on the text-to-image generation. And yeah, it's kind of a tree-laura paper we have on this. And so I think another area that you are doing or showing some work in as in related to robotics? Yeah, so we basically are providing a benchmark called Clever Skills, which is about compositionality at the end of the day. The idea is that suppose that you have a model that has learned some basic operations, moving things out, moving from left to right, lifting things and all those things.
44:56What if we define the task that is composed of those? How good are models performing those compositional tasks? So the clever skills paper, the dataset and benchmark, actually provides such compositional dataset, which has multimodality, and it's basically about compositional reasoning in robotics. And I think it's a very, very, very nice benchmark for researchers to evaluate the compositionality angle. and yeah, it's kind of a combination of various modalities in the dataset. Another connected dataset that we also release here is kind of a life fitness coaching dataset, we call it. And this is more about an asynchronous interaction between your AI agent and in humans.
45:57So instead of just the multi-turn conversations that you just say something and then wait and then get an answer and so on and so forth, you try to make it more proactive, more asynchronous. And in the fitness use cases, this is a particular case. And the live fitness is actually providing such benchmarks and data sets. Roland is behind that. I think you had Roland on your podcast, so he's behind that. as well as clever skills. So it's an interesting data set as well and would be connected again to this type of embodied AI type work, I would say. You mentioned that you're involved in the expo as well.
46:38What are you talking about at the expo? Yeah, so just, you know, there are like, let me just say before that, that there are like a bunch of other very interesting papers on mixture of experts and stuff. But, you know, I encourage folks to look at that. I've, you know, I'm communicating. You're tempting me now to dig into the MOE paper and some of these others, but there's way too many to cover all of them. Absolutely. I agree with you. You have to be selective. But, yeah. So, we have actually an expo talk at Noreps focused on generative AI on the edge, bringing generative AI to the edge device, basically.
47:18Which samples a lot of things that we discussed right now. One is about self-speculative decoding, which is an efficient way of doing speculative decoding on device. We cover those things, but 3D generation models, how we can put them on device, what are the challenges, how we address them. But multimodal languages, multimodal models, sorry, how we enable them, like text vision type models, making them more efficient, as well as all sort of content generation, visual content generations there. And, of course, LoRa as well. So there are a lot of LoRa-related content as well in the talk. So, yeah, Fatih is going to go over a lot of things that came out of Qualcomm Research and was showcased in various summits and demos.
48:13So it's going to be a quite deep dive onto those things, accompanied with a bunch of demos that we will have. So actually, I want to highlight, it's going to be only five demos. So I hope we can cover that here. I think some of the things I want to highlight is, let's start with Mixture of Expert this time. And on the Mixture of Expert side, so Mixture of Expert is a kind of a promising approach for having more efficient models. So instead of just using the old model, You use subnetworks, which is more efficient, plus an expert selection. Now, the challenge here from device perspective is that loading and unloading different experts, it's going to add latency, right?
49:01So you want to be able to cache effectively some of these experts on device based on the tasks that you have to be able to address those challenges. So one of the demo we have is about this MOE caching, basically, and we're going to showcase that. I think it's an interesting demo, and we're going to cover that. We have some video generation-related demos. We will show the first video editing diffusion model on device. So that's an interesting thing. And accompanied by, related to topic of simulators we discussed, topic of basically data augmentation, data generation for driving scenes. So basically putting objects in the scene, replacing objects, but keeping the realistic aspect of the scene.
49:58So it's just not an arbitrary placement, but we want to have some consistency. So that's why underlying that is like a 3D generation and like a 3D Gaussian splatting based rendering that provides this multi-view consistency, if you want, of the model. So we're going to have demos of those type of data augmentations as well as on-device on-device LoRa was one thing that we will cover. And one last thing that I think also would be interesting to highlight is a workshop that we have on Qualcomm AI Hub. I think that would be an interesting thing for various developers. how they can pick models from Qualcomm AI Hub, onboard it on Qualcomm devices, basically, and all the challenges that come with it.
50:51So Siddica was actually on your podcast a while back. She's behind that. So we're going to talk about all these things in the workshop. So I think that will be an exciting also workshop for developers. That should be a great one. I love the idea of being able to develop these models or take these models that you're developing and deploy them out to an array of devices of different types and do testing and not have to have all those devices locally. So it'll be interesting for folks to catch up on that. Yeah, absolutely. Absolutely. I agree. Awesome. Awesome. Awesome. You mentioned Fatih's talk at Expo.
51:41I'll include a link in the show notes to that previous interview that I mentioned. I think it was CVPR and not NeuroPS, but in either case, I will link that because we go into, as I mentioned, the UNet stuff. You mentioned selective decoding. We spent quite a bit of time on that and some other topics that are coming around this time as well. And I'll also be sure to include all of the links to the various papers, the ones we discussed, as well as the other 10 or so. What's next, do you think, when we're back here next year talking about NeurIPS 2025? What are we going to be talking about? So one thing that is, there are a lot of interesting questions right now and people discuss about that.
52:33The question of scale have we hit the scale of data? Yeah, exactly. Are we running out of data? Shall we think of something else? Can we gain more and solve some of the problems like reasoning problems like scaling up? Or like there are other ways like a hybrid design, some sort of neurosymbolic type approach. And the whole agentic AI is another thing that excites me quite a lot. Like what is, you know, how we can build a system and what are the components of that? And what is even a framework within which we can think about these problems? I think these are like very, very exciting topics for the next years.
53:20These problems are open and, yeah, excite me quite a lot. Yeah, it's almost surprising that we're only now talking about agents in this conversation. So many of the conversations I'm having nowadays are talking about agents. Yeah, maybe I should promise a dedicated agent to talk in the future.
53:42Awesome. Well, Arash, it was great to catch up with you. And I hope you have a wonderful NeurIPS experience. And looking forward to next time. Yeah, looking forward as well. Thank you, Sam. Thanks, Arash.
From the publisher
Today, we're joined by Arash Behboodi, director of engineering at Qualcomm AI Research to discuss the papers and workshops Qualcomm will be presenting at this year’s NeurIPS conference. We dig into the challenges and opportunities presented by differentiable simulation in wireless systems, the sciences, and beyond. We also explore recent work that ties conformal prediction to information theory, yielding a novel approach to incorporating uncertainty quantification directly into machine learning models. Finally, we review several papers enabling the efficient use of LoRA (Low-Rank Adaptation) on mobile devices (Hollowed Net, ShiRA, FouRA). Arash also previews the demos Qualcomm will be hosting at NeurIPS, including new video editing diffusion and 3D content generation models running on-device, Qualcomm's AI Hub, and more!
The complete show notes for this episode can be found at https://twimlai.com/go/711.




