In short
Podcast Summary: Gen AI at the Edge: Qualcomm AI Research at CVPR 2024 with Fatih Porikli - #688
Episode Overview In this episode of The TWIML AI Podcast, host Sam Charrington interviews Fatih Porikli, Senior Director of Technology at Qualcomm AI Research. The discussion revolves around Qualcomm's cutting-edge research showcased at the CVPR 2024 conference, touching on various generative AI and traditional computer vision topics. Key themes include the efficiency of training and inference for mobile and edge deployments, and innovative applications of generative AI technologies.
Key Topics Discussed
Qualcomm's Research Focus
- Generative AI Evolution: Qualcomm is at the forefront of advancements in generative AI, focusing on integrating various modalities including speech, language, and visual data.
- Applications in Mobile and Autonomous Systems: Research extends to applications in mobile devices, autonomous driving, virtual/augmented reality (XR), and robotics.
Highlights from Research Papers Fatih Porikli discusses several accepted papers from the CVPR conference, summarizing their contributions and implications:
- Clockwork Units
- Focuses on improving text-to-image diffusion models using a UNET architecture.
- Introduces approximating networks to reduce computational load by over 30% without sacrificing image quality.
- Look, Remember, and Reason
- Develops a multimodal model that enhances object-specific information retrieval in videos using stochastic probing during training.
- Aims to improve the model's ability to learn about object movement and interactions within video frames.
- Edge Relight 360
- A model for generating high dynamic range (HDR) 360-degree images for realistic video portrait relighting.
- Users can create dynamic environments around 3D models based on text descriptions.
- What to Say and When to Say It
- Introduces a new approach for video language models that provides actionable feedback, simulating a personal trainer's role.
- Presents a benchmark dataset, FitCoach, to facilitate research in interactive fitness coaching.
- Math Search
- Establishes a benchmark for visual reasoning over mathematical plots, addressing the challenges of training models to interpret graph-related queries.
- Speculative Decoding for Multimodal Models
- Proposes a method to improve inference speeds in multimodal language models by utilizing a lightweight draft model alongside a larger model.
Optical Flow Research
- Focus on enhancing optical flow algorithms, including:
- Optical Flow Data Augmentation: Using intermediate frame generation to improve training datasets.
- Self-Cleaning Inversion: A new method for refining optical flow estimates while minimizing drift and errors.
Low Latency Neural Stereo Streaming
- Introduces a novel approach for compressing stereo streams simultaneously, achieving lower latency and better efficiency in video streaming.
Demos and Workshops
- Demos at CVPR: Demonstrations include real-time generative models for mobile devices and applications in autonomous driving.
- Workshops:
- Efficient Large Vision Models: Addressing the need for efficient deep learning models for edge devices.
- Omnidirectional Computer Vision: Exploring applications of omnidirectional cameras in various sectors.
Conclusion Fatih Porikli expresses excitement about the progress in generative AI and its applications in enhancing user experiences in mobile and edge contexts. The episode showcases Qualcomm's commitment to pushing the boundaries of AI research, emphasizing practical implementations and efficient technologies.
Additional Resources
- Full episode and show notes can be found at [TWIML AI Podcast Episode #688](https://twimlai.com/go/688).
---
This summary encapsulates the key discussions and findings from the episode, highlighting Qualcomm's innovative contributions to AI and machine learning in the context of generative models and computer vision.
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00Let's say I have a very respectable LLM and it is taking also the ordinary regular feedback would be ah, the user here has successfully completed a squat. But the version that we are presenting, it says, oh, smooth on the way down. You know, it's literally useful information to me.
0:36All right, everyone, welcome to another episode of the Twin Wall AI podcast. I'm your host, Sam Charrington, and today I'm joined by Fatih Parikhli. Fatih is Senior Director at Qualcomm AI Research. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Fatih, it's great to have you back on the podcast. It is becoming a bit of a tradition for us to have you back on to discuss Qualcomm's research at the CVPR conference. The vision team there has been super busy. I don't think we've had quite this many papers to review for any of our conference episodes with your team.
1:14I think there's a total of 16 of which you're an author on three of the main conference papers and another four of the workshop papers. Welcome back and congrats on all the accepted papers. Thank you so much, Sam. It's always a pleasure to talk to you at this podcast. You know, I have been looking forward to that. We are doing a lot of great, interesting things at Coalcom and I would love to, you know, talk about some of them with you. It's a pleasure to be here again. Thanks for having me. you know. Let's maybe start by having you zoom out and talk a little bit about what were some of the big ideas that your team has been working on since last year's CVPR.
1:56It was a very, very busy year. We are witnessing another evolution in research. Very disruptive changes are happening thanks to Gen8VI and Qualcomm is at the front runner of that research big wave and we have been focusing on Genitive AI, we did great work in taking such Genitive AI solutions, LLMs, LVMs, stable diffusion image generators, video questionnaires, and everything. And in addition to those, we are extending to other modalities from speech to language to image to any other things that you can imagine. This is for mobile applications, but also for autonomous driving, for XR, I mean, augmented reality and virtual reality and iot you know robotics anything you can imagine you know kind of we are working on critical applications as well for these gen AI solutions in addition to what we do in computer vision and perception and everything yeah and so i think we'll dig into some of the papers that your team has worked on and we'll maybe organize those in the the way you just suggested.
3:12We'll talk through some of the Gen AI related papers first, and then we'll dig into some of the papers that are more traditional computer vision oriented around perception and motion and other things. First paper that I'd like you to talk a little bit about is the Clockwork Units paper. Can you tell us a little bit about that paper and what it's focused on? Absolutely. As you said, these diffusion models, in this case, it is a text-to-image-generator diffusion model, and in the heart of it is engine, something called as UNET. and UNET because its shape, it is opening and closing. There are higher resolution features and in the middle, it progressively makes the feature size smaller and then it expands.
4:00So this is an autoencoder architecture. We have an encoder and then bottleneck and decoder. So this is UNET. It is computationally expensive because most of the computations we do in text image generation, we repeat this UNET many, many times. maybe 20, maybe 40, maybe more. So this is where most of the computation goes down. Anything that we can accelerate this part will be useful. So we look into UNIT and we realize that not all operations have the same impact. I mean, operations in this UNIT architecture, middle layer, or kind of like the initial layer or closing layer, equally important to the output image quality.
4:39the changes in the higher resolution layers, even small perturbations there will kind of change the output image. But in the middle layers, if we change them small, you know, it doesn't change the output image quality. However, it impacts the composition of the image, overall composition, where the objects are, where the things are. So based on this observation, what we did same is we replace the middle layers in this unit with an approximating network, which is taking the previous features and the current features coming from these larger special layers. And then without running the middle layer, we do very simple operations and we approximate in a way, you know, kind of in this architecture, which provides maybe more than 30 % savings in terms of flux, you know, computations.
5:39It significantly accelerates the overall computation load of the unit. Is it counterintuitive that the lower resolution feature maps at the center of the unit would be more resistant to perturbations? It would seem like a perturbation there would have an outsized impact because there's less information that you're perturbing, if that makes sense. It has an impact, but it has an impact on that. It's a compositional impact. It is not going to change the refined textures. So what we do, we initially run entire units. So we let kind of this diffusion to learn where to put the things based on the given text prompt.
6:26And then once it decided, okay, there will be a face here, there will be a tree here, a flower here, let's say, that object type semantics are not going to change significantly, even if we perturb the middle layers. However, the bigger layers in terms of spatial, of course, these small and big in terms of the spatial resolution of the feature size, feature maps, kind of if we change them, it is going, you will, it will be visible because fine texture is going to change. But if we change the middle layers, object will be maybe kind of slightly moved around or kind of become some different object.
7:03But composition-wise, it's going to be similar. So the idea is that for these middle layers, once the network starts to converge on a composition, you can do a lossy approximation of prior iterations as opposed to doing the full calculation as you roll out the network. You are correct. Yes. Of course, keep in mind that this generative model, it takes only a kind of user prompt, text prompt, and it starts with a random seed point. So you actually don't know what is going to come out of it, right? And we take advantage of this fact. Yes, we changed the composition slightly, slightly also, because as I said, at the initial iteration of the unit, we run the full unit.
7:54Then, you know, we borrow what we computed before. So kind of we take advantage of that fact, you know, it is slightly different image coming out. It's not pixelized accurate if you run the full network. But regardless of the quality, then we look at the FID scores, clip scores. It is not any less than the, you know, kind of the full run of the unit. So maybe in some cases, it's even more preferable, even though we say cut 30 percent, more than 30 percent of the computations. Got it. And you mentioned FID scores and CLIP scores as the valuation metric. Can you elaborate on those and what they measure?
8:32FID score measures the diversity of the generated images. It is a distribution-wise metric. It looks at, okay, what comes out of such generative models? How different they look like? are they generating the same type of images or different type of images? When I type a white bird, it's going to just generate the same white bird in same pose or different type of white birds, different type of pose flying or other things. So that is the FID score. This is an inception score, freshet inception score. The other one is CLIP. CLIP is evaluating how the generated image confirms with the given text prompt, because this score, this is learned, this network, there's a CLIP, VIT, and text encoder, image encoder and text encoder.
9:30they are learned to map into the same space for image and the generated image and the text prompts to be consistent. It minimizes the distance in that mapped joint space. So when we look at these scores, what it means is that the distance FID is smaller, which means you generate very diverse results, which is what we want to do, and also clip scores, they are higher because there's better agreement between this text input and the genetic image. For this paper, you released code as well? Yeah, absolutely. We have now a new initiative at Qualcomm AI Hub, which kind of we take all our models, internal models.
10:22And we also provide some of them to our community, research community and everyone. So you can access Clockwork and many other, you know, kind of models. So the next one we wanted to talk about is look, remember and reason, grounded reasoning in videos and language models. Yeah, we are going over our genetic papers. This one is quite interesting because it is doing something very interesting. These networks, for instance, we have, let's say, a video LLM. It takes video input as an input to the algorithm. Then there's a language model, and we ask a question about, you know, what happened in this video.
11:12And these models are good in terms of capturing the global information, But in terms of the object-specific information, for instance, where was the object in the scene and what was the object type or object identity and how object mode, like this detection, recognition and tracking information is not very well captured by such models. But we have this very interesting stochastic probing approach. What we do during the training, we randomly ask questions about, for instance, about this information that I just mentioned. What was the type of the object? Where was the object? How that object is moved?
11:58So this is, we do such questions. We incorporate such questions in the training of this video. So LLM, so it's a multimodal model. It's a language model that takes the video as an input. And in test time, you know, kind of, we do not need to explicitly probe such information. So then this inference, when we are asking a question, the response would be more efficient test, you know, in terms of the number of tokens, you know, inference would be short also. So for instance, I can give you some examples of what kind of questions we are asking. First, one is blicket experiment. Blicket meaning, okay, I have a couple of objects and one is the indicator.
12:48Think that there is one of them kind of creates a sound and the other one doesn't. And for instance, human beings, we are very good in identifying if I show object A, if it generates a sound and object B doesn't sound, we understand object A is the one that Gen A sounds. But it is not always that easy. I can show both of the, a picture of both of the objects A and B, and we hear a sound, then maybe one of the objects hidden behind the other one, and you still kind of hear sound. And in this case, which object, and then you remove one of the objects, like it could be A or B, we can reason about that.
13:28but it is indirect and it's difficult to directly learn such a thing. Humans, we can do very well, but these algorithms, they need to be trained. So this is one of the experiments or another question, for instance, which object would the person put next? Which object was thrown by the person? We are asking about, okay, what is the object type? is whether it's like a cup or apple or a ball, you know, kind of. So those type of questions we are trying to answer with this model. And so how do these stochastic probes help the model learn about the relationships of these objects and the actions? So, such low-level spatial and temporal details are available in each input frame and in the video.
14:28But during the training, if you are not asking these questions, if you are just taking an input frame from a video and encoding it through, let's say, CLIP or Dynovita or any kind of these encoders and providing tokenized version to LLM, they do not seem to learn such low-level information, spatial temporal details. So what we do, we actually ask questions during training, okay, what is the type of the object on this patch, on this image region. These type of questions are asked to algorithm for algorithm to remember, reason, learn, you know, kind of, I mean, look, remember, reason is the type of name of the paper.
15:15Exactly. That's what it is trying to learn during the training. So in test time, you do not need to explicitly, you know, have an object recognition or object tracking module, you know, kind of, yeah. It sounds a little bit like this tweak to the training process is kind of giving the model additional information about what to attend to during the video, what to pay attention to. Absolutely. It learns what are the important things about what are the objects, what type of objects, what to attend, how they moved, you know, also across the video. It is not just their position and their type, but how they moved in the video.
15:59Yeah, exactly. This is what this network learns. Of course, this network also has, you know, kind of mechanism to incorporate such information through its low level encodings. Yeah. And when you say you're asking questions during training, what does that mean? Meaning you're augmenting your training dataset with a specifically formatted question to get at this low-level information? You are right. Because such models, these visual question-asping models that uses LLMs and some visual encoders, There are these instructions, question and answer pairs, in addition to given input image or video in our case.
16:50Yeah, this is through the training. Keep in mind that algorithm has a temporal attention mechanism across multiple frames and also, you know, self-attention, special attention for each frame. So it kind of can learn that information. It has capacity, but if you are not asking the right questions, it doesn't focus on that information. That's what we are enabling such models to learn through augmenting the way that we are. In a way, you can think that we already trained the network and then we are fine tuning the network. It sounds a little bit like a form of curriculum learning in that you're trying to, or maybe multitask also, like you're trying to teach it to pay attention to these low-level features that are important while you're also teaching it this bigger objective.
17:44Maybe less so curriculum, but maybe multitask. You have a very interesting point. I'm very happy that you mentioned, Sam. In both curriculum learning and multitask learning, if you want to do supervised version, you need explicitly labeled data, right? And also, you like that data to be extracted, available to reason. For instance, I know already where the object is, so there's a detector, and I know how it moves, so there's a tracker, and also I know what kind of object, so there's a classifier. However, we do not want to make it complicated like that. We want distinct information, the network itself, to naturally ground it with that type of information.
18:29So we do not have any explicit detector, tracker, or classifier, unlike multitask learning or curriculum learning. So this is through the capacity of the network, and by asking the right questions in training time, And we allow it to pay attention to such details when it is learning its attention mechanism, like the spatial and temporal attention layers that I mentioned. So the next paper we wanted to cover is called Edge Relight 360. And the focus is text condition 360 degree HDR image generation for real-time on-device video portrait relighting. That's a long, long title. So this is a generated model which generates high dynamic range image maps.
19:17And then we use those maps. We will use those maps in this case, in this paper, as the environment lighting. So since it's 360, we can rotate it around, let's say, target object, which is a face, let's say. and we know the 3D structure of that object because we can generate the approximate 3D from a single image. So given, let's say, a single image or continuous video, we know the 3D structure of that object, maybe a talking face, and then the environment, its background is automatically generated. It's not like static in my case, and also not real, but maybe it's realistic looking, or you can create an environment which is like space.
20:03you know you are on the moon and then when we rotate the environment this 360 image generated image lighting is gonna change because maybe there are lights like sun if it is in the you know kind of moon or kind of in a real indoor environment maybe windows and you know lighting sources so when you you can now you can rotate it it's a 360 image and then the lighting is gonna to change on the person. This is a very interesting, very natural looking cool effect. You know? Yeah. Of course, the challenge is how are you going to generate these 360 images? Because the fusion models, they do not generate 360 images.
20:42They only generate like a kind of regular two-dimensional images. Yeah. The model is generating the entire 360 degree image. Is that, and is that from a textual description? Input is exactly just the text prompt. Actually, we have a demo about this one also. And any general AI model, we use speech interface. I mean, Qualcomm has wonderful speech encoders and everything. And they are very fast and efficient. It is like GPT-4. So through the speech, you describe what you like to have around you, this 360 image. and the challenge is that when we generate this 2D image, the left-hand side and right-hand side is not going to, when we wrap it around, it's not going to, you know, stitch seamlessly, stitch together because it doesn't know that, you know, kind of these two sides should, then I wrap it around, you know, this is a 360 image, wrap it around, they are representing the same kind of, let's say, part of the scene.
21:55So in training, we explicitly imposed it. So we can take any baseline diffusion model from this stable diffusion or SDXL, anything you can imagine, and then we change the data set and we change how we impose this consistency on the left and right kind of parts of the generated images, which allows us, of course, we are using HDRI. high dynamic range image data set, specific data set as well. So we can then have a model that can generate 360 images. So that's what we did. And this runs very efficiently on the phone, you know, on mobile devices. This one sounds like one that you need to see. There's is there a video demo available?
22:47There will be a demo actually on device. You know, you can interact with the demo. please stop at our booth at CVPR. We have many, many demos, actually. Maybe later we can talk about this. We'll link to the papers and available demos as well from the show notes page. The next one in this batch of generative AI papers is What to Say and When to Say It, a video language model and benchmark for situated interactions. Another long title as well. And it's a very interesting title also. And also, let me give you an example of what it does and what other solutions cannot do and how these people contribute.
23:31For instance, let's say I have a very respectable LLM and it is taking also video as an input. And we ask this question, you know, maybe the initial prompt is user should be doing some, let's say squats. And then also another prompt, instruction, provide feedback to the user. We want this model, look at myself and I'm doing some squats and provide feedback to me. The ordinary, regular feedback would be, ah, the user here has successfully completed a squat. You know, this is okay. It is not incorrect. It's correct, but it's not useful. You know, I know I did it. But the version that we are presenting in this paper, what it does, for instance, I'm doing like this squash and it says, oh, smooth on the way down.
24:29You know, it's literally useful information to me because I'm doing something maybe incorrect and it is trying to correct me. Or when I'm doing something right, it motivates me. It says literally good job. You know, excellent. And also, it's like natural, not kind of when you go to gym, you have these people helping you. Like it is personal trainers. It's not like a robot. I'm guessing based on that use case that this grows out of the work of Roland Memesevich and some of the others on his team. They were working on a fitness assistant at Qualcomm and prior to their acquisition. They did an excellent job before.
25:11They had this amazing dataset and they continue, you know, extending, leveraging the dataset for this type of genetic AI purposes. By the way, this paper is also talking about introducing a very large scale dataset, benchmark dataset, which we call as FitCoach. There are more than 470 hours of videos and these videos are organized into five second chunks. And we also provide 1.7, more than 1.7 million question answer peers and also a related model. I mean, we also provide the model, you know, kind of so people can use it like coach model, you know, kind of, and we provide kind of these initial results, you know, numbers, metrics, and everything.
26:06So with this paper, I think Qualcomm is contributing significantly to accelerate research in this area, this visual, video question answering, video models. Talk about what it is in the approach that is generalizable beyond this specific use case. The idea in the paper, is it about presenting this specific use case or is it a technique that applies to other use cases as well? Of course, we have this data set, as you mentioned. That's why we are focusing on this, you know, fitness. This data set is about fitness. However, the model and the lessons learned from this dataset is not specific to just fitness.
26:52It is any type of interactivity, action, interaction that a single person or even multiple people would do. The model also we presented, talked about in this paper, it uses 3D CNN. It is a general purpose, you know, encoder to take the video and, you know, generate useful information, which would then go into this language model backbone. And it is also a general purpose. So kind of, I do not think it is limited to this application only. You can imagine any other things as well, any other activity. Is there a distinction worth pointing out here between fine-tuning and instruction tuning? Meaning, when I think about what you do with the use case, maybe getting a model to think about the motions and analyze them more critically, I might think of as a fine-tuning kind of task.
27:59but being more kind of coachy in the way that it speaks, I might think of it as more of like an instruction tuning kind of. Does that intuition map to your experience? And what specifically of those techniques did you apply here? It is related, very related. Of course, it is not only to improve when we are asking a question to this model, text question or, you know, coming from the speech interface that it would give a response, but it is also learning to provide feedback by itself, just, and that feedback comes from the visual data. So I'm doing things in front of this model and then that model observes me, you know, then provides as a tutor, as a trainer.
28:53There's no specific text prompt that is initiating the model to provide the feedback. Exactly, exactly. Of course, kind of when these models should kick in, you know, provide such encouragement or corrections, you know, kind of is all learned through these high level questions and, you know, fine grade questions in the data set that we provide. So the data set might have segments, it will have some coaching instructions and the model learns to associate that with what's happening in the segment. And then when it's watching a longer video, you know, it might see a pattern that kind of prompts it to coach in a particular way.
29:36Right. And let me give some examples from this data set. What kind of questions they are? For instance, one question could be, is the user going as fast as possible? And the answer is, yes, the user is going as fast as possible. Another one, does the user have good form? The answer, no, because, I mean, not because there's associated video chunk, right? The user doesn't have a good form. And then there are feedbacks also. So these are answers. There's feedback also in this data set, for instance, great job, but watch for your form or deep squats and two punches, get it and go deeper in the squat.
30:15I mean, this type of things, you know, kind of if you can speed up, that would be great. You know, it is not only question answer, but also feedback. Is there anything that was required to do to map from question answer oriented data set to it sounds like the use case scenario does not involve questions? at all? Did I hear you correct in that? What I mean is that when this model is in use on a device, it is offering the feedback unprompted by the subjects. So it's looking at the videos and when it sees patterns that it remembers from its training data set or similar, that kind of prompts it to initiate this feedback?
30:56Yeah, because it contains these annotations for the feedback in the dataset. That is not, for instance, something available outside this dataset. This is one of the contributions of this dataset. There are feedbacks in addition to answers, questions, instructions, and answers is the ordinary way of training such models. But now we have also the feedback. Recently, we heard about this Microsoft recall feature that the idea is to like capture your screenshots from your computer and use that so that you can ask questions of what you were doing and do retrieval types of tasks. One can imagine in conjunction with this type of a model that is, you know, trained on, you know, video or set of screenshots at some point, identifying what you're trying to do and then offering you tips for how to do it more from a productivity perspective than an exercise perspective.
31:56There's a bit of an analogy in a productivity-oriented model? Of course, it is monitoring you, I think. I didn't use recall feature yet. I need to test that. It's something very interesting. But in general, such models, they observe you through these model tests, let's say skin capture images. And if there is something undesired, maybe there's a problem or you mistakenly change deleted some important information or you want to just remember what you did, you know, kind of, there is some feedback from the user. So that information creates this training data. So here are the, you know, what happened and whether the person operation, whoever using this system is happy or not, you know, kind of, it is naturally a part of this, you know, training information.
Read the full transcript
32:56It can be constantly learning or it could be distributedly learning across multiple people. But I think this is a very smart way of getting additional information to further fine-tune, tailor, customize such models, personalize such models. The next paper is called Math Search, a benchmark for multi-hop step-by-step visual reasoning over plots. What's the challenge there? Models, multimodal models, like the previous one was a video multimodal model, and this one is, let's say, image multimodal model. There's an image as an input, and we have this language model, and we are trying to do things around math now.
33:45So these models, they use, they take these input images and there is something called this image encoder, tokenizer, which could be a clip, VIT, or kind of a dyno-way to any type of image, respectable image encoder network, which takes this input image and generates tokens. So it is many of those solutions, they need to scale down the image and they are looking at, you know, patchwise information, texture information in that type of input. For instance, I have a picture of a plot, a plot, literally, maybe it's coming from MATLAB, you know, or some kind of Excel graph, you know. there isn't much texture and there are maybe small but important details like there is a curve, it is crossing some axis, you know, something like that.
34:47And the questions that we want to answer, like how many, what are the, how many zeros does a function t have? Like where does it crosses the, you know, kind of, let's say in this case, vertical axis, or list all the discontinuities of this function. So it has to understand this continuity and everything. Such information is difficult to learn. So this is a dataset paper. It introduced this benchmark dataset. In the benchmark, We provide more than 200 ,000 example plots, and we have 2 ,000 also questions, more than 2 ,000 questions in the test validation part. So we have single function plots. We have multifunction plots.
35:40We have plus within plots. You know, kind of, we have multiple plots. So this is a very challenging data set created, put together to allow research better solutions, better encoders, better fine-grained visual information learning for such multimodal models. And I got the impression with this one that the sole purpose wasn't necessarily to create models that are better at math or interpreting charts, but also that asking the model to focus on some of the details in these math pictures might help models that aren't specifically, you know, purpose to do math better. attend to details in other types of images.
36:39Because of the current image encoders, visual data encoders, they focus on the global information. And with this, we are motivating, giving momentum to further research that will leverage on the local information. And was the evaluation and the development, What models did it use? We have results in the paper on this data set, like how human would perform, how, for instance, there is this lava, which is a model which can take image and text and then do reasoning about that. There is QN, another multimodal model, LMM, large multimodal, some people call it like this. There is also chart lama. There is GPT-4V, the visual version.
37:34So we provide some baseline results for this, including human performance. We show that there is a gap between human performance reasoning on the mathematical math plots with the gap between human and the existing genetic AI models. and the gap is actually quite large, you know, quite big. On average, let's say human does 70%, human accuracy, let's say around more than 70%. But these models, even the best one, which is GPT-4V, is around 25-ish, you know. It shows that there is significant kind of progress to be made, promise, you know, we are motivating people to look into this gap and provide, think about better solutions.
38:27The next paper on speculative decoding for multimodal language models. What's this paper trying to do? There's something called speculative decoding for language models, LLMs. What it does, there's this, let's say, big model, or, you know, kind of, and there's this drift model, tiny model, small model. So big model, we need to go to parameters from memory and to one iteration, then we do another iteration. So it's a big model, it takes time. And the token rate is something, how many outputs, let's say, words, you can think that tokens are similar to words, how many words it can generate per second is an important metric.
39:15We want it to be high, generate many, many results very quickly, not slowly. But there is associated cost of compute with the big model. So what speculative decoding does, it uses a tiny model approximating the performance of the big model. It computes multiple tokens, let's say three, five tokens at a time, then we run the batch big model just once as a batch and we look at the results and if If results are same, then we kind of accept it. We don't need to step by step run the large model. Large model only runs on the batch, which allows us to accelerate significantly. If we look at the results coming from the draft and the kind of batch results from the big model, then we can very easily skip, let's say, two tokens, four tokens.
40:17if you are using three, five, you know, it is a significant speedup. No one so far tested whether, okay, such speculative decoding idea would extend to multimodal. In this case, we have lava. Lava has, we just mentioned, there is like this LLM part, which could be, let's say, Lama 27B or Vikuna, any model, any large language model. And there is this image, and both of them goes into the model. image goes through its own token, so you can ask a question and then it will answer. So the test goal, can we make it faster by adapting speculative decoding like ideas here? And that's what we investigated.
41:01We use a tiny model of 150 million parameters and we observe that we can significantly accelerate up to 2.3 times, you know, using the speculative decoding for a multi-model, not only language model, but multi-model model as well. This graph model could be any model, you know, kind of. It's not this kind of observation. we show that it goes beyond a specific model you know big models draft model it is kind of a useful information it shows that spd speculative coding very easily extends to you know multi-model and just to understand is the is the draft model is it part of a larger model for inference?
41:56Is it like a subsystem or is it a replacement model that's been, you know, like it sounds like distillation in a sense? Draft model is learned to do a good job in approximating, estimating, or predicting, you know, the output of the big model. So we run draft model and we run also the big model. both of them coexist in the inference time. I mean, it's not like we only run draft model and then go with its own decision because draft models sometimes make mistakes. It is, we intend, we want it to approximate the big model, but it is not the big model. It's going to make a mistake. That's why I keep saying on average, sometimes it fails.
42:47So if it fails, then you go around the big model at every, let's say, three, five tokens, you don't accept wherever it is not aligned with the big models, batch processing results. We go back to that earlier token and run the big model. So, of course, if the draft model makes a lot of mistakes, running the draft model and then running big model, like every token is more computationally intensive. It takes more computation. But on average, if draft model does good job, you know, it can allow skipping three token, five token, more tokens, you know. So, yeah, we need to run both of them. In the inference time, we have both draft model and the big model.
43:33Is the draft model used in inference in training or just for like plain inference? Meaning if you're just doing inference, how are you evaluating the output of the draft model to determine whether you need to invoke the real model? So it generates tokens, draft model. We run it iteratively, let's say, five times. We generate five tokens autoregressively, not as a batch. So we have this token, another token, another token, and sequentially. Yeah. And then we have the results. One time runoff to big model comes as a batch. We generate multiple tokens. We do not do it autoregressively. So there's a danger, of course, you know, if they are aligned, if they are same, then looks like we are good.
44:29But if they are not the same, then we need to go to the big model, then auto-aggressively, sequentially, start from that false generation token, update it, and then generate the rest. Because it took a while for me to understand and appreciate how speculative decoding is working. I was thinking, yeah, you do more computation though, right? kind of then what is important in some cases we can run this drift model and the large model in parallel on different platforms that one could be running on nsp or mpu you know kind of another one is running on cpu or gpu so we actually do more compute but then we can very fast generate multiple tokens you know let's say token per second there is a metric called this tps uh if ordinary big model uh token tps could be let's say eight you know on a mobile device with speculative decoding it can go up to 20 40 you know kind of that you want to generate results outputs very fast you know but but you do some additional computation or you need to take advantage of parallel processing.
45:46That is also a fact. By the way, I'm not one of the authors of this paper, but that is how I understand, you know, I will truly invite everyone. And one of the, a couple of authors will be at the conference, you know, as questions, you know, if they are interested to the, our authors, you know. No, it's a really interesting idea. And in terms of the training of the model, Is it analogous to like a distillation type of training or? Yeah, draft model. You mean the draft model? The draft model, yes. Yeah, exactly. It is, there's the large model, big model, and we are trying to, we are distilling to learn this smaller model that generates similar outputs.
46:29Yeah. And distillation is something we do not only for, you know, speculative decoding, but also So for the large vision models, LVMs or BLMs, you know, we take larger models, bigger models, and we distill them to make them faster. You know, in this case, that is also the reason. It seems like this idea of running a distilled model in parallel with a full resolution model and then checkpointing frequently that that's a broader idea than, you know, this particular application of, you know, speculative decoding for language models. Is there something that exists like that in vision or in other domains?
47:10That's a good question. Yes. for instance, this is maybe older, more conventional computer vision networks. There is something called for video activity recognition slow and fast. You look at a bigger model with using much lower frame rate, the original video or smaller resolution frames but more capable model, big model. you run it from time to time and you run the other one, you know, smaller model, more frequently, and you look at how consistent their results, whether to regularize the, you know, kind of the small model outputs or combine them, fuse them. There is that. There is also hierarchical resolution trees, you know, kind of at these trees, you have the small version and you accept small version based on some confidence criteria.
48:12If you are not happy, then you will switch to higher resolution. It could be object detector, for instance, in this case, or a segmentation solution. These are more conventional things, but I see the similar analogies like having a draft model and large model or having a tiny model, but applying many, many times or a big model applying from time to time, you know, kind of something like that. And this thing is actually extends to end-to-end training for auto. There are now these models, very interesting areas. There is this genetic AI-based big models that are doing scene interpretation. And there is more conventional, let's say, vehicle detection, segmentation, tracking, like more conventional And these two, one is kind of high level, providing high level regularization to the conventional pipeline.
49:12And like both of them could be estimating the trajectories where the vehicle is going to be in the future. But the big one, of course, bigger. It cannot run real time, you know, kind of at this point. So it is not running all the time. And but they are running together, you know, kind of. So there are similar things all around the research landscape, I guess. The next paper on our list here is Segmentation-Free Guidance for Text-to-Image Diffusion Models. Yes. What's this paper attempting to do? So this paper applies to text-to-image generation, like stable diffusion. You give a text prompt and it's going to generate something.
49:55it doesn't require any fine-tuning of the model. But it does. When we run such models, diffusion models, we have the conditioning prompt and also a negative prompt. We do actually initial models. They had two runs of the unit, one with the text conditioning. So it is trying to generate things that you ask it to do. And the other one, without any conditioning, it takes the same feature maps and do another run of it. And that is called as, you know, kind of like a class free or, you know, kind of like a negative prompt. So this information applies to each patch of the image. You know, there's some negative prompting.
50:39But let's say we run the network once a couple of times, unit, And we now know a spatial region of the image is more representing a given object describing in the prompt. For instance, the prompt could be a dog sitting on a couch in the kitchen. I just made it up. And after a couple of first iterations, there are some image regions where you would likely to see a dog and other parts, maybe couch and maybe the background, the kitchen. So when we do negative prompting, we are saying the region also penalized for the dog region, penalized for the couch for the couch region. So this is kind of like, it doesn't sound right, right?
51:35We are trying to generate a dog in an image region, but we are also penalizing for that. So this paper, by just changing the attention map coefficients, these attention space are adjusted, you know, kind of solves this problem of, you know, I'm going to generate a dog, but I'm also penalized for the whole thing, you know, kind of. So it doesn't penalize us only for the dog, only for that region. So that's the beauty of it. And when you do that, Sam, when you look at the output results coming from, let's say, stable diffusion. And by the way, maybe I mentioned it, this thing doesn't change the diffusion model.
52:16It applies to anything. It's model agnostic. It's free. It's cheap. It's an amazing idea. So the output images, the quality look much better. What I'm hearing is that it is using this internal ability of the model to do implied segmentation as a way to allow the text guidance to happen in a more fine-grained way, essentially. So there are different metrics to measure how good it is. Some of it we just talked about FID and CLIP. There's Inception Score IS and there's also Peak Score, a network that approximately human learns how humans are responding. Someone sits down and trains such a neural network because that part is automated.
53:02So we look at all these kind of metrics and also we run our own human studies. But we observed that, so here is the base, original model, which is a very good model also. It could be, like I said, stable diffusion 1.5, 2.1, STX, anything. Here is the model, the attentions are changed. Model architecture, coefficients are everything the same, like we are doing segmentation-free. So we asked people, what would be the better one for you? And 60 % of the people with high preference said that I like segmentation-free. And only 19 % said that I like the original model. So the gap is 19 to 60. There's a big indicator that people really like the results coming from segmentation-free.
53:53And also in this paper, what we did that, okay, you want to do this kind of evaluations. of such networks and the input is just a prompt, right? And usually there are these prompt data sets like the MS COCO 30K, which means that there are 30 ,000 prompts in the camera. But then, you know, taking all of these prompts results, like for this model A, model B, model C, and showing to people, it is a big feat. It is very expensive. So you cannot ask, I mean, a person give me your feedback for 30 ,000 prompts. Then in this paper, we are also talking about how to make it better to keep the number of human evaluations manageable, not like 30 ,000, 100 ,000.
54:43But still ensuring that the selected subset of these prompts are representative in terms of the content and fear in terms of the model performance. So in this paper, we are also talking about this subset, like a benchmark plant, data set. So we just made it through the generative papers. And there are even more papers on the traditional side. I don't think we'll be able to get through them at the same degree of depth or coverage. You know, one of the things I noticed is that your team is continuing to investigate the idea of optical flow quite a bit. That's something that we talked about last time.
55:34And it looks like several of the papers are kind of continuing to push on ways to do optical flow better. Absolutely. You are absolutely right. Last time we had another CVPR conference paper, District Flow was its name. It was optical flow. Yeah, we put a lot of efforts in optical flow because it's very important, necessary to improve the image quality. First, when you take a picture, you actually take multiple snapshots and then you do some kind of alignment to get rid of noise, you know, kind of, or improve the high dynamic range. motion information is essential in compression in anything you do in video, you know, understanding how everything is moving.
56:17We are trying to, in these two papers, the first one is the kind of occlusion and consistency area interpolation. In that paper, we are trying to squeeze all the information we can get from the available limited amount of optical flow data sets. Optical flow is very difficult to annotate, you know, let's say I have video, you know, I need to go at every pixel, find which pixel it went to, right? This is, it cannot be done manually. That's why there are lots of synthetic data sets and approximating, you know, computed optical flow manually. This is very, very challenging. So amount of data is limited, Sam, for that reason.
57:03So we are trying to do the best. And in this case, that's the difference from what we published before. We are looking at now, I mean, we do all type of augmentations. We are trying to develop the best models in terms of accuracy and all the leaderboards, you know, Qualcomm is there. But also in this case, we are now looking at, okay, I have two frames. What about if I generate intermediate frames? This is video frame interpolation, VFI, and then, you know, generate the corresponding optical flow also, and then let's use it in addition to the original frames, you know, T0 and T1. Now I have T0.68, you know, something like that.
57:49And this is now going to give me more information, more frames, when I want to go train the optical flow models. Two things in the paper, how you would get the best, these interpolated frames between two frames, and then how you would take it and use it to come up with a better optical flow solution, which is model agnostic actually. And then you've got companion paper in the same, well, not necessarily a companion paper, but it's another optical flow paper called Self-Cleaning Inversion. The first paper was about how to make it, augment data sets even with using intermediate frames. And also, you know, we also talk about a better way of self-supervised training.
58:40That's separate. Here in this paper, CyFlow, which stands for self-cleaning iterations or, you know, inversion. So we are aiming a very efficient optical flow model, which is also accurate. For instance, recent optical flow solutions, most of them are using based on draft-like idea. It starts with initial optical flow or noise and then iteratively refines it, you know, using the features coming from different networks. It makes it better and better. But if there's a problem, there's a drift, they cannot correct it. So in this paper, we are explaining, introducing two concepts. One is self-cleaning iterations.
59:28The other one is regression focal loss to look at where during the iterations in inference time, where the model is failing or has potential to fail because we don't know, you know, because this inference time, but we have indicators, you know, kind of this confidence course like this course to go change the way that we are waiting the loss in those areas and and also we are handling ambiguities and also we show that you know these methods are i mean the networks if you do that you know you don't need to iterate many times it is computationally efficient um it it can run real time you know kind of we have a real-time running version also on mobile devices.
1:00:18So those couple of papers are, well, the first one, the optical flow augmentation is the main conference paper. The side flow is a workshop paper. The next paper that I wanted to learn a little about is low latency neural stereo streaming, also a main conference paper. And it's a wonderful idea. You know, kind of the input is a left, right stereo. It could be multiple also. It doesn't need to be two camera, but stereo or video stream. That's the input usually. But other algorithms do is you can take the left, let's say, and then compress and then take right, compress separately. So obviously there's redundancy here.
1:01:06Some algorithms, what they do, they take the left frame and compress them. Based on the left frame, they compress right frame. So it's one way. And another way of doing, you know, do left T0, then do T2, then right T1, then T3. This interweave, you know, fancy things. But this paper talks about a parallel hypercodec between this parallel autoencoder architecture. and there's this bidirectional shift module, kind of designed to learn the correlation between these left and right branches. So the goal is to do compression together, you know, left and right is compressed together. At the same time, that's why it is low latency.
1:01:56And this is important low latency, Sam, because there are now cameras on the XR devices, headsets and also autonomous vehicles, you cannot just stream the original video, you know, bandwidth limitation and other things, you know, storage makes it, you know, infeasible. So you like to compress them, but you like to compress them with low latency, you know, very, very quickly. And so how you can do that, you know, how you don't create these dependencies from left to right, right to left or, you know, in a virtual fashion, but do it together and still achieve, for instance, bit rate savings. We show that you can save 50 % bit rate if you do it the way that we described in this paper on Cityscape datasets or, you know, 20 or more than 20%, 15 % on KITI 2012 or 2015 datasets.
1:02:54These are benchmarks. Yeah. And it only comes with maybe a little bit more 35 % additional complexity on top of the very simple single stream encoding. Yeah. This is a very interesting paper, useful paper, very practical paper. If you've got a left and a right stream and you just use conventional compression, you're kind of leaving the correlation between left and right on the floor, you're not taking advantage of that. And so you could do better from a compression perspective. And so that's why you want to have a stereo aware compression model. You are right. You actually said it better than me.
1:03:34You know, the thing that I was talking about is bidirectional shift module, this correlation between left and right to, you know, compute it mathematically and, you know, then take it out, minimize the redundancies. Exactly. That's what it improves the bit rate, bandwidth bit rate. But also the fact that we are doing it together, we don't rely on the first left frame compression, then wait for it, then do the right frame. That is more conventional approach. or do left stream than right stream. It's also kind of like, it is delay, you know, kind of. So this delay is important. We are minimizing this delay by removing redundancy.
1:04:27We are going to leave the remainder of the papers as an exercise to the listener. We're going to have them all listed on the show notes page. There are, as I mentioned, a bunch more. Or in addition to a number of demos we mentioned, or you mentioned in passing the demos that you're doing, are there any particular demos that you would want to call out? Yeah, absolutely. You know, please stop by this. Many of those demos are running on device, on phone, you know, kind of mobile devices. And one of them is something we did, LoRa. It's stable diffusion. LoRa is low-rank adaptation. For instance, I can take original model and make it to generate better faces.
1:05:13You know, it is fine-tuned on faces or it can generate, you know, kind of any kind of, let's say, cartoon characters, you know, better. It stylizes with some kind of internal style. But the challenge is that this lower, additional lower coefficients, I have a big model and a small model, lower coefficients, adaptations. If I twist them, if I add them, and then run them. If I change, then I change my mind. I decided to impose more of the style, less of the style. I cannot do it. I have to go back and refuse it again. And these are big models, kind of several billion of parameters. Doing that fusion and then loading it again to memory is going to cost me in terms of time.
1:06:04I cannot do it very fast. It's going to take a couple of seconds on device. So this first demo shows that you don't need to wait. You can really easily switch to change the degree of this, how much you're going to impose. And also, you can very easily switch from one lower to another lower and another lower on the flight. This is a very cool idea. The second one is a mobile, for instance, you can take a picture of yourself, anything, a thing, you know, and then any object, you can ask a question and then it will answer. For instance, you are taking a picture of, okay, what can I cook with these ingredients in this picture?
1:06:52And it will tell you what you can cook and then you can ask, is the dialogue model, how many calories will it be? You know, kind of something like that. And you can imagine, you know, kind of like, you can literally have a conversation with it. And it's running on the mobile. It's available as of now. This is something we recently developed. And there are other demos. For instance, the demos that goes to our autonomous driving stack. Very interesting. We generate new samples, augment the data. For instance, there's a scene and we put animals in the scene, virtual. They are not real. But when we put them, then we use that data to train our stack detectors that we show they are improved.
1:07:36And in some cases, it's very difficult to collect such data. But also things around, let's say, more core computer vision segmentation on device, segmenting anything and tracking anything on device. Or like I said, portrait relighting of the faces or generating avatars, you know, talking avatars on the phone or, you know, kind of AI assistant with the avatar face, you know, kind of, there are cool demos. Yes. All right. And before we wrap, you also, you meaning Qualcomm also coordinated a couple of workshops at the conference. Tell us a little bit about those. Yeah, absolutely. The first one is the first conference, you know, first workshop in this area.
1:08:25it is efficient large vision models, SCVPR. They are becoming very popular, but also, you know, kind of we more and more need them to run on edge devices. So we co-organize this workshop with many other, you know, companies and institution, academia. And please stop by, you will see, you know, how this efficient transformer architecture for large vision models or foundation models, generative models, multimodal models are kind of making their ways into such devices. And the second one is called omnidirectional computer vision. And this is a very important area because it applies to many use cases from augmented reality to any camera system for automotive surveillance, photography.
1:09:22And now these solutions that we look into in this workshop is focusing on such input coming from these omnidirectional cameras. And this is maybe the third or fourth issue of this workshop. Well, Fatih, thanks so much for taking the time to share with us all that you've been up to and what we can look out for at CVPR. Absolutely. I'm very excited about that. Thank you so much for having me, Sam. And this is a great podcast. Thanks for inviting me. Thanks so much.
From the publisher
Today we’re joined by Fatih Porikli, senior director of technology at Qualcomm AI Research. In our conversation, we covered several of the Qualcomm team’s 16 accepted main track and workshop papers at this year’s CVPR conference. The papers span a variety of generative AI and traditional computer vision topics, with an emphasis on increased training and inference efficiency for mobile and edge deployment. We explore efficient diffusion models for text-to-image generation, grounded reasoning in videos using language models, real-time on-device 360° image generation for video portrait relighting, unique video-language model for situated interactions like fitness coaching, and visual reasoning model and benchmark for interpreting complex mathematical plots, and more! We also touched on several of the demos the team will be presenting at the conference, including multi-modal vision-language models (LLaVA) and parameter-efficient fine tuning (LoRA) on mobile phones.
The complete show notes for this episode can be found at https://twimlai.com/go/688.




