In short
Podcast Notes: The TWIML AI Podcast - Episode #738
Episode Overview Title: Distilling Transformers and Diffusion Models for Robust Edge Use Cases Host: Sam Charrington Guest: Fatih Porikli, Senior Director of Technology at Qualcomm AI Research Description: The episode discusses Qualcomm's recent papers and demos from the CVPR conference, focusing on advancements in autonomous driving and depth estimation using distillation techniques.
---
Key Topics Discussed
- Introduction
- Fatih Porikli returns to discuss Qualcomm's contributions at CVPR.
- Focus on two papers:
- DiMA (Distilling Multi-modal Large Language Models for Autonomous Driving)
- SharpDepth (Sharpening Metric Depth Predictions Using Diffusion Distillation)
- DiMA: Distilling Multi-modal Large Language Models for Autonomous Driving
- Objective:
- Provide an end-to-end autonomous driving system focusing on safe motion planning in long-tail critical scenarios (e.g., accident scenarios).
- Key Features:
- Utilizes large language models (LLMs) for structured scene understanding.
- Combines LLMs' world knowledge to improve efficiency in low-level perceptual tasks.
- Moves away from modular systems to an end-to-end approach, optimizing all components simultaneously.
- Performance Metrics:
- DiMA reduces collision rates by 80% and waypoint trajectory errors by 40%.
- Capable of zero-shot learning in scenarios not present in training data.
- Interpretability:
- Offers semantic explanations for decisions made by the AI planner, enhancing user experience during autonomous driving.
- SharpDepth: Sharpening Metric Depth Predictions Using Diffusion Distillation
- Objective:
- Achieve both sharp and metric depth estimation from monocular images.
- Key Concepts:
- Combines generative and discriminative depth estimation methods.
- Generative methods provide high-resolution but relative depth, while discriminative methods give absolute depth but at low resolution.
- Methodology:
- Utilizes a difference map to identify areas needing refinement during depth estimation.
- Employs a UNET model within a diffusion framework for noise-aware training and prediction.
- Applications:
- Monocular depth estimation for various fields, including robotics and virtual reality.
- On-Device Demos
- Text-to-3D Mesh Generation:
- Generates 3D meshes from textual descriptions in under three seconds, showcasing real-time capabilities.
- Image-to-Video and Video-to-Video Generation:
- Efficient video editing and generation on-device, highlighting the potential for real-time applications.
- Future Directions
- Research Trends:
- Anticipates increased use of agentic AI for visual content and reasoning models extending to visual data.
- Shared excitement for the ongoing evolution in the field of machine learning and AI technologies.
---
Key Takeaways
- The shift from modular to end-to-end systems in autonomous driving is crucial for enhancing safety and performance.
- Integrating LLMs into autonomous systems can provide robust interpretability and decision-making capabilities.
- Diffusion models present a promising approach to improve depth estimation, merging the strengths of generative and discriminative methods.
- Real-time applications on-device demonstrate significant advancements in AI capabilities, making complex tasks accessible and efficient.
---
Conclusion Fatih Porikli's insights into Qualcomm's innovative research underscore the rapid evolution of AI technologies and their applications in real-world scenarios. The discussed methodologies not only improve performance metrics in autonomous driving but also pave the way for future advancements in depth estimation and other AI-driven applications.
For further details, refer to the complete show notes at [TWIML AI Podcast](https://twimlai.com/go/738).
Written by AI. May contain mistakes. Listen to the episode to check what was said.
Transcript
Automatic transcript. May contain errors.0:00The world knowledge representation capability of LLM's act as a regularizer or conditioner to generalize the solution. So we don't need to really go every time for each long tail specific, very rare scenario and try to learn them, model them. We want to harness the LLM's world knowledge with the efficiency of this vision-based low-level perception stack.
0:44All right, everyone, welcome to another episode of the Twimble AI podcast. I am, of course, your host, Sam Charrington. Today, I'm joined by Fatih Pridikli. Fatih is Senior Director of Technology at Qualcomm. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Fatih, welcome back to the podcast. You're somewhat of a veteran with us. Thank you so much, Sam. I really love this podcast and I'm very happy to be back. I'm looking forward to our chat. It's been just about a year since we last spoke and we are back once again to dig into Qualcomm's papers from this year's CVPR conference.
1:29And in particular, we'll be digging into a couple of them. the DEMA paper, which looks at distilling multimodal LLMs in an autonomous driving context, as well as the sharp depth paper, which looks at diffusion distillation for computing absolute depth maps and some of the new capabilities that that unlocks, which kind of speaks to multimodal foundation models. Of course, another unifying thread that is often present in my conversations with you and your colleagues at Qualcomm AI Research is the idea of efficiency and distillation. And that's a key theme in these papers. But I guess let's jump in and talk through the DEMA paper.
2:15The full title of that paper is Distilling Multimodal Large Language Models for Autonomous Driving. Maybe let's start with kind of your broad take on the state-of-the-art autonomous driving and in particular the shift towards end-to-end autonomous vehicular systems? DEMA is about autonomous driving, as you just mentioned, and specifically it aims to provide safe motion planning in long-tail scenarios, rare events, and those rare events are very critical events because those represent accidents scenarios as well. So we don't want them to help, right? Previous trends, it was mostly modular systems, they sequentially and independently trained models, components.
3:06And the reason for that, there are kind of smaller data sets and, for instance, for object detection, vehicle detection, lane detection, lane marking detection, road detection, road segmentation. So there are small data sets and it's easy to maybe train such component than, you know, kind of build a larger system. But when we do that, it is, you know, kind of easy to train such things quickly. But since we are now focusing at each module, their own performance goals for better segmentation, better detection, better tracking. it doesn't mean that at the end, you know, altogether they will work in harmony.
3:50And our goal is to drive safe, right? Drive our destination. I mean, when we are driving in our brain, we don't really segment the things, detect the things or track them explicitly. But so this is the shortcoming of such approaches. And end-to-end now is a breakthrough in a way that training objective applies to all components at the same time. So we are not really focusing on test-specifically trained modules on limited dataset. But now, you know, kind of we are trying to optimize everything, all the processing with the end goal in mind. That's why it is called end-to-end. Frequent listeners to the podcast will know that this idea of physics-based versus model-based or modular versus end-to-end is kind of a battle that's been raging in the machine learning community and these various communities like autonomous driving and robotics for quite some time.
4:54We've been talking about related themes for many years here on the podcast. There is a big reason for that because end-to-end systems doesn't only offer better, more robust solutions for a wide variety of long-term scenarios, but also they are better in terms of interpretability and their KPIs are much better. So DEMA is an end-to-end solution. It establishes the new state of the art, literally 25 states of the art for autonomous driving and AI. Can I pause you there? Because you said something that when I think about modular versus end-to-end or when I reflect on the conversations that I've had with folks in which this comes up, it's usually positioned as given enough data, we think it will perform better because we're optimizing against the thing that we really care about.
5:57But modular is still useful because of explainability and interpretability, because you can look at, you know, the output of each of these modules and use that to get some intuition about why decisions are happening at the planning layer. But you just said that end-to-end system has some interpretability advantages. Can you explain that? That interpretability is more the one that I implied is semantic interpretability. For instance, we have AI planner. There is a planner at the end, right? We do all such low-level perception things that I mentioned to make a decision to whether accelerate, slow down, change lanes, you know, or turn, those type of things.
6:41And when you are sitting in a vehicle, you know, kind of, yes, like maybe modular system can tell you how many people are in the scene or, you know, kind of what are the locations to the positions of the kind of other vehicles. But kind of when it is, it decides to change the lane, you know, you don't know why it's happening. It is a very stressful experience if you do not know why vehicle is making such decisions. So Entrance Systems and DEMA in particular, the way that we kind of like incorporated the language model into overall system provides explanation. Look, I'm slowing down because there's congestion ahead or we are approaching a zebra cross.
7:30So analogous to explaining thought traces in a reasoning model, the model can to some degree explain what it's doing. Yeah, high level explanation, useful explanations. But you are right also, you know, kind of like the modular system, we know kind of their detection results for instance, the thin advanced system. And sometimes they do not need to reveal that because maybe that is not critical at that point. And that's why I described it as a bit of a battle between these. Okay. I'm trying to be like a judge, but honestly, I'm more kind of supporting end to end at this point. Yeah. All right. So, sorry, I interrupted you.
8:20You were talking about... I was talking about DEMA's kind of position advantages. So for DEMA, I mentioned it is the new state of the art. It can reduce the collusion rate, which is a very important KPI for autonomous driving system, 80%. It's respect to the 2024 or latest 2025 baseline SOTA solution. So it is a big improvement and also 40 % reduction in, for instance, waypoint trajectory estimation and other KPI. Let's say I know how vehicles should drive, how close is planner making better decisions. So that is also kind of a 40 % reduction in the error rate. So more than 40 % improvement in accuracy for long tail.
9:15So in the paper we show, it's a long paper, and we have many results, comparisons, and, you know, kind of I invite everyone to take a look at that. So it says the new SOTA, we compared with the previous SOTA, including VAT, Vectorized Autonomous Driving, which is an amazing paper also. And this NEMA can also do two things that they are not in the training data. For instance, there was a zero-shot scenario about three-ponged turn. We didn't have any such example for planning, but then when we gave it for testing such examples, it was capable of, you know, kind of doing the right planning. And then I mentioned that it can answer questions or the system can proactively, you know, kind of system prompted and then provide explanations while it is driving itself.
10:11You know, it can tell you why it is making such decisions. So overall, it is, you know, there's a low-level perception state. There's an end-to-end planner, AI planner state. They work together. It integrates all the perception coming from the cameras and other sensors, the vision-based front-end low-level perception into this language model, LLM-based, you know, kind of AI planner. This improves generalizability. It improves the robustness. It is now the explanations are semantically grounded. And also, we have two versions. You know, we can run this AI planner, a transformer-based model, still leveraging the language model or language model itself, as the overall AI planner.
11:03So both of them are in the paper. Both of them are better than the existing SOTA. Is the transformer-based model the distilled model? That is the distilled model, right. We are using the model to distill it, yeah. When I think about the idea of using LLMs in an autonomous driving context, the elephant in the room, if you will, is that these things are slow, right? inference for LLMs is very expensive. And I'm presuming that's where the distillation in DEMA comes in. You know, kind of. Yes, absolutely. You have a very good point, LLMs. At this point, of course, we are talking about the systems that we can run on the vehicle, not like on a very expensive H100, you know, GPU.
11:56that would be more expensive than the vehicle itself. So using this reasonable price accelerators, the token rates are not that high. Maybe I can say that depending on the model, size of the model, anywhere from 30 token per second to maybe 1 ,000 token per second. But running everything end-to-end using an LLM would be much more compute intensive than the transformer model. That's why we have the transformer models. We are saying, well, if you are interested in efficiency, here's the transformer model. Still leverage this still using the language model, but there is also a language model version of it, a multi-model version of it.
12:40You are right. Let me take a step back and kind of recap where we are. So we've got this autonomous driving problem. We're moving quickly towards autonomous driving solutions, a promising direction for autonomous driving is to introduce LLMs into the autonomous driving loop. Why even LLMs? Given we kind of jumped ahead and talked a little bit about the challenges or the cost of LLMs, like what is the role of an LLM in this loop and why do we even want it? Like when I think about what I'm driving. I'm not verbalizing questions to myself about what I'm seeing or what I'm doing. Why do we think an LLM should even be part of an AV system?
13:29Yeah, that's a good question. We were also asking the question, this question to ourselves. And the reason we want to use LLM is that they represent world knowledge, many things. So they capture all the knowledge, or most of the knowledge, yes, they hallucinate at the end when they are generating this muscle, but they are much better than the things that we manually crafted. So there is something there that would make it richer. The word knowledge representation capability of LLMs act as a regularizer or conditioner to generalize the solution. So we don't need to really go every time for each long tail, specific, very rare scenario and try to learn them, model them.
14:27We want to take advantage of this world knowledge of the LLMs to do it automatically. So that's why we want to harness the LLM's world knowledge with the efficiency of this, you know, vision-based low-level perception stack. And so is that world knowledge access via traditional prompting, or are we talking about doing something like taking the embeddings of an LLM and somehow directly accessing them to unlock this world knowledge? By the way, DIMA is not the first end-to-end solution. There is UNID and VAD. I mean, it's, you know, let's say big brothers, but we do, I think, better than them by incorporating LLMs there.
15:16So those models, they try to represent the scene in terms of vectors and replacing these, you know, kind of dense representations to reduce the cost and also, you know, kind of like provide better improved capability. So there is this thing called scene representation and I will kind of, we can also consider it as tokenizing. Tokens are very common term for language models, LLMs. So what goes into those models separately, let's kind of think them as tokens and visual data. And this visual data is coming from multiple cameras, six cameras, eight cameras, and for each camera or better we are taking them and low-level perception projecting them into a bird's-eye view map and then we have things that ego vehicle does and then we have things for other agents other vehicles and the things moving in the same pedestrians and also there is map we have a map so altogether so let me hit pause there so the ego view is kind of this bird's eye view and that is familiar.
16:31Yeah, I've got that in my car. Like it's got cameras on the mirrors and in the back and in the front and it kind of projects into this top down view. Top down view. That's kind of the... Exactly. So top down view vehicle, ego vehicle is in the center and everything is around. Looks like 2D from top down. And so the other, did you say the view from other vehicles? No. the tokens or their kind of not necessarily the views, but the useful information that they kind of provide into. So maybe localizing other things in the scene or something. Yes. And also their previous trajectory, their headings, that type of information, not necessarily available in one bird's eye view picture, but because they are about other vehicles, like maybe also how they intend to drive, which is not available in Berzavu.
17:36So this system is not only, you know, looking at a frozen Berzavu, but, you know, thinking on behalf of the other vehicles, agents, how they may move. Are they accelerating? Slow down, you know, do something unexpected. So we are trying to anticipate what the other person is going to do, other vehicle is going to do, like a real person, right? We are on a traffic stop, and then I'm looking at the other vehicle, I'm thinking it's going to go ahead, or it's going to wait for me to go ahead. So in other words, the LLM input is kind of a string of tokens that represent this top-down view from the perspective of the vehicle, but also almost like other features, you know, representing the vehicle's trajectory, what it thinks other agents in the scene are doing, whether those are vehicles or pedestrians, and like somehow mapping out the scene.
18:36So it's like a featurization of the space that the vehicle is operating in. Exactly. That is what we call as scene representation and tokens. Everything now kind of standardized in terms of tokens. We have this ego vehicle tokens, you know, coming from its own Qformer and then other vehicles tokens coming from their Qformer and Intense. And there is this bird's eye view scene at that. We are tokenizing it also. And there is MAP also. There's tokens for that. But they are not all the tokens. When we are training with LLM, these are going in, but then there is also a question, right? Or a prompt, text prompt going into there.
19:22And then this language model, what it does, provides an answer. And then we are, you know, giving in training time, we have the answers. We have this, you know, scene representation tokens and we have the question. So we are updating this language model. At the same time, we are also looking at these tokens are also updated. You know, this ego or agent or bird's eye view or map tokens, they are also updated. These tokens goes into language model. And then within this language model, when we are running layers, each layer has cross-attention and also self-attentional layers. We self-advice those tokens.
20:11So in a way that you can think that these tokens are now changing. In a way that they are more consistent and they are more useful and they are answering, for instance, this input prompt better. So at the end, we will have updated representations of those tokens. So some raw tokens goes into and then better tokens comes out. So those better tokens are used to update two versions again. Transformer versions, we can distill the planning transformer to capture, you know, how planning transformer will have some additional compute. So raw tokens will become better tokens. And then the other tokens will allow us to make better decisions.
21:02So that's why we are improving to performance these planning constraints. And I should note that there's a really good image or figure two in the paper that is kind of a block diagram of all the things that we're talking about. And hopefully if you're watching the video, we will have dropped that in so that you can follow along. Certainly if not, or if you're watching the audio, then we'll be including a link to the paper in the show notes and you can pull that up. part of what I'm trying to do here is reconcile what you're saying with the image which is you know it's obviously a static thing and so I'm trying to understand how the things that you're saying happen in time like so you've got you know I think your data set consists of at a given point in time, six images from the cameras on a vehicle, right?
22:04And you just have many of those. Is there additional data associated with that or is it just those images? We represent all of these current six images and let's say T minus one, six images. It could be T minus N also, but it looks like T minus one is sufficient. In terms of their tokens, there is this scene encoder And then... That scene in Kodo is kind of like the feature extractor that we talked a little bit about, right? It generates tokens, you know, at the end there are these tokens for the six images, current images and then previous images. We have those tokens already in the LLM. So again, two versions to kind of distinguish.
22:48If LLM is making the AI planner decisions, Once those are in the KV cache long context of the model, so model retains those. And then when it is going to make, for instance, a waypoint estimation or a driving decision, a planning decision, it will leverage those previous tokens up to the buffer size of KV cache. And then the transformer version, we use previous token instance and current tokens and update this planning transformer. It's a transformer, but it's not like language model, but transformers are the building blocks of language models, by the way. It's not explicitly asking a question, you know, it doesn't have that capacity, but it's with tokens.
23:38So those tokens, what I mentioned, are now in training time, we learn better tokens. And now we take it and we learn an adapter, you know, kind of for this planning transformer. So next time in the transfer... Meaning so you can get the tokens that come out of the scene encoder and give them to the planning transformer directly and get a prediction as opposed to going through all of the machinery of the LLM. Exactly. So that version is transformer version is more efficient. Now, C9 encoder, again, generates the raw tokens. It goes through this adapter and then it goes to the planning transformer.
24:17And that is, there's no LLM anymore in that version, but, you know, kind of a better transformer for planning. And this is the core idea behind distillation in general. It's like you've got this, you know, surrogate model in this case, a transformer or the student model that I'm referring to in this case that is, you know, it's a function approximator. And so the distillation process is kind of teaching it how to approximate, you know, the function. In this case here, the teacher model, it's output tokens given the beam tokens from the scene encoder or the function that we're trying to approximate.
25:02Absolutely. Of course, the devil is in the details, how you do the selection, how you do surrogate tasks, how you do VQA, is it part of it? What are those data sets and what are those functions? We tried many things and more details are in the paper. You mentioned surrogate tasks and VQA. DQA. What are surrogate tasks in this context? For instance, they are mostly around trajectory prediction. Some of them are about also mass token prediction, future BV token prediction to encourage the LLM to learn spatiotemporal cues, useful for planning and scene editing. Scene editing is a part of it because we have to hypothesize where the other things, agents, are going to be.
26:01So, I mean, this is in training time. How the surrounding agents, other vehicles will impact the vehicle's future path, you know, kind of. So this is what I meant by scene editing tests. But that's a part of the surrogate test or trajectory prediction. and there are a couple of more also. And so they're not necessarily tasks that you want to use, but rather they're exposed as tasks so that you have a separate loss function that you can incorporate so that the model attends to those things. Is that a fair way to think about it? Exactly. We are not explicitly saying here's the trajectory, here's the plot form equation.
26:49We are using that to make sure that the temporal aspects of drawing is captured through those surrogate tests, like prediction. I said prediction, right? The scene itself, the other agents, and these are future predicting. And also leveraging on what kind of it observed before. So the model would learn a better representation itself. But it is not like we are telling a model to go optimize for these future instances, something like that, or change it. It is through the training, it learns itself, changes its coefficients to better understand the dynamics of the traffic scenarios. And what's the role of incorporating a VQA component, visual question answering?
27:40Yeah, we want this capability, that's an important part also, this work knowledge to be a part of overall kind of solution. We do not want to just drive, steer NLM towards surrogate tests, you know, the feature prediction or, you know, generating better tokens for the planner. But we want it to be grounded also. So that is keeping it real, connected to the semantic information. For instance, those questions could be, okay, in this scenario, what are the safe actions to take for ego vehicle? That is literally the question and the answer. It could be the action is to brake gently to a stop. So for this input, we want model to be still able to generate such answers.
28:38given the prompt and all the visual data. So that is going to allow us to retain the semantic word information of this LLM without destroying the LLM. And also the second version, we can ask questions and then it can answer, you know, kind of, again, two versions. Transformer version doesn't answer the questions, but the LLM versions in inference time can answer the question. So it has two purposes, keeping it grounded, semantically grounded, still not destroying the world knowledge, but then also having capability to answer questions if you are running LLM as the planner. uh and so you've talked about the kind of the the outcome here so the state-of-the-art performance on some of these long tail on both trajectory uh error and kind of long tail collision detection or avoidance.
29:44What you didn't mention yet is any kind of SLAs or performance around the distilled model. Is it fast enough yet to be able to use in a real-time AV loop, or does more work need to be done to make the model more efficient? Distilled model we can run as of now on Qualcomm accelerator faster than real-time. It's a new solution. I mean, this paper is not for publication research sake, but it has a very practical purpose also. We hope to have a vehicle running around soon and showcasing this technology. Most likely it's going to be at CES, But, you know, kind of, I don't know the future, you know, kind of, but it's there.
30:45Given what you've done with this paper, where do you see this kind of direction of research going, this particular line of research? What's next? So we see that, you know, such models are useful. And other companies are also looking into that. And I see one trend of running both transformer and language model together. Language model is, as you know, kind of as we talk about, is slower, compute intensive. So maybe not that fast at this moment. So it is on a longer horizon, you know, in a lower frame rate. It is estimating, doing its own thing, you know, estimating waypoints and, you know, some useful AI planning decisions that the current AI planner using Transformer is also running on 3D or 2D perception stack.
31:35And then these are combined. So there is redundancy in the system. If it is something unexpected for the low-level perception stack, at least we have high-level understanding that would provide some safety mechanism for the long tail. So this is like a trend. But many companies, research labs are looking into that. So I think it will become more popular. But as important as this one, using language model in the vehicle to interact with the vehicle, decision mechanism, change its behavior, customize for a geospatial location, like for a city or for a country, because rules are different and driving dynamics are different and preferences are different.
32:23Maybe you want more aggressive driving or more conservative, you know, kind of. Yeah, they are also happening. I think kind of, yeah, I can say that we actually, we are in the process of having an even more capable version of DEMA, which is domain adapting to some of the scenarios that I mentioned. Very cool. Very cool. All right, so shifting gears to the second paper that we want to dig deep into. Oh, deep. No pun intended, is sharp depth. And the full title of this paper is Sharp Depth, Sharpening Metric Depth Predictions Using Diffusion Distillation. So distillation, again, a theme here. My impression is the general space that this paper is exploring is monocular depth estimation.
33:20So you've got a 2D image, a photo, and you want to essentially reconstruct it as a 3D, in a 3D space. Is that the right way to think about this? Very good, yeah. You said it very well. And so talk a little bit about, you know, the background here and state-of-the-art prior to this work, and then we'll dig into what this work continues. Absolutely. So this is monocular depth estimation, but more specifically, this paper chart is about metric depth estimation. You don't know whether it's a tiny chair or a huge chair or, you know, some reasonable chair. So there are two ways, two separate, you know, type of algorithms.
34:05One are maybe, let's call them as generative models. And those are fine-grained. they provide very sharp molecular depth estimation because they can use synthetic data. There's a lot of such data like Marigold and Lotus and many molecular depth estimation algorithms are like that. So they have also dense ground truth because it's synthetic, it's there. But their skills, it is point. It is not like inches or millimeter or anything like that. You don't know what that point each pixel, how big it is. It could be one meter or one millimeter, you know, very different. So there are... It's relative within the reconstructed image, but it's not absolute.
Read the full transcript
35:00So you can't measure with it. Yeah, it is relative. So those are generative solutions. And there is discriminative solutions like metric 3D and Unideb. And when they do provide these estimations, depth estimation, they either inches or millimeters, centimeters. However, they are trained with such data doesn't exist in quantity because usually LiDAR sensors, images and LiDAR sensors or structured light sensors are used. So first of all, they are very low resolution and not big in quantity. It's real measurement. So there's a challenge. So when you use both solutions, you get kind of correct depth estimation from the camera or 3D accurate in terms of scale.
35:55However, they look like very blurred. In the paper, you may see some examples, you know, all the details. are missing, if there is a chair, legs will be missing. Genitive method, because of synthetic data they are leveraging, they will show very nice details, very fine, but we don't know how far away the chair or how big the chair is. But these discriminative methods, they will give exact, absolute distance, accuracy is great, but then very low resolution. Maybe you will not even see the chair. So there are two kind of existing solutions. So this paper, Sharp, that bridges these two approaches integrating metric accuracy with the detailed boundary preservation of these generative methods.
36:45It is a generative method. It is taking advantage of, for instance, diffusion, denoising diffusion. and it has actually in the architecture, it has a UNET model running in the background. I'll jump in to note that anyone that wants to dig into UNET a bit, I think we talked about that in our last CVPR review and maybe even the one before that. I think we've been, UNET has been a theme that you and your research group have been working on for a while now. So UNET, actually kind of the name, U-shape, is the shape of the architecture. It has been very popular. Of course, there is also a trend of now using DIT, diffusion transformers, which has a different structure.
37:39But UNET is still one of the choices for genetic AI, for visual data. And there are lots of applications of this. So Sam, you may wonder why we want to do monocular depth estimation. Should I talk about this? Yeah, I'm imagining something like, you know, I'm in a room and I want to take a picture of something with my camera and know how big it is. That would be pretty cool. Absolutely. Then you can place, you know, for instance, you go by from a furniture store and you know how exactly it's going to fit. It's not relative because now you know the actual measurements. So this virtual tryout, room planning, object placement, 3D model reconstruction, you can use your font to create a real size with metric absolute scale models.
38:38Of course, people are using for robotic navigation because we know exactly how far away occluders and other objects in the scene. and robots are interacting with the scene, when it is reaching out something, it knows where it is exactly, not how many pixels. It doesn't know whether it's far away or near. Yeah, there are a lot of applications in, for instance, immersive property, walkthroughs to understanding things in 3D space, even surgical assistance, you know, kind of you can do minimally invasive procedures procedures using a single camera and because they need to be very tiny right you may not be able to put like two cameras a bit white baseline into a very tiny blood vessel so but you can squeeze something like a camera single camera but then you know how far away everything is in the sense of maybe i mean these are real applications like drones or you know kind of So they say monitoring to construction sites and other type of things.
39:45Parking, automated parking is a part of it also. You mentioned that part of what's happening here is you're kind of fusing a generative approach with a metric-based approach. And it's still not clear to me. like it seems like you need some piece of absolute data somewhere in order to start to do this like you know if you're taking a picture with the phone you know you might need some like lidar you know that says the distance from objects or like maybe stereoscopic where you've got the distance between the lenses or some piece of uh piece of additional information from the real world to ground you.
40:35How does this method overcome that gap? Yes. So we are running two things to start the overall process. One is this kind of discriminative metric depth model is running. So we run it because it exists, right? Such models I mentioned. We run it. we get an estimation of everything in the scene, their depth estimations, but those are metric estimation, but it is just not high resolution or, you know, kind of sharp enough. Then we also generate the other one. Is it clear, you know, there is such an algorithm already running and this sharp depth can use any of those, you know, This is an overall methodology to incorporate these two different types of models.
41:30We are not reinventing the wheel, but we are putting those wheels in a way that they will go alight and whatever we are driving goes better, faster. So you've got this course measurement that we know how to do, and then you've got the fine-grained relative thing from the generative side. and the idea is that the latter is good enough or, you know, they're each good enough so that together you can get fine-grained depth predictions if you put them together in the right way. Absolutely. And the thing that you mentioned right away is the big question. How I'm going to know that each one is each because I don't know this thing, you know, One is blurry, the other one is sharp.
42:27I don't know which one to trust. So what we did, our intuition is, we take these two estimations, and then we compare them. We have this adaptive substitution, scale-level substitution, which generates a difference map. So the intuition is, in this, let's say, difference map, the regions with minimal differences, are more reliable in terms of their metric depth estimation. And while other areas with larger kind of differences will require maybe updates. So this is actually kind of allowing us to, I mean, keep in mind that this is a diffusion solution with UNET. and UNet, when we run UNet, there is the input.
43:25And for text-to-image generation, input is just Gaussian noise. In this case, it is not. In this case, it is this, let's say, pure Gaussian noise. It is changing depending on the difference map. We are conditioning on this, so minimal differences. We don't like the diffusion process to change them. Big differences, we like it to change. And then we learn in training time a model that takes such difference map and makes the best use of it to refine itself. So there is fine tuning going on in training time. So we have one model. We run the model. We generate this fine invariant depth, but that is not good enough.
44:19We have the kind of metric depth version. So we compute this difference, differently face how much diffusion is going to make a change in the depth estimations. And in training time, this is optimized because we have, you know, kind of the, even though low resolution, we have some metric estimation. And also, we can use existing metric data sets. We also show in the paper that you don't really need a lot of it. We are using maybe 100, 150 times smaller than the amount of data used to train such discriminative models. So kind of in training time, we learn a better diffusion model, and then we just use it in inference time.
45:11So overall, take a look at the paper. It is accurate. It is shorter, you know, in terms of the resolution, but it is a metric estimation. Recently, I've been seeing a lot of interesting applications of diffusion models beyond kind of the, you know, stable diffusion image, you know, text to image. and they all seem to revolve around creative ways to manipulate the noise to do interesting things. And so this is that same kind of idea. Yes, in this case, we are controlling the noise. You are absolutely right, by the way. We can use Diffusion framework to make a robot's planning or trajectory estimation or solving a puzzle.
46:03Diffusion has many applications. In this case, we are controlling to know if we are right. Ultimately, is this supervised or totally self-supervised? Self-supervised, because we have this metric depth model. Even though low resolution, we can, in the trim time, we can just use to learn, you know, a better shock net. But this thing can also, like I said, can use leverage much less supervised data. So we can do SFT supervised fine tuning. And we show, you know, the difference in the paper, you know, how they will compare. Both of them are still, you know, providing very sharp depth maps, monocular depth maps, but metric also accurate depth maps.
46:49So you show both supervised and unsupervised versions of the model? Exactly. And the reason we showed both of them, so people can go, you know, if they design their own metric depth, they can also incorporate in this framework, leveraging the ideas that we talk about in the paper, and then they may come up with even a better version of, you know, kind of a sharp depth. But the idea is, I mean, this paper facilitates further research in that sense. Very cool. Are there additional areas you foresee in terms of future research along these lines? Of course. This is monocular depth estimation. But you don't do monocular only if you have multiple cameras or if you have video, right?
47:41You can do structure from motion if there's video or multi-view stereo or multi-view depth estimation. But the whole idea would be the same. The mechanisms like how this noise aware gating for using the difference map is explained in the paper. I didn't go into those details. Maybe a little bit math heavy and I may need a book. There might be, of course, extensions to video and multi-view as future work. We are actually looking into that and also maybe kind of alternative ways of computing these noise maps and cross-attending them into the model. Yeah, this is active research, you know, kind of iterating.
48:31Sometimes they generate amazing results. Sometimes, you know, kind of they ask to learn and, you know, refine the idea. Great. So in past years, we've talked about, you know, a broad set of Qualcomm's research and demos and workshops and tutorials in our CVPR show and touched on a bunch of the individual papers. This time we wanted to go deep into a couple of papers, but there's still a lot of other, you know, papers as well as activities that you presented at the conference. You know, let's maybe choose a few of those to touch on and then we'll refer folks to your blog post to get the full picture.
49:25You know, what did you do for demos this time? In addition to 11 papers, we had maybe more than 10 demos. With me, I'll talk about three of them. The first one is going to be text-to-3D demo. The second one is going to be either video-to-video or image-to-video, JNATVI demo. These are all multi-model JNATVI demos running on device. And the third one is going to be multimodal VQA, if you have time. And text-to-3D is a model literally user speaks through audio interface or text prompt. And you describe an object like a cactus sword, you know, kind of a hippo wearing a sweater, you know, anything you can imagine.
50:13And it will generate on device a 3D mesh and also associated texture map, like color, everything, in less than three seconds. I mean, such models, we are not the first to come up with such models. But in terms of the existing models' computational complexity, there is an amazing model called MVDREAM. Yeah, it takes 194 minutes, not seconds. And there's DreamFusion. I mean, maybe quality-wise, you know, maybe we are much better. it takes 22 minutes, you know, kind of. And now we are saying that, hey, you can do it on your phone in three seconds. And the demo, let me describe the demo. So it is a game, video game, and then the player goes into this table.
51:10And then this is a game where, you know, the person is fighting with other things in the scene, you know, kind of like monsters and other people. and it says, it describes it's Batman, maybe a battle axe or sword. And that's what I imagine about, you know, kind of a dragon or cactus sword or watermelon sword, you know. Then it generates things accordingly in less than three seconds. Then the person grabs it and then continue playing, killing the monsters in the scene. That was the demo we were explicitly showing, but the thing that it can generate anything and this is the first on-device text to 3D demo literally the first time we also showed and it allows offline personalization at the edge you know it uses fusion models again to image generation in particular SSD there is SDHL version of it as well both of them generating high resolution data so it generates either one image or six images multi-view then And we run another unit, another SSD, to take it and map it to a mesh.
52:28So that's the text to 3D. For image to video and video to video, there are two JNTF AI video generation demos. We showed both of them next to each other. So video to video is like video editing or stylization demo. it takes text as an input conditioning. It could be, let's say, I have my, you know, selfie and then I see that, oh, change it to, let's say, pencil drawing style or moment painting style. But it's not about faces. It could be anything, any object, anything or make it look like Albert Einstein, you know, kind of looks zombie, you know, The challenge, of course, is how to do it in a way that it is temporarily consistent.
53:22So this is another unit-based model. It runs quite efficiently on device. It runs at 12 frames per second speed, generating 512 by 384 video frames. So it's an example. I can say that when we started this model, the original model was taking seven seconds. And there's an existing model, you know, kind of someone outside Qualcomm developed it. When we put it on device, it was seven seconds. But after we optimize, and I can talk about how it's done, now it is running around 80 milliseconds. Like this is 90 speed up. Yeah. This is the video-video stylization. Image video is even more aggressive. When we started, you know, kind of, it was taking 2200 teraflops.
54:21We decreased more than 500 times. How does that compare in terms of time? Time, the original model even didn't fit out of memory and very slow. I remember playing with one of the demos at a Google event. I think it was their new Veo 3 model and image to video in particular. And I mean, it looked great, but it took a really long time. Like it was, you know, go away and we'll send you a notification when it's done kind of timeframes. To hear this happening on device is a whole new idea. And it is happening almost real time. That's crazy. Now, the images, it sounds like they're a little smaller, but...
55:10Image, the video demo we had, it is generating 10, 24, 5, 12. It's a funny resolution, but it is not quite a high resolution, actually. Without any video frame interpolation, we can go to, you know, very fast interpolation to increase the frame rate. But without that, we can generate two seconds of video in three seconds. So you can imagine, you know, this is very fast. And yeah, but the challenge is, again, these models are big models. They are compute intensive. How to heat it down and run efficiently on device. But in Qualcomm's secret sauce. I mean, many things we do are aiming for power efficiency, speed, memory efficiency, making sure people can use on their own device.
55:58They don't need cloud connection or anything like that to generate such videos or content. Very good. So when we're back here in about a year, what should we expect to hear from you in terms of things that were hot in the next 12 months in computer vision? What are you excited about? There are two things. Two wildcards. You know, I can say that, you know, I'm betting on them. One is, yeah. If it is the correct term, you know, kind of, or two jokers, I mean, I'm really excited about it. One is agentic AI extending to visual content. The second one is visual reasoning models. So the agent AI is, I mean, system capable of autonomous planning, reasoning, and interacting with environment and user, multi-turn, like AI assistant.
56:56But then such systems, for instance, it is running on XR Glass, you know, Ray-Ban. And then you are, you know, giving it, okay, make a dinner reservation for the people in this scene based on their preferences. You know, this is literally the question. And now the system has to go understand who are those people based on exchanges like text or emails or anything we did together before. Understand their preferences in food, what kind of cuisine they like, Italian or sushi or something like that. And then go and find a restaurant and then make a booking based on their availability. It has to access their calendars.
57:44So this is happening. We had an agentic AI demo, I didn't mention. Components are there. So, but it is now using, I think next year we will see a lot of things through screenshots or through cameras. Such content will be a part of agentic AI. Second one is reasoning models. You may remember Sam Dipsick made a huge, you know, kind of impact on how people are training models. They should say that, okay, you don't need huge compute to train models. And these models now are going to, smaller models are going to perform as good as big models. Because what they do in inference time, they don't just give an answer, but they think.
58:27I'm innovated. They generate tokens, respond, respond. And they say, oh, wait a minute. Maybe there is another train of thought. Maybe I will follow it. So this is like chain of thought models, view of thought models or graph of thought. There's a lot of XOT of thought models happening. But now what I see that they are extending to visual data. It's already happening. There is multimodal COT. There is compositional COT. Many things are coming. We are looking into that. What does it mean for visual data to do reasoning? So I think it's going to be the other area, very exciting to see. we will see who is going to get the best paper, you know, in those two areas.
59:14Awesome. Awesome. Very good. Well, Fatih, thank you once again for taking the time to share what you've been up to and some of your team's papers from CVPR. It's great to catch up. It was a real pleasure. I'm very excited about what we do. And it's always great to talk to you, Sam. I love this podcast. Thanks so much. Thank you.
From the publisher
Today, we're joined by Fatih Porikli, senior director of technology at Qualcomm AI Research for an in-depth look at several of Qualcomm's accepted papers and demos featured at this year’s CVPR conference. We start with “DiMA: Distilling Multi-modal Large Language Models for Autonomous Driving,” an end-to-end autonomous driving system that incorporates distilling large language models for structured scene understanding and safe planning motion in critical "long-tail" scenarios. We explore how DiMA utilizes LLMs' world knowledge and efficient transformer-based models to significantly reduce collision rates and trajectory errors. We then discuss “SharpDepth: Sharpening Metric Depth Predictions Using Diffusion Distillation,” a diffusion-distilled approach that combines generative models with metric depth estimation to produce sharp, accurate monocular depth maps. Additionally, Fatih also shares a look at Qualcomm’s on-device demos, including text-to-3D mesh generation, real-time image-to-video and video-to-video generation, and a multi-modal visual question-answering assistant.
The complete show notes for this episode can be found at https://twimlai.com/go/738.




