High-Efficiency Diffusion Models for On-Device Image Generation and Editing with Hung Bui - #753

28 Oct 2025 · 52 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Episode Summary: High-Efficiency Diffusion Models for On-Device Image Generation and Editing with Hung Bui - #753

Podcast Overview

  • Podcast Title: The TWIML AI Podcast
  • Description: The podcast discusses advancements in machine learning (ML) and artificial intelligence (AI) and features insights from top industry experts.
  • Host: Sam Charrington
  • Episode Guest: Hung Bui, Technology Vice President at Qualcomm

Episode Highlights In this episode, Hung Bui shares insights into the latest advancements in high-efficiency diffusion models for generative AI on-device, specifically focusing on image generation and editing. Key topics include:

Background of Hung Bui

  • Education: PhD focused on multi-agent systems.
  • Career Path: Experience at notable institutions like Google DeepMind and Adobe Research.
  • Return to Vietnam: Established the first AI research lab in Vietnam, VinAI Research, which was later acquired by Qualcomm.

Key Topics Discussed

  1. High-Efficiency Diffusion Models
  2. Challenges: Generative models, particularly diffusion models, are computationally expensive due to their iterative sampling processes.
  3. Innovations: Introduced SwiftBrush and SwiftEdit, enabling high-quality text-to-image generation and editing.
  1. Distillation Framework
  2. Methodology: A multi-step teacher model guides the training of a single-step student model to enhance efficiency.
  3. Architecture: Incorporates a secondary 'coach' network to align the student's denoising function with the teacher's, allowing the model to bypass multiple iterations.
  1. Practical Applications
  2. On-Device Agents: Discussed the potential for personalized agents that operate securely on local devices, leveraging private user data while maintaining privacy.
  3. Inference-Time Scaling: Explored the use of reasoning models and the balance between performance and computational constraints on mobile devices.

Technical Insights

  • Model Development: Transition from larger models to more efficient, smaller models without compromising performance.
  • Performance Metrics: The smaller sub-4 billion parameter model outperformed larger models due to optimizations in training.
  • Generative Art Quality: Quality of generated images with SwiftBrush and editing capabilities with SwiftEdit have shown to meet or exceed benchmarks.

Future Directions

  • Expansion of the AI Residency Program: Continued focus on attracting young talent through residency programs to foster research in the region.
  • Integration with Qualcomm: The acquisition has allowed for smoother integration and access to more resources and expertise in the field of AI research.

Conclusion Hung Bui emphasizes the importance of efficiency in AI models and the exciting potential for on-device personal assistants that leverage generative AI while prioritizing user privacy.

Key Takeaways

  • Efficiency in AI: Critical for deploying advanced models on consumer devices.
  • Research and Development: Continuous innovation is needed to balance the capabilities of large models with the constraints of smaller, mobile-friendly models.
  • Talent Development: Programs like the AI residency are vital for nurturing the next generation of AI researchers.

Further Reading

  • For detailed insights and research papers mentioned, visit [TWIML AI Podcast Episode 753 Show Notes](https://twimlai.com/go/753).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Thanks so much to our friends at Qualcomm for their continued support and sponsorship of today's episode. Qualcomm AI Research is dedicated to advancing AI to make its core capabilities, perception, reasoning, and action ubiquitous across devices. Their work makes it possible for billions of users around the world to have AI-enhanced experiences on devices powered by Qualcomm technologies. To learn more about what Qualcomm is up to on the research front, visit twimlai.com slash Qualcomm. So we're realizing this open-width model, 7 billion parameter model. We're still getting complaints from the community in Vietnam that, oh, this model is too big.

0:39We can't fit it on our GPU. And we said, okay, fine. Let us go one more step. Try to reduce the number of parameters. Try to half the size of the model to less than 4 billion parameters. And with a couple of improvements over the way we train it, we noticed that this model, less than 4 billion parameters, actually perform even better than 7 billion parameter model.

1:15All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington. Today, I'm joined by Hung Bui. Hung recently joined Qualcomm as VP of Technology through the recent acquisition of VinAI Research, which ranked in the world's top 25 industrial AI labs based on research output in top conferences like ICML and NeurIPS. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Hong, welcome to the podcast. Thank you so much, Sam, and it's my great pleasure to be here. So we've got a bunch of really interesting topics to dig into, including your research into topics like diffusion models, image generation and editing, and more, and of course, how to make all of that efficient on mobile devices.

2:04To get us started, I'd love to have you share a little bit about your background, which includes time at places like Google DeepMind and Adobe Research, among others. And tell us a little bit about how you got into AI. Actually, let's see, I did my PhD almost 30 years ago. And my PhD was actually on the topic of multi-agent system. And back then, it was also a very interesting AI topic. The reason I got into AI is just by curiosity. During my undergrad classes in Australia, I learned about things like Turing test. And I was very curious to myself that how we were able to program a machine to pass a Turing test.

2:48And back then, I have to be honest that I don't think I'm going to live to see the machine passing the Turing test, which is kind of like taking for granted now. Tell us a little bit more about some of the work that you've done in your career. So I think my first job, first real job after academia, was at a place called the AI Center at SRI International. SRI, Stanford Research Institute? That's right, yeah. It's formerly known as Stanford Research Institute. and back then, you know, we're talking about 20 years ago, I got a chance to work on a project called KLO and back then we tried to build, you know, guess what, a personal assistant.

3:37It was running on a desktop and we were trying a lot of things, you know, a lot of machine learning techniques, a lot of probabilistic reasoning techniques to understand the intent of the user and so on. And, you know, we even have a system that can actually record all the user actions on a desktop screen. And, you know, it was, I think, a very interesting project. I think back then it was sort of like the largest AI project of the day. And, yeah, I think one of the, you know, artifacts of that project was a little system known as Siri. Right. Right. And after that, I started to move to various industrial research labs in the Bay Area.

4:19I spent some time at Nuance Lab on Natural Language Understanding. Also some time in Adobe Research where we start to look into applications of machine learning in various areas of the business. And this is, I think, this is the time that Gerative AI starts to become popular with models like Variation Autoencoder and GAN and so on. And it's also correlated with my next move to DeepMai. And I was part of the DeepMai team, also still based in Mountain View, the Bay Area. And I think that was a really interesting time looking back. But yeah, I think, you know, in 2019, I actually got an opportunity to, you know, move back to Vietnam, which is, you know, the place I was born and had a chance to set up, you know, the first AI research lab in the country.

5:14Tell us more about that. Like you, you're coming from some of the top industrial research labs in the Bay Area, you know, the place where it's all happening. Like what prompts you at that point to leave all that and go back home? I mean, I could tell you that, you know, it was a very exciting opportunity. I, you know, I think a lot of us, right, when we're sort of like, you know, working for, you know, like big company in the U.S. or so, I think we all have a little dream, right, you know, to return to the home country and make an impact. And I think that's kind of like a big opportunity presenting itself for me.

5:54But I think, you know, in hindsight, I should admit that I also, you know, took a big risk as well because it wasn't known whether it's actually possible to actually run a, you know, like a proper one-class AI research lab in, you know, a place like Vietnam, right? Because it isn't known that, you know, AI research is going, it's like, it is something that requires huge investment, especially around talents. So, yeah, big risk, but, you know, I'm glad that I did it and, you know, it paid off. And how did you craft a research focus or direction for the lab? I must say that, you know, I must credit this to the really interesting research that I was doing back in the days where I was at DeepMind, right?

6:44So I think it is quite natural that we'll continue to follow the same direction in what is called back then deep generative models, which is kind of like the precursor of generative AI today. So from day one at VIN AI, we start off working on things like variation autoencoder, generative adversarial network, and of course, like autoregressive models, BERT model for Vietnamese, for example. or book model for tweets. And those are kind of like, you know, just the natural research topics that people will latch on back in those days. It just turned out to be that all of those topics, you know, are super important and it really prepares really well for the evolution of generative AI that follows until today.

7:35You mentioned that talent was a big challenge in starting a lab there in Vietnam. How did you address that? I remember myself that during one of the first few weeks when I was here in Vietnam, I stood in the Hanoi office. I looked out of the window and, you know, I kind of like, you know, just checking the view of the streets of Hanoi. And that's when it hit me that, OK, I'm no longer in the Bay Area.

8:08And then, right. OK, so how do you, you know, how you actually build a team here? And, but I think, you know, one thing that I know for sure is that, okay, I have to be able to convince people to move to Vietnam and work in Vietnam, right, for this effort to be successful. And then I have to be able to have a good balance between the experienced guys, right, who kind of like, you know, been there, done that, but are still willing to move back to Vietnam. With the, I would say, the young talents, the people in the country, right, the talents are still in the country. Whether, you know, just really, really smart young guys might not be that experienced in AI or, you know, things like differentiated models.

9:02But, you know, smart enough so that you can actually train them quickly. So, you know, it kind of like that was the strategy that, you know, we went about in building a team. So we will be able to hire a couple of, you know, I think really strong research scientists with many years of experience working in the U.S. and also other countries. They're willing to, you know, move back to Vietnam and work with me to build up the team. And then for attracting the young talents, then we actually started the first AI residency program, not only in Vietnam, in Southeast Asia, actually. And that has been a great way for us to attract and recruit the best young talents in this part of the world.

9:56And they come and work with us, become a full-time employee with the company for two full years. And yeah, we've had, I think, a really fantastic opportunity to work together with those young talents. They're just so smart and they're learning things so quickly. So the lab eventually became known for its work on efficiency. So mobile device, getting models to work on mobile devices. And that ultimately led to the acquisition by Qualcomm. I think that's a clear shared interest. From your initial research focus on deep generative models, how did this focus on efficiency come about? Going back to 2019, this is also the time when people start to look into how to scale up these models to work with larger and larger quantity of data and how to get model of larger and larger size, The number of parameters keep on increasing and demonstrate that the bigger the model, the more capabilities.

11:08And, you know, in Vietnam, I think we try to follow the trend. But very quickly, we are limited by our access to computational resources. We have, you know, the investment to have access to a small GPU cluster locally in Vietnam. right but you know very quickly you know we cannot compete with the giants in big tech and and and because of this right so knowing that technology is like deep generative model is important or generative AI is important but yet you know we kind of like you know we know that we are being constrained by the access to to computation resources to to to do to train these models and that That led us down to the only path that you could possibly go, which is how you'll be able to make these models more efficient, which means that we have to figure out a clever way to make the smaller models work, get more juicer out of the smaller models.

12:15And yeah, I think that's the reason why efficiency is a very natural focus. Talk a little bit about some of the research that that path led you to. Was it primarily focused on topics like quantization and related ideas? Or how did you approach it? I think the first thing that we're looking at is what we can actually get by smaller models. Right. And let me give you an example. ChatGPT was released. This is roughly November 2022. And this, you know, obviously caught all of us by surprise. But we were already working on a model like BERT for Vietnamese. Right. And then it's very natural for us to figuring out, OK, so can we actually pre-chain something like ChatGPT but using only Vietnamese data?

13:17And we know that this is at a time where CharchPT, I think CharchPT3 is close to 200 billion parameter. And we know that there's no way we will be able to rig that scale in terms of size with the training resources that we have. So we kind of run and perform an experiment to see what we can actually do with a very small model, something about only 7 billion parameter back then, right? And this is actually going against the trend when other companies were trying to get larger and larger model. We kind of like, okay, we ask ourselves, okay, what is it it can do with like a 7 billion parameter model pre-chain completely from scratch, right?

14:08Using Vietnamese-only data, right? So this is kind of like, you know, first it started as an experiment. Was this based on the GPT-2 architecture or what was the model architecture for it? Did you come up with that independently as well? The model itself is well understood. It's just like we're not entirely sure, you know, 7 billion parameter and the Vietnamese-only data, like whether we can actually produce anything interesting. And so obviously we grabbed all the Vietnamese data that we could from the internet crawl and then we fit into the modern architecture. And I think it took us a couple of months using our compute.

15:01And the end result was something that actually surprised us because we thought, oh, you know, we actually have a model that can speak Vietnamese really well. You know, you can answer questions in Vietnamese, you can, you know, write letters in Vietnamese, you can, you know, like write poems and all of that, right? So things that you see with the earlier version of ChatGPT, we actually saw it in Vietnamese with only a 7 billion parameter model. Then the next thing that we did is that we asked the question, what if we reduce the number of parameters for this model even more, right? So it's kind of like a counterintuitive back then, right?

15:44So we're asking ourselves questions like, can we get more for less? Here is just less number of parameters. Or can we get a simple performance with less number of parameters? And all of this is because it's the focus on efficiency, less number of parameters. I remember that back then, even with 7 billion parameter model, the other people in Vietnam, the other team in Vietnam, they're still complaining. So we're realizing this open-width model, 7 billion parameter model, we're still getting complaints from the community in Vietnam that, oh, this model is too big. We can't fit it on our GPU. And we say, okay, fine.

16:26Let us go one more step, try to reduce the number of parameters, try to halve the size of the model to less than 4 billion parameters. And with a couple of improvements over the way we train it, we noticed that this model, less than 4 billion parameters, actually perform even better than 70 billion parameter model. So we already know that there are a lot of room to optimize the model, the way we train it, and getting more performance out of a model that is even smaller. Can you talk a little bit about some of those optimizations and how you tweak that training recipe in order to get decent performance out of the smaller models?

17:11So one thing is, how do you get even more data, right? Now, we're not going to get more data, obviously, right? But we can iterate through the same data multiple times. And back then, we were already thinking, okay, can you actually use synthetic data to even drive more pre-training, right? You found everything you could find. Okay, so later on, of course, other people found the same thing for the internet. But back then, because we limited ourselves to Vietnamese, then we already hit that boundary. Okay, so we figured out, okay, if you iterate over the same data set more, perplexity keeps going down and the model keeps getting better, right?

17:50And also there are a couple of minor adjustments on the model optimization side that we did. But yeah, I think at the end, we have a sub-4 billion parameter model. And it starts to fit into just the basic GPUs that people in the labs, people in the universities, even now they could actually use. And so that was a very welcoming addition that we support the local community here by just giving these models, giving them out as open-wits. I'm just thinking about how so much of the innovation in models is driven by data sets. And here you are trying to create models that perform well on Vietnamese language.

18:39You collected kind of the, you know, just the raw corpus of internet information in the Vietnamese language. But, you know, when I think about, you know, traditional kind of NLP and, you know, model research, there are so many benchmarks that folks are trying to optimize models against that teach the models, you know, new and different things. and you didn't have all that for Vietnamese. Did you create any of that? Or did you find that just, you know, kind of the raw training on the data that you scraped was enough? We don't have data for Vietnamese for evaluation and so on, right? So we had to create some of that ourselves, of course, right?

19:31But I think lucky for us, a lot of the corpus for English can be auto-translated into Vietnamese, right? And I think looking back, machine translation between English and Vietnamese, and in particular, we actually also were also working on machine translation between English and Vietnamese, right? That was working well enough. So that lets you bootstrap, take an English QA data set and translate that into Vietnamese. Yeah, exactly. Exactly, exactly. And so how did you get into the image generation side of things? So text in languages isn't the only focus in the lab. Right at the beginning, we got people working on, for example, GAN model for image generation.

20:16Then, of course, the thing that comes around is diffusion-based approach. The quality of image generation that methods like diffusion model it can actually produce, you know, at the time it was getting more and more impressive. And so, of course, you know, because we were working on image generation using GAN, a lot of people in the lab starts to, you know, experimenting with the noise and diffusion. And that's, you know, it's a very natural thing for us. Regarding efficiency, it's a little bit different, right? I mean, the goal is the same, right? How you'll be able to, you know, get text generation and also here's how you'll be able to generate an image, but in a more sufficient way.

21:01So interesting enough, for image generation, the size model isn't that big, right? Compared to sort of like, you know, large language model. But for image generation, especially a denoising diffusion approach, the main thing was that you need to run these denoising steps for many, many times, right? So it's a sizable network, but you have to repeat it for many times, sometimes between 50 or 100 time steps. What that means is that if we have a big compute, you still have to wait. Back then, I think if you run this with ChatGPT, you still have to wait. It's not real time. This is noticeable length.

21:51And if you have a small compute cluster, you have to wait even more, right? So for us to be able to do experiments, then we have to look at ways to just be able to generate image in a much shorter time step and significantly reduced latency. And that drives us down the path of asking ourselves the question, can you actually get a model that can share an image with as few number of steps as possible, or even with just one step? So imagine that if you can just do this in just one step, your latency is all of a sudden boosted up by almost two other magnitudes. Is it still diffusion at that point, or are you needing to create an entirely new approach?

22:42We are working on the assumption that you already have access to the noising diffusion model, right? So you already have access to this model that can actually produce beautiful images, but it'll... Through the multi-step process. Yeah, exactly. So you assume that you have that. And you would treat that as a teacher, right? And you use kind of like a distillation approach, right? Maybe I should explain a little bit about intuition here. So distillation is usually thought of as, you know, you have a teacher that is a big model, and this big model encapsulates a lot of the knowledge about the task at hand, and you want to distill that knowledge into a smaller student.

23:30right but i think you you in here when you when here the problem is kind of like a step distillation because you don't have a big model it's just it's more of the same size you need to run it for multiple steps right and and here you kind of like have to you know uh distill that multi-step knowledge right into something that that has just one step meaning because your diffusion model is already relatively compact. It's not a huge model. You're not in this distillation process necessarily trying to shrink it. You're just trying to eliminate the steps. In fact, the architecture of the student and the teacher could be almost the same.

24:16It's just that the teacher, you have to run multi-step, but the student, just run step. And so another way you can kind of think about it is that, so first of all, the teacher is given. So this is already a really strong denoising diffusion model that has already been trained with lots of data, right? And it can actually generate beautiful images with, let's say, 100 steps. And it is given. But the way it is given to us is through a denoising function. because you're going to have to repeat this denoising function like 100 times. So that's fixed. That neural network is fixed. We're not going to change that.

24:59We would have to learn a student network. So that's distillation. That's a standard distillation. For any student network, you can also estimate a denoising network for that student. This is a function of the student network. And now you're forcing this denoising function for the student network to be the same as the denoising function for the teacher network. By forcing me the same, I mean, we're going to minimize the loss. So that's your constraint. You're going to minimize the loss to get them to be as close as possible. Exactly. So, yeah, the teacher network is fixed and constant, right? But this secondary teacher network or the coach network needs to be learned.

25:41What's the intuition for why you need the secondary coach network? Like if the teacher has all of the knowledge about how to generate the images, what is the secondary teacher actually doing? Great, great. Very good question. Very good question. So the answer could actually get to the heart of why you need that additional instrument, right? Right. So remember that in the first approach, right, we are asking that the denoising function for the teachers, right, has to work really well on distributions that are generated by the student. Which would be true if the student would generate the same distribution as the teacher would.

26:30So this is what we choose if we manage to successfully learn the distribution generated by the teacher. So I think there's like an agreement that, yeah, okay, this thing should be zero, you know, at convergence. Okay. But during the beginning of the process, right, the distribution generated by student is widely different from the distribution generated by the teacher. Okay. Yeah. Because the student's just not very good yet. Yeah, it's not very good at it, right? So basically the signal from the denoising function from the teacher might not be a very strong, even a good signal to follow, right?

27:15And that is the intuition, right? So the secondary network is kind of acting as a bridge between the students, the early student distributions and the teacher distribution? Yeah, yeah, yeah. Yeah, so that's exactly right. So early on in the process, right, we would want to explicitly estimate the denoising signal for the student. And then we minimize the difference between that denoising function versus the denoising function of the teacher. And that's kind of like providing the bridge of guiding the process during early stage of the optimization process. and so that initial work again this is swift brush and that is focused on um you know we've got image generation diffusion it works great takes a long time you know 100 steps so that's a lot of kind of inference so how do we make that more efficient well let's skip the 100 and do like one shot from noise to the desired output image.

28:26You know, even with all of the explanation of how that works and the intuition, like it's hard to believe that it actually works and works well. You know, talk a little bit about qualitatively what kind of results you saw. Again, you know, we're getting very good quality in terms of image quality, right? And also quantitatively as well, right? With all the benchmarks that we measure especially with some additional improvement that we later on did for the second version of Street Brush. We call it just simply a slightly improved version of Street Brush. The scores that we're getting is almost as good as a score.

29:04Sometimes it's even better than the score of the original teacher. And which benchmarks are we talking about? It's a bunch of benchmarks. It's a very standard benchmark on image quality and also diversity. And yeah, And this is kind of like a standard benchmark that you will see in any papers in this area. And you look at a score for this one-step model. We notice that we're getting the score sometimes as good as the teacher itself. The challenge is to get this thing to converge. The challenge is to be able to get this training pipeline to be in a stable condition and initialization and also the training pipeline.

29:54So it actually converged. But once it converged, we noticed that results are really strong. Yeah, so then, you know, like, okay, you can ask the question, okay, now that you can do text-to-image generation in just one step, it's really efficient and so on, right? can you also do text editing in a similar efficient manner? And that gets to this topic of image editing, which I think is a topic of a recent paper that we just published at CVPR this year. That was a Swift Edit paper. Yeah, so Swift Brass, which is just generating the image. But then Swift Edit, right, is to be able to also edit images quickly as well.

30:43in just like one step, right? So that's the key. It has the Swift Brass, one-step image generation, Swift Edit, one-step image editing. That is the goal. So, you know, do you want me to get into how it's working kind of like a high level? Yeah. Assuming that you already have a very efficient one-step text to image model, getting to text editing, you know, it's a lot more simple, right? If you don't have that, then forget it. But you already have a one-step text-to-image model, then text editing, you can start to conceptualize things in a much more simpler fashion. So the way we architect this model is as follows.

31:32So you have an original image. Of course, you have a text form. and you want this text prompt to kind of like, you know, operate on the image, right? And get into the final image here. Okay. But the way we want to do it is that first let's get this text prompt into noise, right? And then we already know how to get from noise to the image, right? Given the text, right? So then the challenge here is then how to get from image to noise, like condition on the text, okay? Yep. So to get from the image to noise, right? It kind of like, okay, you know, we could set up various laws to do this, right? If we have real images, right, as training data, right, you can set up what is called an inverted network or inverted model, right?

Read the full transcript

32:25And again, this is a one-step model. So this is like a neural network architecture that would take you from image, right? to of course the encoder in that image latent space right and then from the representation of that image to noise but you should do it in such a way that if you apply sweet brush again remember you have access to sweet brush right if you apply sweet brush again it would take you from noise to a latent representation of the same image right so you know like latent representation of that image Z, go to noise, apply ZBrush, and you get another latent representation. And you can simply say that, okay, this two latent representation has to be the same.

33:15And that gives you a very natural loss. You can also go to the image cell. You can go, okay, from this Z latent representation, you can use the decoder. Of course, all of this coming from ZBrush to generate the actual image. and you can say that the original image here and the image that you generate by applying SIPRATS to this intermediate noise has to be the same. And this is if you have access to only real training data. But you can actually do more because you have SIPRATS, so you can synthesize any images you want. So you can do this without access to any real data as well. So you can go from noise, applying sweep brush, going to a particular image or the latent representation of the image, and from that apply this inverted network going back to noise.

34:19And you can say that, okay, the noise you're starting from and the noise that you ended up with as a result of applying the inverted network has to be the same. So these two epsilon and epsilon hat here has to be the same. So that's just yet another loss function that you can have. There's another signal for training this inverted network. So having access to an efficient one-step model allows us to train this inverted network to invert from latent representing Z to noise and do that quite efficiently because you can differentiate through this network quickly. It's just one step, so you can differentiate through this network really quickly.

35:01And you can also combine signal from real data, also signal from completely synthetic data because the synthetic data can be generated really quickly with this one-step network. So all of a sudden, you have a combination of really just very like a loss function that's highly intuitive. and most importantly, you can implement it and train it very efficiently. All thanks to already have access to this one-step image generation model, which is Swift Brush, which is, I think, the call to making this Swift Edit possible. We'll link to these papers in the show notes and I encourage folks to pull them up.

35:41In particular, the Swift Edit paper, the images are super high quality and it's surprising how high quality they are given that process. Yeah, yeah, yeah. And it also works really fast, right? So we can actually run this on a standard one single GPU. And, you know, it, it takes us like, I think, a fraction of a second, I think a quarter of a second to do one, one image editing, which is real time, right? I think even before you finish typing, you already see the resulting image. So that work is all focused on efficiency. I know you're also working on kind of agents and making that type of model efficient enough to run on mobile devices.

36:29Can you talk a little bit about the way you're thinking about that space? The way I think about this is that I think we all want to have agents, assistants that are personalized, right? And what that means is that this agent or assistants, they need to have access to our private information. Otherwise, how can it actually be personalized, right? And where are the private information reside today, right? Well, you can argue that a lot of those private information actually resides on our personal devices. But that also means that privacy suddenly becomes really important. And so the way we think about this is that we want these agents or assistants to be able to do as much of the compute the workload on device as possible.

37:23Why? Because, you know, first of all, it's very close to the personal data that's sitting there. And second, and more importantly, right, if you can actually process all that information on device, right, and you don't have the risk of exposing this private information to third parties and so on. And of course, you know, if there are tasks that actually require information, you know, from the internet, right, and then these agents can also, you know, collaborate with, you know, a bigger agent, right, a more sophisticated agent with access to broader knowledge from the internet, right, that can actually reside on the cloud.

37:56And so how does that broad direction or vision translate into specific research projects? So first of all, I want to kind of say that this relates very closely to all the stuff that we talk about on efficiency. To be able to run this model to support this agentic behavior entirely on device, right? It means that, again, you have to look for ways to have not only large language models, but this is large multi-modal models that can take in both text and also multimedia as well. Those are the things that you actually have access on a device today. And you want to run all of that on a local device.

38:43So that means a lot of experimentation with smaller models. models that ranging from forbidden parameter or even less. And various ways to get these models to perform efficiently in terms of the rate, and which you can actually ingest tokens, pre-filled token rate versus encoding, but also decoding token rates. Another thing is that I think we start to kind of like look into what are the source of information right uh that are really important for this on-device agent right so i i mentioned access to private information on a device right and making sure that you process the information securely on the device itself but but then i think very soon you would need to look at the information that you're getting not only from your phone but from other wearable devices, for example, smart glasses.

39:52That's very rich in terms of multi-modality. That just opens up a really interesting space of research. How do you enable this model so you can actually understand video that's coming through your smart glasses. You can actually understand the content of the screens, right? That's actually the user is looking at in the phones and so on. And yeah, to be able to kind of like, you know, distill all that information to a compact representation, to be able to do that in a way that's efficient so that, you know, you kind of like, you know, don't consume a lot of battery, right? And then store that information somewhere so it can actually be retrieved.

40:36And so all of that is just a lot of work that needs to be worked out, both in terms of, again, the locking of information, compressing this information, bringing it to a form that can actually be ready to be retrieved later. And for retrieval, you can think about this as, okay, well, yeah, something like BRAC, but you do BRAC, but not only have access to information on the internet, but also information on a device, right? But if your device is being represented in a different way, the data is different, so you're going to have to make it work for this new distribution of data as well. When I think about agents and the kind of broader innovation that's been happening there, the rise of reasoning models has really changed that game and what's possible.

41:31That's very inference and compute heavy. I have to imagine that that poses a big challenge to running those kinds of models locally on the device. Are you thinking about this idea of inference time scaling and what that's going to mean for on-device models? Thanks for bringing that up. I mean, inference time scaling is a really important topic. You mentioned that, yeah, it's a challenge to run this kind of technique, inference time scaling on-device. But the way we should think about it is that it's both a challenge and an opportunity, right? Why I say it's an opportunity is because there has been multiple works in the literature showing that if you take a small model in terms of number of parameters and you apply test time scaling, right?

42:26And you measure on a particular kind of task, let's just say math, right? And with test time scaling, the small models actually perform a lot better than models that are significantly larger, right? And so test time scaling is a way that you can actually make the small model, right, to even beat a much larger model on a specific task. And I think this is really important because all of a sudden it enriches the capability of the small models. It's an interesting give and take. So like the test time scaling implies that you're exploding inference and that's a, you know, a constraint on a mobile device.

43:08but it also inherently allows smaller models to match the capabilities of much larger models, which is a tailwind for you trying to get this running on a mobile device. Yeah, exactly, exactly. And of course, regarding the challenge of how to make this work efficient on a mobile device, and this is kind of like more compute bound. It's no longer a memory bound, it's like compute bound. this is a topic that we are looking at very closely how to again make this test time scaling work more efficiently for example how do you do it assuming that you have an up about in terms of access to compute it's almost like a fixed compute budget so you have a fixed compute budget and what other strategies how do you allocate your resources across exactly it's interesting because I was going to ask the degree to which test time, you know, test time scaling and getting these reasoning models working on constrained devices is different than just the inference problem of, you know, having any LLM inference happen efficiently.

44:22And so it raises these kind of, you know, meta issues. And a good example of that is if you've got a fixed overall budget, whether that's compute or latency or whatever, and you're able to do some kind of planning that this is going to require some number of inference requests, like how do you optimize where you spend your compute is an interesting way to formulate a research question there. Yeah, I think you said it pretty well. And in a sense, this test time compute kind of like combine the probability that has been learned with an LNM, which is kind of almost like a forward predicting, with the kind of like an estimation of what's the future reward is going to be.

45:18So it's kind of like combining probability and utility in a sense. So it's kind of like already gone over what the original L &M was designed to do. So in that way, I think it allows us to move into a much, I think, richer problem, which is how we're able to find a particular answer path that maximizes expected utility. and that's a very rich framework, right? So, you know, things like compute resources constrained, you can see that it's possible that you can actually formulate this under the framework of optimizing for future expected reward. And so this idea of agents clearly kind of opens up many different research directions as you know and it kind of serves as maybe a grounding kind of application area are there others that are high priorities for you i would say that um yeah in in general um efficiency it's important for many broad topics and i i think it's this is something that people also have realized.

46:42You know, like I think for some time, you know, people probably don't focus too much on this issue of efficiency. But I think recently, I think, you know, there's a lot of focus on this particular topic. So I probably don't need to say anything more. So it's been, I think, six months since the acquisition. just what are your biggest kind of lessons learned through that process and what are you most looking forward to as you look ahead? I must say that we were pretty lucky because you know issues like efficiency right and on device models is something that is just something that Qualcomm and Qualcomm AR research, those folks, have been working on a similar problem for quite a while as well.

47:45And there's a great depth and breadth of expertise within Qualcomm AR research itself. So, because of that, I think integration has been a lot smoother. We don't have to change the objectives in terms of the way that we've been working. It's more or less the same focus. So for us, I think it's just more of a matter of learning about the capability within Qualcomm AI research and see how we can actually best help and enhance the capability of the group. So, yeah, I think, you know, we were pretty fortunate. And also, you know, like having access to the talents and the resources of Qualcomm ARI research is, you know, like making us even more excited.

48:47And talent was one of the big challenges that you mentioned when you got started. Are you continuing the residency program? Ah, yes, absolutely. I think this is, yeah, so this, the AI residency program today, we, it is something that will continue to reinforce. And, you know, like a little bit of a history of the residency program. We started this, having the first batch in 2019, that's six years ago. So, yeah, and at any point in time, we have about between 40 and 50 residents in the lab, right? 40 and 50? Yeah, between 40 and 50 residents. Wow, that's a lot bigger than I imagined. Right, yeah.

49:44And how many researchers total in the lab? The total number of people in the lab is about 90 people, right? So the number of residents is almost half or even a little bit more than half of all the research and engineers. Yeah, wow. Because they stayed with us for two years, right? They have enough time to contribute significantly to our research and also engineering process. And even now, right, the people who have gone through the residency program, you know, almost like close to 100 of them. And a lot of them are actually in top AI PhD programs in the US or Europe, Australia, and so on. And I think that has been a really nice tradition.

50:30And we want to continue to keep it that way. The program now would continue under the new branding of Qualcomm AI residency program. And I think we just hired the first batch of research residents. And we continue to look for ways to improve and expand the program as well. And we are about to recruit the first batch of engineering residents, so AI engineering residents. And I think this will also provide the opportunities for the young local talents to have this unique experience to be part of an AI research lab of a big tech company like Qualcomm. I don't know that we have a huge listener base in Vietnam, but in case there are folks listening that might be interested in the program, is there a page that they can go to to learn more?

51:25Yeah, we have a landing page for the Qualcomm AI residency program. You can find out about it from Qualcomm's website itself. And from that, we have a link to recruitment. Well, we'll find the landing page and stick it in the show notes. All right. Okay. Awesome. Awesome. Well, Hung, it's been great catching up with you and hearing a bit about your journey and the projects that you have worked on and are embarking on as part of Qualcomm. Thank you, Sam. And again, you know, thanks a lot for RTPT to, you know, share my thoughts and opinions here. Thank you. Yeah. Thanks so much.

52:07Thank you.

From the publisher

In this episode, Hung Bui, Technology Vice President at Qualcomm, joins us to explore the latest high-efficiency techniques for running generative AI, particularly diffusion models, on-device. We dive deep into the technical challenges of deploying these models, which are powerful but computationally expensive due to their iterative sampling process. Hung details his team's work on SwiftBrush and SwiftEdit, which enable high-quality text-to-image generation and editing in a single inference step. He explains their novel distillation framework, where a multi-step teacher model guides the training of an efficient, single-step student model. We explore the architecture and training, including the use of a secondary 'coach' network that aligns the student's denoising function with the teacher's, allowing the model to bypass the iterative process entirely. Finally, we discuss how these efficiency breakthroughs pave the way for personalized on-device agents and the challenges of running reasoning models with techniques like inference-time scaling under a fixed compute budget.

The complete show notes for this episode can be found at  https://twimlai.com/go/753.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
High-Efficiency Diffusion Models for On-Device Image Generation and Editing with Hung Bui - #753The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 52 min
Listen in VO