Why Image Generation Needs More Than Bigger Models with Fatih Porikli - #773

12 Aug 2026 · 57 min · 17 chapters

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Why text-to-image (and related) generation still needs more than “bigger models”: better controllability, identity diversity, planning vs rendering, and efficient high-resolution generation on-device. It highlights Qualcomm AI Research work presented at CVPR.

Guest

Fatih Porikli, Vice President of Technology at Qualcomm. His team presented 20+ papers at CVPR.

Key claims

  1. Realism isn’t enough; models often fail at controllability (e.g., multiple people with distinct identities).
  2. Better training objectives (not just more data) can fix identity “cookie-cutter” duplication.
  3. Splitting tasks (planning/composition vs rendering) can outperform forcing one model to do everything.
  4. High-res generation is constrained by memory/latency; solutions should run on smartphones without changing base models.

Notable examples/papers

  • DISCO: reinforcement learning fine-tuning with explicit intra-image and inter-run diversity objectives; adds count accuracy; uses GRPO. Benchmark “Diverse Humans”; unique face accuracy ~98–99% vs 10–20% gaps for baselines.
  • R2CAN: “architect” predicts face locations/composition; “artist” renders while preserving provided identities; GRPO rewards include composition alignment, pose, count, identity retention, and human perception.
  • PixelRush: 16MP generation via cascade upsampling + latent-space patchification with semantic noise injection; ~10 minutes to ~20 seconds.
  • InverFill: image inpainting/editing by initializing diffusion with inverted-noise for the whole image plus masked noise, reducing boundary artifacts.
  • Video generation: “Pre-Mode VAN,” “Hybrid Attention,” and “Attention Surgery” for faster on-device video generation.

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Chapters

Tap a time to open that second in VO

The Future of Image Generation

0:00 to 0:15

Exploring the challenges in controllable and efficient image generation.

“Thanks so much to our friends at Qualcomm for their continued support and sponsorship of today's episode.”

The Future of Image Generation

0:31 to 2:15

Exploring the challenges in controllable and efficient image generation.

“resolution or local generation, and quality, speed, and memory quickly become constraints.”

Advancements in Text-to-Image Models

2:37 to 4:27

Discussion on progress and remaining challenges in text-to-image models.

“Yeah, that's a big difference that I see between, I think, even this conversation and the conversation we had last year, the performance and capability of text-to-image models has improved significantly.”

Challenges in Controllability and Quality

4:27 to 5:50

Identifying key areas for improvement in image generation models.

“And then you mentioned in there quality, so fewer artifacts, sharper images.”

DISCO Paper Overview

5:50 to 7:27

Insights into the DISCO paper addressing identity diversity in image generation.

“The name of the paper is DISCO, Resolving the Identity Crisis in Text-to-Image E2I Generation.”

Reinforcement Learning in Image Generation

7:27 to 9:33

Exploring the use of reinforcement learning to enhance model diversity.

“But the missing piece was that those models, space models, amazing models, had not really learned to create truly distinct identities.”

Curriculum Learning and Efficiency

9:33 to 11:34

The importance of curriculum learning in stabilizing model training.

“something called as group relative policy optimization, GRPO.”

Optimizing for Multiple Objectives

11:34 to 14:00

Debating the challenges of optimizing multiple attributes in AI models.

“And training, fine-tuning such models would be also kind of very affordable.”

Exploring Model Limitations in Image Generation

14:00 to 15:10

Discussing the challenges of asking a single model to handle multiple complex tasks in image generation.

“Maybe we are also asking a single model to solve too many difficult problems at once.”

The Concept of Model Routing for Specialization

15:10 to 17:40

Introducing the idea of model routers to improve image generation by directing prompts to specialized models.

“And that's what we explored in the other paper, R2CAN paper.”
Show all 17 chapters

Evaluating Diversity in Generated Images

17:40 to 21:20

Discussing how to evaluate image generation models based on diversity benchmarks and scoring systems.

“The other one is better text generation.”

The Artican Framework for Improved Generation

21:20 to 24:20

Exploring the Artican framework which separates planning and rendering in image generation processes.

“Yeah, we talk about that in the paper kind of even there is a nice flow diagram.”

How the Architect and Artist Models Collaborate

24:20 to 28:00

Explaining the collaboration between the architect and artist models in the image generation pipeline.

“Then kind of decides where, for instance, if there is a person is going to be in this image where it should appear.”

Episode Discussion

28:00 to 42:00
“and then the artist, the fine-tuned model, renders it.”

Exploring Latent Space in Image Generation

42:00 to 48:30

Learn about how image generation leverages latent space for faster processing and better quality.

“That's how diffusion, denoising diffusion, text-to-image models work.”

InverFill: Enhanced Image Inpainting Techniques

48:30 to 54:02

Discover how InverFill improves image inpainting while addressing common artifacts.

“And we can do everything into a bigger latent space, but we are petrifying and do it maybe much faster.”

Discussion on CVPR Papers and Video Generation

54:02 to 56:00

Gain insights into recent advancements in video generation and notable papers from CVPR.

“So as we suggested starting up, you know, Qualcomm always has a ton of papers at CVPR.”
Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:00Sam Charrington:Thanks so much to our friends at Qualcomm for their continued support and sponsorship of today's episode. Qualcomm AI Research is dedicated to advancing AI to make its core capabilities, perception, reasoning, and action ubiquitous across devices. Their work makes it possible for billions of users around the world to have AI-enhanced experiences on devices powered by Qualcomm technologies. To learn more about what Qualcomm is up to on the research front, visit twimlai.com slash Qualcomm.

0:54Sam Charrington:resolution or local generation, and quality, speed, and memory quickly become constraints. The next frontier in image generation is closing this gap between plausible images and precise results. It's making these systems controllable, efficient, and reliable enough to consistently produce a high-quality rendition of the image you actually ask for. One researcher at the forefront of this work is Fatih Parikhli, Vice President of Technology at Qualcomm, whose team presented more than 20 papers at this year's CVPR, the Computer Vision and Pattern Recognition Conference. Here's Fatih explaining why image generation still has plenty of hard problems to solve.

1:28Maybe we are also asking a single model to solve too many difficult problems at once. Think about what happens when you generate a scene with several people. Now, the model has to understand the prompt, but I'm asking the model, and then decide how many people should appear, determine where they should be placed, like the composition of the scene, reason about their interactions, because if there's a person, if there's another person, most likely there is some connection. Preserve the identity of the person. We can give, okay, this is my daughter, this is my son, and I want them to be in the picture, not like any random person.

2:07And finally render everything together in all a single process.

2:11Sam Charrington:So maybe this is too much. I'm Sam Charrington, and this is the TwiML AI Podcast. For over a decade, I've been exploring the ideas and innovations shaping the future of AI through conversations like this one that help you understand what's real, what's next, and what matters. Let's jump in.

2:37Sam Charrington:Yeah, that's a big difference that I see between, I think, even this conversation and the conversation we had last year, the performance and capability of text-to-image models has improved significantly. And not just performance and capability, but I think accessibility. Like now, ask whatever your favorite LLM is to generate an image, and it will do a really, really good job. And it does kind of beg this question of the computer vision community, like what's left to do? If this problem is solved to this degree, what's left for us to work on? That's a fair question. T2i models, these text-to-image generation models, or image-to-image generation models have become incredibly good at producing very realistic images.

3:30The lighting looks natural right, the details look right, and the overall quality can be amazing. As many things we do, there's initial excitement, There is great work coming up. But still, if you look into that one dive deeper, you realize there are many things to be accomplished. One was, you know, like the controllability. Generating multiple people in the same image, people look almost identical, faces kind of blend together. And also, we showed in the past, we can run such models on users' devices. You do not need to rely on a cloud service provider. I think most of the models still are limited to 1K, 1K resolution, but then you want to go beyond that.

4:17That is a challenge, which has not been actually addressed before.

4:21Sam Charrington:Okay, so what I'm hearing in there is that we've made a lot of progress, but there's still work to be done. And when you think about that work, some of the big buckets include controllability, the ability to really get the models to focus on the way you describe the task or focus on the output that you want. And then you mentioned in there quality, so fewer artifacts, sharper images. And then you mentioned efficiency. We need to keep up with our ability to run the latest and greatest models on the device. So those sound like three chunky buckets for researchers to continue to work in. Yeah, that's a very good depiction, Sam.

5:07Thank you. You know, kind of you said it very well. It's not only Qualcomm, but the community also trying to make sure the models, such models, genetic AI models, can also match the guidance or the quality or controllability expectations or efficiency expectation goes off kind of like a real, like a camera. But you are taking pictures using your birds, your prompts. So there are still many things to be done. And that's why, you know, we publish, present many papers addressing such challenges as CVPR this year.

5:49Sam Charrington:The first one we're going to talk to actually talks about this example you gave, the diversity of facial attributes. What's this paper called? The name of the paper is DISCO, Resolving the Identity Crisis in Text-to-Image E2I Generation. I'll mention to folks listening that a couple of things. One, we'll have links to all these papers in the show notes, but also we will, I'll try to get some of the illustrations from the papers into the videos. If you're not watching on YouTube or watching video, look for that because all of these papers have really good illustrative examples that will help with following the conversation.

6:34Sam Charrington:And so DISCO and Artican, when you think about them relative to these buckets that we've talked about, which of these buckets are they really going after? They are about controllability and how to train a T2i model better aligned with the guidance from the user. Let me go a little bit deeper about DISCO, if it's okay, San. When we look at, you know, kind of T2i, models and our own models also. What was interesting that this is not an image quality problem, the problem that I mentioned before. Like you are asking the model to generate faces and certain number of faces and it keeps generating same faces, almost identical faces over and over again.

7:26So quality-wise, image quality-wise, if you look at the pixels and noise and everything, it looks realistic. But the missing piece was that those models, space models, amazing models, had not really learned to create truly distinct identities. Because existing training objectives focus heavily on the realism and matching the user prompt, but they don't explicitly encourage diversity between people. And that is very important. So that observation led us to a simple question. What if the identity or facial appearance or any diversity, you know, kind of itself becomes an optimization objective? So that led into the idea behind Disco.

8:16instead of creating a completely new T2I model. We kept the underlying model and fine-tuned it with reinforcement learning. We designed rewards that encourage several things simultaneously. For instance, different people within an image should have distinct identities. You don't want to kind of duplicate, create duplicate faces. That's something we call in the paper as intra-image diversity. And across different runs of the same model with similar prompts, same prompt, we shouldn't generate, we shouldn't keep producing, generating the same faces. So this is inter-run, inter-image diversity. So these are explicit new objectives when we train, fine-tune the model.

9:06And also, you know, we want model to generate the correct number of people. If I'm asking generate two people, that should be two, not three. VRX also, you know, kind of incorporating that in the objective. And of course, we still have the previous, you know, image quality objective there. So putting all together, to put everything all together, we use reinforcement learning, something called as group relative policy optimization, GRPO.

9:39Sam Charrington:Taking a step back, it strikes me that, you know, yes, this as an objective diversity of faces is an important one if you're generating images with multiple faces. But it strikes me as, you know, one of many, you know, possible ways that you might want to or characteristics of an image that you might want to influence. And it seems like tuning the objective for all of the possible ways that you might want images to kind of generate correctly seems, you know, not just difficult, but like anti-bitter lesson. Like, is it, you know, why not just collect more data with lots of faces and train the model with better data?

10:34You know, that is what we are doing. But we are sure that you do not need a lot of data. And you are right. You know, kind of there might be many attributes, diversity, facial, you know, kind of diversity is one objective. A number of people is another one. But then let's say we want to generate a certain action of a person or pose of the face or location of the person in the image or their body pose and everything. So for those, but we are saying that, well, you can incorporate such objectives in addition to overall, let's say, perception or image quality objective when we are training such models.

11:18But you do not need a lot of data if the novelty of the paper in this work has to incorporate everything into a reinforcement learning framework that would make it possible. And training, fine-tuning such models would be also kind of very affordable. By the way, there is a difference between the quality of the input images going into such models. There is another thing called this curriculum learning, starting from clear scenes and gradually increasing complexity makes learning, this reinforcement learning, much more stable.

12:01Sam Charrington:I remember conversations I've had many years ago about curriculum learning, and it was always a theoretical improvement. It's exciting to see it being incorporated into practical algorithms, training algorithms nowadays. Yes, you are absolutely right. Curriculum now is making impact. not only, by the way, our papers, but there were several other papers I see talking about how wonderfully curriculum learning makes, let's say, multimodal models better. One broader lesson from the DISCO paper is sometimes the model simply needs the right objective in training, like the things that I mentioned. And also, it is not mainly an architecture limitation, but it is how you are providing these objective and training data to the algorithm.

12:59The way I would summarize your answer is that, yes, there are lots of different attributes that you might want to exert some control over.

13:11Sam Charrington:And, you know, there is a kind of a, you know, a mental or human cause to going after each of these attributes. But from a perspective of the training process itself, it's fairly efficient and nothing like when you think of like traditional fine tuning. It's, you know, very data efficient and you can apply techniques like curriculum to make it compute efficient as well. You know, Sam, something I like to also add because it is related to what we are talking about now. So optimization, of course, is right objective. Optimization at the right data is definitely very critical. That is what Disco paper is talking about.

13:56But we may also kind of think that maybe, I mean, these are models, genetic BI models. Maybe we are also asking a single model to solve too many difficult problems at once.

14:10Sam Charrington:That's kind of asking the question, could we possibly optimize for all the attributes that we care about? So in doing so, would we be asking the models to do too many things at once? That's a good point. Maybe we shouldn't. And also, I can give you an example. Think about what happens when you generate a scene with several people, going back to that core example. Now, the model has to understand the prompt, but I'm asking the model, and then decide how many people should appear, determine where they should be placed, like the composition of the scene, reason about their interactions, because if there's a person, if there's another person, most likely there is some connection, preserve.

14:54And we also, I didn't mention about this, the identity of the person. We can give, okay, this is my daughter, this is my son, and I want them to be in the picture, not like any random person. And finally render everything together in all a single process. So maybe this is too much, And that's what we explored in the other paper, R2CAN paper. Instead of asking one model to do everything, what if we separated, for instance, planning from rendering, similar to what human artists might do? And that kind of led us into this paper.

15:33Sam Charrington:Before we dive into Artican, thinking about that comment applied to Disco and again, this idea about different attributes, possibly a thing to keep in mind is how we've matured the way we think about like using model routers. so you know maybe you have you know a suite of models for the different attributes that you care about and then when your prompt comes in you've got some kind of router that says oh this you know the image that's required here you know will probably have a lot of people let me route it to this model that's been tuned for that as opposed to trying to again you know have a single model that's optimized for all the attributes you want.

16:25Sam Charrington:It's a different approach to what you've taken with Arcane, and we'll dig into that. But it is a way that, you know, that Disco could be put into production practically. Okay. You brought up again another, you know, amazing perspective, and we are working on it. This sounds like more like an agentic orchestrated image generation framework, part-time, right? Depending on the input prompt, maybe we want to generate, for instance, a realistic scene or like a cartoony scene or something, maybe there's a text or some human-generated graph or something like that in the image. So for all of it, different attributes like diversity, like facial identity, we may have specialized models, specialized processes within those models.

17:23So how to pull the right one is we are working on such agentic also pipelines. Depending on the input, it goes, it's, or determines the right tool. This is like, if we consider this call, one instance of, let's say, diversity is one instance. The other one is better text generation. So it can go find the right one depending on where we want to apply them and then orchestrate the final generation. I think this is going to be the ultimate solution. One, maybe size doesn't fit everyone. If you really, really want to generate something amazing, top of the line, so we need such specialization. Yeah, that's a very good perspective, Sam.

18:15Sam Charrington:So I will, again, the images with regard to Disco are really interesting. Like I never really thought of, I don't know that maybe I've just never asked for an image with multiple people, but there are some really good images where you have a prompt that, or several prompts where you're asking a variety of different models, GPT, Nano Banana, Flux, Hydream, a bunch of them to generate multiple people. And you're absolutely right. You look at these images and it's like the same person cookie cuttered across the entire group. It's kind of surprising that the models do that. When you talk about or when you think about evals beyond this idea of like running across the different models, did anything in particular jump out at you in terms of, you know, evaluation for this model?

19:12Oh, yeah. For instance, for this group, we look into existing benchmarks and then we pull together a new benchmark to specifically help people facilitate further research here and also, you know, kind of provide some standardization. We call it as diverse humans. It's also kind of available in the paper, the link for that data set, that benchmark. And, you know, we invite everyone to take a look at it. So we look at that, we created this benchmark and we evaluated this code on that benchmark and also on the other benchmarks data sets. We see that when you explicitly impose such objective, diversity objective, the score, for instance, unique face accuracy detection score, significantly improves.

20:10This score is around 98, 99. But, you know, kind of like the models that, these models where we started, that doesn't have such explicit diversity objective, they are very low. There is maybe more than 10, 20 percentage gap. So kind of this solution that we also talk about in the paper sets the new sort of for diversity in text image generation. In addition to like human preference score, which is very important, evaluating the quality of the output is also superior to the kind of like the base model.

20:53Sam Charrington:And it's interesting how you built the reward function. And I'm imagining that you would do a similar thing from an eval perspective. So one of the components of the reward is what you call intra-image diversity. And that is, you know, looking at one image, are the faces the same or are they different? and essentially you take your image, you apply a face detector that kind of identifies where the faces are with bounding boxes and then you kind of extract those faces and then you embed them and then you can just do similarity across pairwise similarity across the faces and kind of get the distance from face two and one face to the next to determine if they're, you know, all the same faces or essentially to get a measure of the diversity of the faces.

21:45Exactly. Yeah, we talk about that in the paper kind of even there is a nice flow diagram. I mean, I invite everyone to take a look at it. As you said, there is a face detector applied and this is off the shop, you know, kind of. And then we take these detected architecture faces, we embed them into some, you know, kind of space, canonical space to be able to compare them. Then we compute such pairwise similarities and then it gives us a score for that image. What would be the diversity score? Of course, that is one score. And we can do the same thing. Let's say I generated one image and I have another image with the same prompt.

22:34I have another image with the same prompt, but the initializations are different seed points. So kind of the prompt is same. We run it and then we end up with a different image.

22:45Sam Charrington:So you were starting to talk about the next paper, Artican, and essentially how this model kind of breaks down the problem of text-to-image generation. Talk us through that. So what motivated about that one, Disco is great. We, as a pivotal example, we, of course, talk about this identity, diversity. but then in the other paper we are making the point that maybe it's too much for a model to try to do everything going back to you know agentic flow again maybe it's easier to approach some of these genetic challenges like a human being we do not like try to solve everything ourselves. Well, not just ourselves, but even in this domain, art, like you, an artist approaching this problem wouldn't necessarily think about it pixel by pixel.

23:50Sam Charrington:They think about what's the subject, what's the background and... There's a planning, right? There's a planning aspect to it, yeah. In this paper, we build on it, build on that idea instead of asking model to do everything we separate planning from rendering similar to how a human artist would work. So we have two components. There are architect, which is, it doesn't generate pixels, but instead it creates the structure or composition for the scene. Then kind of decides where, for instance, if there is a person is going to be in this image where it should appear. And if there are multiple people, how they should be arranged.

24:33You know, it should look natural and, you know, realistic, then artist starts with that composition structure and generates the final photorealistic image while, of course, one objective also preserve identities if you provide identity. So, yeah, this is like planning an architect than an artist type of framework.

24:58Sam Charrington:The underlying technical approach is also using RL just like Disco. It's also a GRPO-based approach. You are right. We also have this reward function leveraging on different objectives that we want to optimize. In this school, we had intra-image and intra-group, like inter-image at human perception score and count accuracy. You see, here in R2CAN, we also have, well, we have the composition. Can we put the right face into right place in the composition type of objective, which is a part of the reward function for GRPO, and also pose of the face? Because now we are composing. I mean, we don't want people to look this way, the other people look the other way.

25:51If, let's say, we are taking a group picture, we could, you know, kind of, we could say that. So kind of those type of things are all combined into the reward function goes into reinforcement learning. And then we kind of like optimize that one. The model now learns how to pull all of it, optimize all of it at the same time.

26:15Sam Charrington:So to make sure I understand this, you have a prompt. the one in the example I'm looking at is three best friends riding unicorns on Mars. You pass that to what you call the architect. The architect says, okay, I want three faces, and they're going to be here, here, and here in this image, and kind of plans out this canvas. And then that essentially becomes a kind of grounding for your reinforcement learning loop. And so the images are all actually generated, but they're kind of optimized to ground to this canvas representation. Is that the right way to think about it? Right. So that model, that architect knows where to position such faces and artists rendering the base model where we started, now fine-tuned to actually follow the guidance based on the locations coming from the centroid of those face areas.

27:25But there are two things to keep in mind. There is a training phase offline. So this model, the architect and the artists are optimized to do better. But in inference time, you know, first we take the prompt and if you want, for instance, a spatial person identity to be in that image composition, we provide those images and architect generates the locations of the faces. And that is, there's no learning there. It just, you know, builds on what it learned before. and then the artist, the fine-tuned model, renders it. So this GRPO is on the fine-tuning offline ferry.

Read the full transcript

28:14Sam Charrington:And so that's an important distinction for sure. So the output of the architect is essentially three tuples, three XY points in the case of three best friends on Mars. Three people, yeah, exactly. Right. And then that becomes input to the artist, and then the artist just does a single-pass inference? Or, well, I guess it's a diffusion model, so it's iterative. It's going to iterate internally, but you are right. It's, you know, like single-pass in the sense that, yes, it's not going to go called something else, you know. It's going to do this. Yeah, it's interesting. Why is that interesting? I guess it's interesting that it works because the artist only gets, like, locations.

29:06It doesn't get, in the training phase, the artist gets actual faces, right?

29:13Sam Charrington:Not just centroids, or does it only get centroids in training also? If we want to say that there's a special person, there are special tokens representing a kind of like special person. If it is given, then we need to provide in training phase those target phases. But, you know, such phases can be generated. And that's what we did also by another model. You know, these are not necessarily real phases. I mean, to automate the training process, the training is very efficient. So in training time, those phases are given, the prompt is given, and then target is generated. And then we evaluate whether the identity of the generated faces matches with the initial identity in addition to other objectives, like count, accuracy, and pause, because pause could be part of the prompt, and also co-op.

30:12So in training, we are evaluating all of them, pulling them into this reward function, and then reward function through GRPO is optimized. And once this is optimized, now we, artists, know how to generate better, render better images given these input centroids. Architect already run and provided those locations.

30:41Sam Charrington:and for this one in evaluation you have a a bunch of prompts and then a bunch of faces that go along with those prompts and you're evaluating several things in the images one are the identities of the provided faces retained in the output image and also is whatever the descriptive scene or the prompt action reflected in the output image. Right. We have to do all of it, right? We want to make sure that we are still aligned with the input prompt. For instance, if you say three best friends riding unicorns on Mars, like the example in the paper, still we see three people and there are unicorns, you know, kind of.

31:31And it's like the scene is like Mars. So that is alignment to the input prompt. We still impose that. We also want to make sure that the generated image is high quality. This is human perception score. And there are three people, not like two friends. So configures there. And also, okay, I mentioned three friends, but specific friends, not like random people, not like two people riding unicorns, but three, you know, kind of Sam, me, you know, another person kind of like so we actually provide some examples of our faces so we also want our pictures to be also our faces to be in the generated picture so that's the identity face matching all the objective and also another one we want them to maybe look at the camera if that's a group picture i mean this is again one example of how to improve the image generation It doesn't mean that this is the only solution or, you know, kind of like welcome is showcasing what is possible.

32:38Sam Charrington:Is the identity retention like a big part of the kind of the value of this approach? Or do you see it as a valuable approach independent of identity retention? Oh, absolutely. Again, the composition itself would look much more natural because we ask architects to tell us about the structure of the image. Definitely. But, you know, kind of like I said before, the message is bigger. The message is, hey, community, everyone working on image generation, text image generation models. look what you can also accomplish if you really think smartly about the objective function, like the disco paper, and also simplify the task, like R2CAM paper, not trying to do everything at once, like I'm going to generate the composition and the render, but divide those processes into more manageable parts.

33:37Sam Charrington:So we've got a couple of papers as well focused on more of this kind of quality bucket. And Disco and R2Can also are both controllability and quality. But the next couple we've got here, PixelRush and Inverfill, that are really squarely focused on output quality. Talk a little bit about what these two are doing. Absolutely. The first two papers, like you said, were about making image generation more controllable and incorporating different objectives and making it easy for model to do the things that we expect it to do. The papers you mentioned, the pixel rush and invert fill, shift gears into a different objective now.

34:22We ask how to generate images much more efficiently than even what we did before. Because Qualcomm proposed and presented many papers at the previous CVPRs, how to run such models much more efficiently on mobile phones. But now we are saying that, can we even push it to the next level? because before we have been talking about, let's say, 1K resolution, and that's the current sort of, even the cloud models are kind of limited to that resolution. But the question is now, can we do like 4 megapixel image generation, 16 megapixel image generation? The challenge is not only, you know, kind of like how fast you can run the model.

35:09By the way, such models, existing solutions may take anywhere from seconds, like 50 seconds to minutes. They are not that fast. But there's a memory change also because when the image resolution gets larger, we need to retain this diffusion process, the latent features in the memory somewhere on the device. So that requires big footprints. So how can we run a model such that we generate an extremely large image, high-resolution image? And of course, the quality has to be also still real high-resolution, not just kind of like high-resolution up-sampled image, but real details, a lot of details are there.

36:03And it would run at a reasonable time. You don't need to wait like 10 minutes, which is the current, you know, kind of, I mean, except our solution, the time such model states requires. And it would run on, let's say, a memory available on a handheld device, like a smartphone. Yeah, that was our objective for the, let's say, PixelRush. And also we didn't want to change really the existing models significantly. So that is something to note. There are wonderful image generation models, and some of them kind of like by many big companies. We all know those models. We didn't want people to go and have to fine-tune those models.

36:57But we are saying that, hey, you can still use any of those models that you have and then follow our pipeline we discussed in the paper. So you can use that model, still generate, let's say, four times, 16 times more pixels.

37:14Sam Charrington:And so talk a little bit about the generation process. Like what's different about the way you've approached this? Yeah, absolutely. So it again starts with a prompt and there is this base generation like any model. It could be like a flux model. And then it generates, let's say, a base image. But base image, what I mean by that, let's say, 1K image. And then we have this cascade upsampling stage. That is the part that is known about this paper. That's why it's a CBR paper. this cascade up sample what it takes? It takes this image which is like RGB pixels, not latent space and then it creates for instance using any image super resolution solution, it could be cubic up sampling or it could be something smarter it generates let's say higher resolution image.

38:14So when we do that, okay we have now let's say 16 megapixel in the image, not one megapixel. We have like a lot of pixels. And then we take that image, Sam, and then we apply an encoder, a VAE. Then we go into a kind of latent space. In that latent space, of course, I mentioned that we want to, we are concerned about the memory. We now divide that latent space into manageable chunks. we patchify them so then we do improve those latent space features but then we are still in the latent space we go to through a VAE decoder in this case to the pixel space something we need to be very careful here there are solutions also using patchification like I'm going to take and create a kind of like patch, then another patch, another patch.

39:17And when you do that, you create artifacts actually, you know, like scenes, visible.

39:21Sam Charrington:Meaning when you do that in your origin space, you create patches. What's different here is that you're doing it in the latent space. Absolutely. There are many reasons. One is latent space is much smaller spatial dimensionality than the original pixel space. The other one is in latent space, we can induce nodes. And that is a smart way of leveraging noise. Yeah, that is the other reason, yeah. And so you have in the image describing the cascade, coarse latent refinement stage and then high quality latent. are these two separate latent spaces or is it one latent space? Like, what does this mean?

40:16Sam Charrington:Help me understand, like wrap my head around. Latent space is the latent space of the kind of original model. It is not a separate space. So what does the refinement stage do? Refinement space that adds some guided semantic noise and then it iterates a couple of times. That is the new part. That's the part that we provide. So it starts with these latent features. It gets some additional kind of flexibility through this injected semantic noise. We want that because we don't want suddenly, you know, kind of, okay, we are generating a big image, but then there isn't enough texture in the image or high resolution, semantically meaningful texture in that image.

41:11So that's why we want model to have flexibility to add things. We are in the latent space and then we induce such noise. I mean, we add literally some guided noise into this latent space because diffusion models, the ones that we are using, they do diffusion in the latent space and they start with noise, like random noise. And then it iterates, iterates, iterates like many, many times. I'm talking about like a high-level idea of the diffusion. And then every time it estimates some noise and removes it from the original previous noise. And then it clarifies step by step. And then it ends up with a final image, which is now noise-free.

42:02It looks like a real image. That's how diffusion, denoising diffusion, text-to-image models work. so kind of we are in the same space in the latent space is not the pixel RGB but these are spatially lower let's say input is 1k 1k this is maybe 128 128 and of course at every let's say point we have a vector representing features for that patch corresponding to a pixel so that is the latent space?

42:37Sam Charrington:Is the dimensionality of the latent space larger than what you might use for smaller images or is it the same? Yeah, latent space dimension is same. That's why this method is still very efficient. I mean, we could of course go to larger latent space. But that has a cost. But the memory challenge is there, right? And also it will take forever. The baseline actually does that, you know, kind of like they just make the latent space larger proportional to the target output resolution, which is not what we wanted to do. We wanted to run everything much faster. Yeah, that's why latent space is much smaller, like the original 1K inch latent space.

43:21Sam Charrington:So conceptually, the way I am thinking about what you're doing is you generate a low resolution image. You apply an off-the-shelf upsampler to get a larger image, but that image is going to be kind of blurrier and not all that great because that's the state of the art for upsampling. And then you essentially go into latent space and you end up in the blurry place, and then you apply diffusion to kind of generate out a higher definition. When we go to latent space, there is a nuance and it's important. So we upsampled the image input and this is BR RGB, like this is a color image, right? We have now a larger version of it.

44:14It is large. If I go to the latent space of the same proportional size to the depth size, it's going to be very slow. So what we do, we partition the input image into like original size patches. Let's say I started like 1K, 1K image. Now I have like 4K, 4K each side. So I divide it into 16 parts. So I have like 16 1K, 1K images.

44:40Sam Charrington:Do you project each patch into a latent space? Absolutely. Okay. Yeah. Got it. The nuance, yeah. Okay. And then? So then how do you avoid artifacts at the borders when you project out? So these images are not really super resolved in the way that we wanted because they are just up sampled, right? It didn't add any semantically meaningful details. We still have the original prompt, by the way. And that's why we add noise for each latent space. Yeah, that's why I think about it as like applying diffusion, you know, text-driven diffusion to this latent. Yes, yes. And so it allows us to create more texture.

45:31The algorithm is going to definitely generate much nicer looking output image than the input image, given input image before we went to the latent space. So in the latent space now, we've provided this freedom for algorithm to generate better semantically meaningful texture. But now I have 16 patches, right? And if I stick them together, arrange them in the original arrangement, that will be... Because they did it independent from each other, right? Because of the memory. And of course, we allow some overlap, small overlap, but it's not going to solve it. You know, when we pitch Y, by the way, another nuance is those patches are overlapping slightly, but it's not going to solve our problem.

46:16So what we also realized that when we are doing blending across such, you know, latent space data representation for these 16 patches, we like to, again, do another noise injection when we combine them together. But in this case, it is not like all over the patch, same in the latent space, kind of like sampled from the same distribution. We kind of allow algorithm to generate more noise towards the boundary, but in the center of the patches, you know, kind of like be maybe at a lesser degree. now we generate this high quality latent, which also at the same time now resolve added texture, but also resolve the boundary, this same issue.

47:13So you kind of like, I'm imagining that

47:18Sam Charrington:as opposed to like independently processing patch by patch, you're like pairwise or crosswise refining these patches so that on the other side of that refinement process, they're aware of each other in the sense that the edges are aligned and blended and all of that kind of stuff. Yes, but when we generate this high-quality latent, most of the problem is already solved. It will allow when we go to VAE decoder, generating this high-resolution image, which is now artifact-free. I mean, we solve the problem, handle the test at a space in the latent space where it is easier to do. Because if we go to the image space, if we start blending into image space, we will most likely blur around the seams, which is another artifact we don't want.

48:20Well, it's amazing that that works. Oh, yeah. I mean, again, like all our motivation in this case, let's achieve a quality level like we have a big memory and a lot of compute power. And we can do everything into a bigger latent space, but we are petrifying and do it maybe much faster. When I say much faster, it's not like two times faster. It's maybe 35 times faster from, let's say, 10 minutes to around 20 seconds type of acceleration.

49:01Sam Charrington:And it's interesting that we're talking about kind of boundaries and images and like reconstructing them correctly, because that's part of what the Inverfill is doing also, isn't it? In InforFill, we also have a semantically steered noise injection like the PixelRush. But fill is not T2I model. It is an image in painting model. So this is image to image generative AI. For instance, you can touch an object, create a mask, and then remove that object. or you can bring another object, new object into the scene and then you can create. This is very practical, Sam, because many times we take pictures, but in the background, maybe there are things that we want to remove.

49:53The challenge is, well, I can still see it, you know, kind of you create a texture. Sometimes, okay, it is meaningful, but then...

50:01Sam Charrington:There's those artifacts that you see in the background, like the sand on the beach where you remove the person who didn't need to be in the picture was kind of funky looking. And maybe the texture is not really compliant with the rest of the image or I see literally the artifacts around the boundary. So now, yeah, it is like a toy, you know. Yeah, you can do that, but I'm not going to really use it. But we are saying that, hey, you don't need to be there. You can do much better object removal or image in painting, image editing. And that's what this Inverfield paper is talking about. And so what's the core idea behind this paper?

50:43Sam Charrington:I'm imagining that you're using diffusion, you know, maybe multi-step or one-step diffusion somewhere. So in Inverfield, what we do, okay, we have this input image and we allow generating noise in the background and then we have also this noise in the mask, like the bird that we want to remove. But what I mean by like this noise in the background, the rest of the image. So in denoising, we started with, during training, with a noisy image and then we progressively removed noise and end up with a clean image. In image generation, we also do the same thing, right? We start with the noise and then end up with a clean image.

51:40But think about the other process, like the way that we actually train those models. Those models, when we train, we start with the clean image and then add noise, noise, noise, and at the end, it becomes like noisy image. So think about the reverse process where we actually train the model. so we can take this input image and map it into noise so we are progressively inverting a real, let's say clean image into noisy versions so this thing is very well studied, understood and it's very fast, like 60 milliseconds we can take a large image and then end up create going through the reverse denoising and a noisy version of it.

52:30So this noise is not random noise anymore. It is specific to the input image. So if I change the input image, the noise is going to be different. So this is about entire image and then we have this mask of the bird. Now I added noise to there for the bird because I want to allow algorithm to generate a new bird compliant with my field prompt, text prompt. So now I changed the way that we kind of create this image. But I don't need to change the original model. I can still, you know, I changed the way that I initialize the diffusion in this painting. And when we do that, when we start with, you know, like this inverted noise plus the mask and new noise within the mask.

53:27Now, first of all, we can retain background, but we also allow background to slightly impact the foreground, like the mask itself. So it allows seamless harmonization and it generates high quality images. But most importantly, there's no boundary artifacts anymore. because we have two noise, you know, like random noise and inverted noise within the same when we started the image. We are not starting from like just noise within the mask.

54:02Sam Charrington:Awesome, awesome. So as we suggested starting up, you know, Qualcomm always has a ton of papers at CVPR. We can't cover all of them. We've covered a handful of the most important image generation papers. but there were also papers on video generation. There were a ton of demos. Anything in particular you'd want to call out in terms of other things folks should look for? Absolutely. I'm very proud of the three video generation papers because it makes video generation accessible to everyone. You can run such models now on your PC, laptop, or your phone, you know, kind of those three papers are, I just mentioned their name.

54:52The first one is a pre-middle VAN. VAN is one of the best open source available, publicly available model. So we are showing in that paper, hey, now you can actually run it much faster and you can like five times faster and then put it into, squeeze it into your memory limited device. You don't need to run it on the cloud. The other paper is hybrid attention, rehired. So kind of recurrent hybrid attention paper. Again, there's this challenge of where to apply attention. So it kind of does a much better job and makes it faster. And also attention surgery, because attention takes a lot of compute.

55:38So these are for video generation amazing papers. please take a look at those project pages also. We have project pages for these papers.

55:48Sam Charrington:Well, Fatih, as always, it's been great catching up with you and hearing about what you all are doing with regards to computer vision and all the cool things you did at CBPR. Sam, thanks for having me. It's really a real and always a pleasure to be a part of your amazing podcast and I'm very excited to be a part of it, talk about artwork. I only mentioned a few of them and I look forward to kind of meet again, you know, join your amazing podcast. Yeah. Thank you so much for inviting me. All righty. Thank you.

56:41Thank you.

From the publisher

Text-to-image models have become remarkably good at producing realistic images. But realism isn’t the same as correctness. Ask for several distinct people, a specific composition, or a high-resolution image generated locally, and today’s models still struggle in surprising ways.

In this episode, Fatih Porikli, Vice President of Technology at Qualcomm, joins me to discuss what remains unsolved in image generation and several approaches his team presented at CVPR to address those challenges. We explore why better training objectives can improve controllability, how separating scene planning from rendering may lead to more reliable image generation, techniques for generating 16-megapixel images efficiently on edge devices, and new methods for eliminating the visible artifacts that often appear in AI-powered image editing.

Along the way, we discuss reinforcement learning for image generation, agentic image generation pipelines, on-device AI, and what the next phase of progress in generative vision systems is likely to look like.

🗒️  Full show notes: https://twimlai.com/go/773

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
Why Image Generation Needs More Than Bigger Models with Fatih Porikli - #773The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 57 min
Listen in VO