Personalization for Text-to-Image Generative AI with Nataniel Ruiz - #648

25 Sep 2023 · 44 min

Ask about this episode

Ask anything about it. ChatGPT or Claude reads this page and answers with the times it was said.

Connect VO and ask about every podcast you hear, including the moments you saved. Add to ChatGPT · Add to Claude

In short

Podcast Episode Summary: Personalization for Text-to-Image Generative AI with Nataniel Ruiz - #648

Podcast Details

  • Podcast Title: The TWIML AI Podcast
  • Host: Sam Charrington
  • Guest: Nataniel Ruiz, Research Scientist at Google
  • Episode Number: 648
  • Release Date: [Insert Date]

Episode Overview In this episode, Sam Charrington interviews Nataniel Ruiz about his work on personalization techniques for text-to-image generative AI models, particularly focusing on the DreamBooth algorithm. The discussion covers the effectiveness of DreamBooth in generating personalized models based on a limited number of user-provided images and explores various technical aspects, including fine-tuning techniques and evaluation challenges.

Key Topics Discussed

  1. Nataniel Ruiz's Background
  2. Originating from Bolivia, Ruiz pursued his education in France and the US, culminating in a PhD from Boston University.
  3. Formerly worked on adversarial attacks and deepfakes, pivoting to generative models and personalization technologies at Google.
  1. Introduction to DreamBooth
  2. An algorithm designed for personalizing generative models with a small set of images.
  3. Enables "subject-driven generation," allowing users to create unique images of a subject in various contexts using text prompts.
  4. Key features include:
  5. Fine-tuning on a few images.
  6. Avoidance of overfitting through techniques like early stopping.
  1. Technical Insights
  2. Fine-Tuning Process:
  3. Involves adjusting model weights based on a limited dataset of the specific subject.
  4. Uses personalized tokens that help differentiate between subjects.
  5. Challenges:
  6. Language drift may occur, causing the model to lose generalization capabilities. Strategies are in place to minimize this issue.
  7. Evaluation Metrics:
  8. Ruiz discusses the creation of a dataset for evaluation and the metrics used, including CLIP and DINO embeddings for assessing similarity and prompt adherence.
  1. Further Developments and Related Work
  2. Subsequent Papers:
  3. Ruiz briefly mentions other projects, including SuTI (Subject Driven Text Image Generation) and HyperDreamBooth.
  4. HyperDreamBooth:
  5. A new method that aims to make the DreamBooth process faster and more parameter-efficient.
  6. Introduces hypernetworks to generate model weights, allowing faster fine-tuning and improved preservation of model prior knowledge.
  1. Future of Generative Models
  2. Ruiz expresses excitement about the rapid advancements in generative AI and hints at further developments in personalization and model efficiency.

Key Takeaways

  • Personalization in AI: The work done in DreamBooth marks a significant advancement in creating AI models that can generate custom outputs based on limited user data.
  • Technical Innovations: Techniques such as prior preservation loss and hypernetworks show promising advancements in efficiency and quality for model personalization.
  • Research Landscape: The landscape of generative models is expanding, and the potential applications of personalization techniques are broadening, hinting at future innovations in the field.

Conclusion Nataniel Ruiz shares valuable insights into the evolution of personalization techniques in generative AI, the challenges faced, and the innovative approaches being developed to enhance these models. This episode offers a comprehensive look at current trends and the future direction of AI personalization.

For more detailed information and show notes, visit [twimlai.com/go/648](https://twimlai.com/go/648).

Written by AI. May contain mistakes. Listen to the episode to check what was said.

Hear the part that matters, and keep it.Open this episode in VO. Double tap your headphones to save a moment as you listen.
Get VO free

Transcript

Automatic transcript. May contain errors.

0:09All right, everyone, welcome to another episode of the TwiML AI podcast. I am your host, Sam Charrington, and today I'm joined by Nataniel Ruiz. Nataniel is a research scientist at Google. Before we get going, be sure to take a moment to hit that subscribe button wherever you're listening to today's show. Nataniel, welcome back to the podcast. Thank you so much. Yeah, it's been a while. It's been like three years, maybe. It has been about three years, maybe a little bit more, and I'm really looking forward to chatting with you about what you're working on nowadays, which is personalization for generative AI models.

0:41Before we jump into that, I'd love to have you share a little bit about your background with our audience. Absolutely. So yeah, a lot has happened since then. A lot of my research topics have changed, but I think for very interesting directions. My background is I was born in Bolivia. When I was 18, I left for France and I did my undergrad there in computer science in Paris at Equipity Technique. And then I came to the US, I did a master's at Georgia Tech, and then I did my PhD at Boston University, which I just finished in March this year. And then I joined Google right after. So I was doing a long internship at Google while in my last year where we did a lot of the work that kind of set me on this path right now.

1:17And then, yeah, I joined Google full time as a research scientist. So I'm working there right now. Thank you. Last time we spoke, you were we were talking about deep fakes and adversarial attacks. And that was part of your graduate work. Are you still working on those topics? Yeah. Funnily enough, it didn't even make it into my thesis. My thesis ended up being about simulating images and video in order to train and test machine learning models in efficient ways. That was kind of like the topic. So a lot of I tried to like make it very kind of related, all the papers that went into that. So the deepfake and adversarial attack against deepfakes work didn't really kind of relate as much into that topic.

1:56But I did some cool work there. I think we were a little bit early with that idea, which the kind of like core idea is to use an adversarial attack on images in order for people that are trying to like kind of generate deepfakes out of those images. to basically thwart them from generating deepfakes. And so it's basically an adversarial attack. Like the most basic idea is to do an adversarial attack against a conditional image translation network in order for the network to completely fail when used on this specific input. And in that way, kind of, we were a little bit early because the methods that were going around then and there didn't work very well with very few images.

2:32The outputs were in that realistic. And I think that's kind of had a bit of a revival now with all of the incredible generative models for images and video that have been popping up. So this is actually kind of taking a little bit of a renaissance, that idea. And now there's a bunch of work. For us, it was a bit hard to get this published or understood at the time. And now it's very, I think, easy for people to understand it when they see the quality of the generations now. So they understand the value of needing to protect identity and from being replicated, et cetera. So there's a lot of work from Alexander Madri's lab at MIT that has been kind of reviving this issue, which is really cool and really important, I think.

3:10And so your focus today is on, broadly speaking, as I mentioned, personalization for generative AI models. Talk a little bit about just the problems you're exploring and trying to solve there. Yeah, so kind of chronologically how this happened, how I got interested in, I guess, is I had work at Apple that's called MorphGan, which was basically, this is kind of what set me on the deepfake disruption route, with adversarial attacks, because we realized that if you train a model on, like this was again at the time, it was like 2019, I think, you train a model on a lot of images, like paired images of faces, then you were able to kind of like manipulate these faces and generate like one shot kind of deepfakes of people.

3:53So it was pretty cool in order to like, you know, have your avatar and be able to like talk, but obviously has some repercussions with respect to privacy that set me into that route. So that's where I had some experience in generative models, and then kind kind of got recruited here at Google for an internship on personalizing generative models in order to generate objects. For example, you want to generate this specific can in different scenarios, different contexts, and you only have a limited amount of images of that can or this dog or something. So that's when we set out to do that. But this was right at the moment where Dali2 had come out and Imagine had come out from Google.

4:26So we got early access to Imagine internally, and we were able to kind of set out that route of like personalization for generative models at the time, for diffusion models that were generating such good quality images. And surprisingly, you know, simple techniques work incredibly well with diffusion models. And that's kind of how we came up with Dreambooth, which is basically personalizing a diffusion model or a generative model for a specific subject in order to generalize it in different contexts and situations. And then we discovered they could do different styles, et cetera. And so that kind of set out my current direction, which is personalization of generative models.

5:00And is Dreambooth, it's the beginning of the title of the paper of that name, but is it also a product? Is it a model? Is it a system? What all is Dreambooth? Yeah, I would say Dreambooth is kind of like a method or an algorithm. That's what I would say. And it's like, the idea is personalization of a generative model with few input images, the personalization, yeah, of a generative model using few images of a subject for the subject driven generation. That's something I think the term we came up with, subject driven generation, which is basically you want to grab the model and you want to query it in different ways.

5:35And then you want to generate novel pictures of that subject. So it could be like a cat or, you know, like some type of - So I upload a few pictures of my cat or my dog and say, okay, now show me Sparky in Paris at night with a orange sky or something like that. And the idea is that this algorithm will generate images based on that prompt with that subject. Exactly. So that, I mean, I think you explained it very well. So that's like, it was kind of like a new problem because methods at the time weren't able to solve this problem before. And it was kind of a problem because the capabilities of current models weren't there yet.

6:13And then when the models had come out, like the diffusion models like DALI 2 and Imagine, And then it was kind of our contribution there is to create a method that is able to kind of tackle this problem. So like formalize this problem, tackle this problem. And also in the paper, we have a lot of like scientific contributions, such as like how do you measure success in this scenario? And we created the first data set that is large for this problem. So we kind of set out a new direction for research in this area. Are there things that, for folks that are tangentially familiar with diffusion models, but not familiar with their details, are there elements of the details, the workings of diffusion models that are important to fully understand what you've done and the way that you've used them?

7:04So diffusion models kind of like, I think I guess already existed like about seven years ago, I think, in 2016, if I'm correct. I'm 100 % sure. But then like in 2019, I think there were a bunch of papers that really made them work very well. And then 2019, 2020, and then, you know, more and more like training on larger data sets, getting some of the techniques right. And then they were kind of, now they've superseded like GANs for the best generative models. The core idea is you're basically learning to denoise images. So I guess a very natural application scenario was to do like denoising or super resolution as applications with this, which basically, you know, the core idea of the diffusion model is you grab an image that's clean, you noise it intermediately with a specific amount of like noise, and then you train a model that tries to denoise that image in one step.

7:58And then in inference, you go from, you know, if you want to generate new images, like people realize that you could generate new images after you train this model, if you just grab full a full noise, also full noisy image, and then you can go fully into the clean image regime, iteratively denoising that image. And that actually works really well, which is kind of like, it's an interesting kind of concept, because it's an asymmetric thing where your training is different from your inference, but it works extremely well. One thing that really makes it work is like the new techniques that have been developed like inside of Google and OpenAI and Berkeley in the last like three to four years.

8:32And there's a bunch of papers on that. And then one of the things that really makes them amazing is that they're much more stable than GANs. They're easier to train basically because, you know, GANs, you have a lot of stability issues. You can have mode collapse. You have to kind of be an expert in order to get your GAN to train well. And a lot of people are trying to do that. But I think with diffusion models, you have much more flexibility in what you're doing with the architecture, what you're doing with certain things, and it's kind of easier to get them to work. And that can change everything.

9:01I also think it's like objectively a really good idea. So that's why it works. Is it fair to think about the way you set up the problem as, you know, the traditional way of doing image generation is you are kind of conditioning the image generation on some texts? Do you think of the subject-driven generation as additional conditioning or a constraint or something else? Like, does that question make sense? Yeah, it makes a lot of sense. I think like you can think of conditioning the whole procedure of generating new images for this kind of cat or something as, yeah, condition on the data set. But in practice, what we do is we don't kind of create a new conditioning kind of like pipeline in the network.

9:45What we just do is we just do fine tuning on the network, on these images, basically. So that's how it like technically this is how it happens. So I wouldn't call it like, yeah, conditioning per se, but you know, more recent models are trying to do faster Dreambooth and they do take this conditioning route. So conditioning on a small data set of images. What we did, which is the contribution that we have for Dreambooth is like, okay, you have the set of images of the cat. Then what happens if you fine tune them? Like, okay, so we figured out a way to fine tune these diffusion models on the set of images of this cat by just adding like a personalized or a personal token that is rare that kind of identifies this cat.

10:22So it can be like a string of characters that is not, that doesn't have a strong prior in the model. And we condition like these. So we train the model using this kind of small data set of cat images with prompts A, V, where the V is like the rare identifier and then the name, or so the class name cat. So the model can localize, you know, very well actually and if you do early stopping with this very simple technique of fine-tuning the model on the small data set with these prompts that we came up with then it works really really well and you can start generating the cat it doesn't overfit because you do some early stopping and it was able to like memorize this subject so then it can like regenerate in different even sometimes different poses and you know with different lighting and different circumstances and also like different styles so a lot of very surprising things that we get out of these diffusion models.

11:15You make it sound very simple. So yeah, I guess, honestly, I love that about it. I think it's something that I think we were kind of first to see that this was possible with diffusion models. And because of several things, probably like we don't have like a clear answer on why this simple algorithm works. But, you know, possibly it's because these models are like huge, you know like much larger than previous models they're also trained on much more data it's trained on pairs of text and images so a lot of those reasons and also like some properties of diffusion models can like help with avoiding memorization because if you try to train again on very few samples then you get mode collapse very easily so you start generating the exact same samples very quickly with diffusion model this happens more slowly and it can have something to do with kind of maybe how the training is constructed, where you kind of train the fusion model to denoise partially noised images.

12:14And maybe that makes it, you know, overfit less fast, you know. So those are all hypotheses on why it works. But we kind of were the ones that figured out, you know, oh, this works. And this super simple idea works. We don't even have to fine tune a subset of the parameters, find which parameters to fine tune. We can just fine tune the whole model on this small data set with this very simple caption and it works. So I really love that the idea is simple. And it has like, I think that's where it really exploded because people could use it pretty easily or implement it pretty easily. And it just worked.

12:44You know, what I would have imagined to happen is, you know, less about overfitting and more about the generated images being kind of reminiscent of the subject, but not really fully capturing the subject. But the samples that you show, like the subject is there and it doesn't seem like it's distorted. Yeah. It does a really good job of recreating the subject in the scenario. and granted that you just said that you know we don't understand fully why this works all that kind of stuff any any thoughts on why that part works why the the subject is recreated so you know carefully yet without the overfitting yeah like the subject details are very well preserved i think for the one the model that we were using imagine at the time it was a great model trained on a very large data set so i think what it can do is like it already has like a concept of For example, for dogs, it was pretty easy to train it on dogs because it already has concept of these, you know, specific dogs.

13:45So it can mix its prior with the details that you're trying to teach it, which are specific kind of like details in the fur or colors or shapes, et cetera, of that dog. So it's kind of leverages is very strong prior in order to learn these things. And for like pixel level diffusion models, which is basically you don't have any kind of like down sampling into latent space at all. I think there is no real issue because you could overfit to specific images very easily. You just like train it all a long time on an image. So it will learn how to generate almost anything you give it, especially if it has this like already pre-trained, like it's a large pre-trained model.

14:22So it will be able to memorize things much faster than a random model, I think. And for the latent case, for the, for example, like stable diffusion, which is a latent diffusion model, then the only kind of obstacle that I see it from like actually reproducing the subject with good fidelity would be the encoder, the autoencoder, right? And the autoencoder is able to kind of reproduce most images with pretty high fidelity. It's not perfect, but it's able to do it. And then with SDXL, actually, it is like much better autoencoder. So it can reproduce many, you know, most images almost perfectly. So what's SDXL?

14:56Oh, Stable Diffusion XL, which I think is has been released like a couple months ago. So yeah, I think I would have expected it to maybe be able to very easily reconstruct the subject, but then it can't reconstruct it in other poses or in other contexts that easily, or it would have had trouble. I think the amazing thing about it is that it does learn the subject and it's able to preserve these capabilities. That was kind of the biggest surprise. And then also like do it in different styles and different accessories. Like we just kept discovering new and new things that the model could do when it was personalized.

15:31And can you talk about the training process for diffusion models generally and then how the training process for Dreambooth? is Dreambooth using an off the shelf diffusion model? You said Imogen, imagine. Yeah, that's the first model we used. And then we have also experiments on stable diffusion, I think 1.4 when that came out. And then since then, Dreambooth has been kind of re-implemented open source on all of the open source models, which are not many. So like stable diffusion, all the stable diffusion versions and stable diffusion XL. So, okay. So you're taking the off the shelf model and fine tuning it with the subject images.

16:13Yeah. And it's like, I took, so we took a bit of inspiration. We took inspiration from computer vision a little bit, but one thing where this was kind of already working for it was language models. So you had like a very large language model and then you fine tune it for a specific task. And then it did really well for that task. And one thing that I was noticing is that when you did this, it wasn't just that it was learning new things and it had to learn a ton of new things. It was that the prior of the large language model that had already been pre-trained was so strong that you were bringing out kind of knowledge onto the surface.

16:47And I think that's a pattern that you see nowadays is like these models have been trained on so much data and they're so large and they're basically so good almost everywhere, but they have to be good on average. So then once you fine tune in a specific scenario, a specific subject or a specific kind of for large language models for a topic or something, then you really get, you dig out that prior. You maybe teach it new things, right? But you dig out that prior and you kind of specialize it a little bit. It's basically the model that was already good. You give it a little bit more knowledge to memorize, and then you also bring out that prior, that part of the prior that you're really interested in for your specific task, and then it does really well.

17:23So that's one of the things that inspired me was the large language models. Yeah, and our team in general, I think. And so does the, you know, when you do this pre-training, you mentioned you've got the, you kind of come up with this unique identifier that you use as part of the prompting. To what degree is the model able to retain the distinction between my cat that we've fine-tuned on and cats? Like if you, you know, said a group of six cats, you know, with my cat in the center, like, is it just going to generate a bunch of copies of my cat? Or does it actually, you know, is it actually able to retain the, you know, the ability to generalize cats?

18:02Yeah, exactly. So this is one of the questions we grapple with at the beginning is that if you fine tune it even for a few iterations on a specific subject with this like prompt AV cat, then it'll kind of tie the word cat to your cat too. And then it won't be able to generate other cats. So it kind of becomes, so we call it that language drift, where it forgot the meaning of the word cat. So we have some techniques in the paper. So like one specific technique that's called the prior preservation loss that avoids it from losing its prior. And it was an interesting idea where, again, like Dreambooth is the core idea of like personalizing generative model for a subject, right?

18:37And the simple algorithm of fine tuning for that subject. And this is kind of an extension on that. It has some advantages, some disadvantages. but definitely one thing that it does do is it avoids language drift. So what we did is you just generate a bunch of images of other cats and you also mix those in with a different prompt, so a cat without the identifier, and you train the model on both kind of data sets of your cat and other generated cats. So this is one of the first situations where you could generate your own data with the model and feed it back into the model to achieve something.

19:09So I think that was kind of cool. We call that autogenous because it's like generating its own, like training samples. And I think that idea actually influenced some future ideas. For example, the Suti paper, this subject driven text image generation via apprenticeship learning is basically like, that's like a fast dream thing that I participated in that was at Google, like really good work by some colleagues. And the idea was to generate a ton of like dream data sets. So you have certain subjects of like one, you know, this cat and then this bag or And then you don't have enough images of them in different contexts.

19:43So like generate a bunch of data sets of them in different contexts. So dream booth data sets, and then retrain the whole model with a conditioning kind of conditioning pipeline. So like condition the model on small data sets and train this in a supervised way to generate new images. And basically you get fast dream booth for like a wide array of subjects and it could be specialized, but also, you know, in that paper, we show that it works for like a wide array of subjects. And the more data you can generate, the better it can become. And when you say a wide array of subjects, meaning the subjects are predetermined?

20:18No. As opposed to... I mean, like you can insert a new subject, like you can use it on a new subject and it has a very high chance of working, especially if the number of subjects and variety that you've trained it on are large, which they are in that paper. Okay. Yeah. So one thing that's interesting there is like exactly that thing. So we were talking about that prior preservation loss. It's just like, wow, we reuse the images of the model to train it. That's the marking of like, wow, those image generation models were already really good that we could use their own data to train them, train like for some things on.

20:47And then that has been reused in other work, I think, to some measure of success. That method sounds, you know, super hacky, but kind of amazing that it works. Yeah, yeah. I mean, you could use real images, but it just takes too long to look for real images. So that's why we use images of the same model so that it could just fit in a training loop basically. But there are some caveats, by the way. If you keep training a model on a lot of generations by the same model, you can get artifacts that kind of start reoccurring. So you can get a feedback loop. It's like basically if you train LLMs with, we've been hearing all about this, train LLMs with LLM outputs, then they become more and more limited and more and more biased and the outputs will start to become really bad at some point.

21:34But for one round of training, it works. Is the prior generation an automated pipeline that looks at the image that you upload and figures out what images to generate? Or was that all done manually? You could do it manually, but you could also do it in an automated way. I think the only thing you need is just the class of the subject. And I think more and more... That's what I was kind of speculating at. Yeah, and you can have that as user input. But I think more and more people also have realized that for some things, you don't need the class of the subject. So I don't know, some people also have stopped using that in some situations.

22:09We kind of came out with the first version of Dreambooth, I like to think of it. And then a lot of people online have improved it more and more. So for example, we train all of the parameters in the model. And then there's been like kind of Dreambooth versions that train like only the cross attention and self attention layers, which is kind of just realizing that you can do Dreambooth, but with less amounts of layers. You don't have to train the convolutional layers of the network. But it's still the same basic algorithm. And then there's also doing Dreambooth, but with lower rank adaptation. So instead of training the full matrix for a linear layer in the cross-attention self-attention layer, then you have a low-rank approximation of that, which is the LoRa paper.

22:51And then that has become the main way of doing Dreambooth because it's so parameter efficient, but it's still using the core idea of Dreambooth, which is basically the subject-driven generation and then prompting. And then also another extension of Dreambooth that was kind of like found out by artists very quickly after we released, which was super cool. Like one of the coolest things was following the art community, honestly, and their excitement about Dreambooth and other algorithms. So they just figured out that you could do it on style with enough images and with careful enough prompting, you could actually do Dreambooth on style.

23:23So you could learn a style And then you can have new images of other things in that style. But then there's also, we have like, you know, more recent work that does focus on style. Meaning subject-driven image, like how's it, how's what you're describing different from style transfer? Yeah, it's very different. So we called it subject-driven because the only thing we have in the Dream with paper are subjects that are kind of in different contexts with different accessories, in different styles, but it's still like you're learning the subject. But what people figured out very quickly is that instead of subject-driven, you can just change it to concept driven and your concept could be a style.

23:56So you learn a style, you personalize the model for style, and then you generate new things using that specific style. So let's say you have, you're an artist and you have a very distinctive style, like it's like watercolor and like certain colors that you use a lot. And then certain, I don't know, like lines that you, you know, you usually highlight certain things more than others. And you have like a portfolio. So you have a hundred images and people have realized that they can fine tune on those images using Dreambooth style. So like very similar to Dreambooth, but just for style. And then now they can generate new things, like let's say a dog in my style.

24:29And then they were able to do that. And this is kind of the genesis also of like future work that inside of Google that I also participated in that is able to like do personalization for style. So this is different than style transfer in the sense that you're not grabbing an image and saying, I'm going to use this style and transfer it to this image. So like make this image in that style. The content is not preserved, like the content is fully generated by the model, the diffusion model or the generative model. And the style is the thing that you had learned and that you're reapplying into this image, basically.

24:58Just as it knows, as the model knows how to generate something in the style of Van Gogh, then you just added a new one, you know, Nathaniel style, even though I'm not an artist and my style is probably not very good. And if you were comparing two pipelines, one that does generic generation and then apply some style transfer method to this dream booth style or i forget the what is that one called style booth is does it have a name a distinct name or is it just the yeah so so dream with people started using dream with their style very early on but then we have a paper at google called style drop which uses a different yeah so that uses a different architecture that's not a diffusion model so it's a transformer based image generation model Muse.

Read the full transcript

25:43And basically, but that one is like, that model is really, really good at learning style. And you only need like one image for it to learn style. And it's similar to DreamWorks, but we do use like adapter layers in order to fine tune the model. So it's like parameter efficient and we have some other improvements. Like, okay, so like read the paper, like there's a bunch of other cool stuff, but the core idea is basically that you can use one image and it already learns the style from that image. And then you can reapply it to different, to new generations. So yeah, there's definitely is like, yeah, style drop is really, really amazing in that respect, especially that model, you know, helped a lot.

26:18But yeah, like if you compare style transfer to like generating a new image and then doing style transfer on that image, I haven't seen that many results on that. I think that's still a good idea. But I think for style transfer, I don't know that much about recent work that does really good style transfer. And I think using the fusion models where style transfer is probably, there probably exists a lot of stuff that I'm, that I've missed or something, but I would say like, yeah, style transfer with diffusion models is probably the way to go right now if you were to do that. And I haven't seen any comparisons for now, but it's still a good idea, I think doing in a two-step rather than one step.

26:52But the one step is just conceptually very, very cool because it's like the model learned how to paint like this person, you know, like just realistic style or like, you know, like style can be expanded to so many things like certain types of lighting, you know, movie stills, et cetera. You mentioned earlier what you thought some of the core technical contributions of the Dream Booth work were. I think one of the things you said was evaluation and kind of the way you evaluated it. Can you elaborate on that a bit? Yeah. So we released an initial version of the paper before submitting to any conferences with kind of like the findings that you could do this and it worked really well.

27:33and the results were really, really pretty. But the work was to then make this conference paper. It ended up getting into CVPR and it got a student paper honorable mention, which was really cool. So we got to present it in front of thousands of people. That was a crazy experience. And one of our main contributions there was, okay, this is a new task. No one has really tackled this task successfully and we present this method to do so. So since no one has done it, so how do you evaluate it, right? It's such a subjective thing. You look at images and you say, these are pretty and this looks like my subject and it followed the prompt.

28:09You know, it says exactly what I want. But like, that's a human being kind of picking and choosing which ones are good. So we did a very, very large scale. So that took a long time. We basically, so I was able to like gather a data set of 30 different subjects. It doesn't sound that huge, but actually getting different images of a subject in different contexts. So like I went around taking a lot of pictures, but I also like we grabbed some pictures from other members of the team and other data sets, et cetera. So we made this like very specialized data set for different subjects that are very varied in different parts, like in different contexts and different poses with different lighting.

28:46So a lot of requirements to make it challenging. And then 30 is enough for now, I think. Like maybe bigger data sets will be good in the future, but 30 is enough for now because you have to train a model to get results for each of these subjects. new methods are faster, but at the time it was pretty slow. So it took like a lot of compute to like get all of this working and then generating a bunch of images. Then we kind of had a set of prompts that we put out with our dataset. All of this is public, by the way, I think it's like GitHub and then just dream booth in GitHub. We'll just type that and then you'll find the dataset.

29:18And yeah, so we came up with like a bunch of prompts that people could use. So recontextualizing it like in front of the Eiffel Tower and stuff like that. But then also like with accessories, like pets with accessories, like wearing a red hat or something. So a lot of different things like test how good your model was personalized for the subject with a lot of different prompts. So how flexible, this is called editability. So like how editable your subject is basically when you generate it. And then we kind of had to come up with like metrics. So like just or choose or come up with metrics that were good.

29:50So we came up with, so there were already some in use at the time. So clip, you can use it for image similarity. that can be for subject similarity. It's not the best metric. And then there was like, but it is one metric. And then there's like clip text, which similarity between the prompt and the final image. That was another metric we used for prompt following. And then we also use the dyno, kind of like dyno cosine similarity between two dyno embeddings. That was better in my opinion than clip image. So that's one thing we added there in that paper. But I think since then there's new metrics that have been popping up that are good for like higher level semantic similarity.

30:26Like, is this subject preserved? But then also that look at more middle level details, like are the details of the subject preserved? Not just, is this like a corgi? Is this my corgi, right? And I think there's still a long way to go for these metrics. Super interesting field of research. I'm not working on it currently, but I do hope like there's more work that comes out in that direction. I think there's like one work called DreamSim that is pretty cool. Actually uses simulated images to generate a metric. So again, this idea of like these models are so good that you can simulate images to make a metric that will then be used in the future for simulated images, I guess, which got to be careful a little bit with that.

31:02But actually, this metric is much better than other things that exist already. So, yeah. One of the extensions of Dreambooth is a project called Hyper Dreambooth. Can you talk a little bit about Hyper Dreambooth? Yeah, sure. So that has been my main work in the last. So I think we released it about two, three months, two months ago, I think. And that was my work before joining Google and like right after joining Google and my research work. And mostly like, so Dreambooth is really cool. You can personalize a model, you can get your subject in different, you know, circumstances, etc. But it has some weaknesses.

31:38One of the weaknesses was that it was not very parameter efficient. So the models you would have to save were pretty large, like a gigabyte or something for stable diffusion. It wasn't a big problem in many cases, but for many applications, you can imagine you would like smaller models. So that's one thing that LoRa applied to Dreambooth, basically just works. Because actually, with very low rank, you were able to get good approximation of your subject, and you didn't have to do this full parameter tuning for the full model. So parameter efficient tuning, definitely a thing in fine tuning of diffusion models and generative models in general.

32:11So that became kind of a mainstay. And we have like even smaller model in that work called Lightweight Dreambooth, which is 100 kilobytes instead of like several megabytes for LoRa, which still like it's not like that makes a huge difference in terms of like application. But it's still really cool scientifically and conceptually that you're able to reduce the number of parameters to very low counts, like even lower than we previously thought. And you're still able to personalize the model well. So I think we got up to like 30 ,000 parameters. So you only have to fine tune 30 ,000 parameters of a model that has millions of parameters in order to personalize it.

32:47So that was really cool. And the kind of technical aspects for this maybe are a little bit annoying to explain. Like without a figure, I guess, with a figure, it's pretty easy. But in essence, what we do is for each LoRa, so LoRa is basically you have like a full rank matrix, right, for the weights. And then you can decompose this matrix into two matrices that have lower rank. And then when you multiply them together, they're full rank, obviously. And what we did is we just decomposed these matrices into two other matrices, one that is fixed and another that is trainable. And the one that is fixed, we randomly initialize it with orthogonal vectors so that it kind of spans like a subspace in that full parameter space.

33:30It's basically like an incomplete basis. so it's like now you're playing inside of like a lower dimensional kind of subspace that is even lower dimensional than the lower than the lowest rank of laura which is one and you're still able to personalize the model well in that case so you're still able to find local minima that generate your subject with the details and also follow prompts so are editable so that was really cool so we found a way to do that and it's very kind of weird to me that it works but it was cool And why we did that is because, yeah, go ahead. I was just going to ask if there's like a geometric interpretation of or intuition around why this works.

34:12I mean, I can imagine it roughly just like you're going down like this very high dimensional space. I mean, you can't really imagine it, but you can conceptualize it. You go from this very high dimensional space. There's even lower dimensional kind of like manifold for the lower space with rank one. And then you can even go further down by exploring random directions there. It's like a smaller one, like an even lower dimensional one. And even then there you can find. So it's like basically you find a very lower dimensional manifold that contains these local minima, I guess. But why that works, I don't know.

34:49I think probably because there's a ton of local minima that actually satisfy what you want, which is basically having the subject be expressed in the model and then being able to like edit it. So that's kind of like what I think, but I don't think there's like a good geometric explanation on why that works. But it is surprising because very few parameters you're able to personalize a small for a subject. So then why we did that is not because we're greedy and we just want to make the ball smaller and smaller. It's because we were interested in something fast. So like the other bad part about Dreambooth, so like the other shortcoming of Dreambooth is that it's a slow method.

35:27So you have to train it for quite a thousand iterations, maybe or hundreds of iterations on your data set. Even if you've tuned the parameters really, really well and you've spent a ton of time tuning like the highest learning rates, learning rate schedules, et cetera, you're never going to get below a certain threshold. And what we were trying to do is get something to be much faster. So we were able to get it to be like an initial prediction of the subject and then like a little bit of fine tuning, which is like 10 to 20 seconds, basically fine tuning. And then you're able to get a personalized model instead of like five to 10 minutes with a dream booth, with traditional dream booth.

36:01So this is still not as fast as other work, which is like instant, like Suti is instant or like instant, just takes the time for like one generation. So it's like an inference step. But I think with both of these elements, you have like very strong properties for this method, which is basically, so I haven't explained what the method is, so maybe I should explain what the method is first, and then I'll explain why the properties are cool. So the method is basically, we use a hypernetwork in order to condition on an image of the subject, so we do it for faces, and then you have this hypernetwork generate weights that will personalize the model.

36:38So you basically generate a personalized model using this hypernetwork. And this hypernetwork is? HyperNetwork is a network that outputs model weights. So in this case, we would want the HyperNetwork for HyperDreambooth to output model weights that will make your model personalized for a specific subject. So then we'll be able to generate your face in different styles, for example. So I think one cool thing about HyperDreambooth is that we do have, so we have a prediction using the HyperNetwork for the model weights, and then we have a very short fine-tuning phase that's like 10 to 20 seconds long.

37:09And then you get the subject details quite right. And then one of the things that we did realize is that compared to competing methods, a lot of the methods retrain the full diffusion model or some parts of the diffusion model. So they lose the prior of the model, which is something we really didn't want to do. So we wanted all of the styles to be conserved. So that's one thing that we really made sure of. That's why we don't touch the model. We freeze it. And then we have these very kind of like low norm weights that we're trying to generate for the model that are then integrated into the model.

37:39So that's one cool property about HyperDreambooth. another one is like with this short fine-tuning phase that you get the subject details quite right basically. And that's another very important thing because people are very sensitive to you know if their face is not exactly the same in the stylized images that they're generating they won't be happy. So that's one thing that we were very careful of doing like with high consistency having like the same person be generated. So I think hyper dream booth I mean it's not the be-all end-all of like fast dream booth methods. I think there's so many such like good work out there, but it does have some properties that we were very interested in like conserving.

38:13And overall, I think, so it's one of the first works that does like hyper networks for this personalization of models, especially in the diffusion space. And I think this idea of hyper networks for diffusion models is really, really good because you can, you know, use it for many different things. It doesn't have to be limited for subject driven generation. It can be like for, you know, editing things or for inversion, you know, to get the same image again and stuff like that. So I think this hyper network idea is actually like a powerful one, because if you only limit yourself to like learning or generating a part of the text encoding vector, then that is not expressive enough.

38:46Sometimes you need to touch the model weights. Like that's one thing that we're learning is that the model weights are very important to like insert a new subject or a new concept into the model. So that's why hyper networks kind of are natural because the hyper network generates model weights and that will be able to like personalize your model. So I think this idea is like larger than just for Dreambooth, but we did apply it to Dreambooth and we thought it was like really cool that it worked and hopefully, you know, we'll see people using it soon. And just to recap all that, the initial Dreambooth work fine-tuned on the subject images, fine-tuning is essentially what you're trying to do is update the model weights.

39:25And so HyperNetwork, our hyper dream booth is, hey, let's try this other approach to updating the model weights, i.e. the hyper network training. And the benefits are twofold that it's a lot faster, but also because you're just kind of adjusting the weights via this delta as opposed to fine tuning, there's less disruption of the underlying model. It preserves the model's properties better. Yeah, yeah. Compared to other methods that are fast. And yeah, basically the core thing is it's faster, but then it also compared to other methods that try to be fast. I think one of the key things is we train on very little data and it works, like 15 ,000 images, so it can be easily replicated as opposed to other methods that use like millions of images.

40:08And then another thing that happens is that, yeah, it preserves the model prior very well, so we don't lose some of the styles that we really want to preserve. Got it. And so we will link to the papers that we discussed, Dreambooth, Hyper Dream Booth, Suti, Style Drop as well, all these in the show notes. I think you also wanted to shout out another work that you were involved in, Platypus. What was that one? Yeah, sure. That was very interesting because this is not really in my wheelhouse. I've been doing computer vision for all of my research career, but some students at Boston University, before I left Boston University, while I was still interning at Google, they were very interested in fine-tuning language models.

40:50So I can say that I have some work now in fine-tuning language models and we had the best open source large language model that was evaluated on Hugging Face, like open leaderboard for the whole world. And we had that for like two weeks and now everyone's using the data set that we created. And it's basically just kind of language, large language model, like llama type that is fine-tuned on a data set that is specialized for reasoning. And, you know, there's some data, some details on like how, how we got there and like how the students got there, which was very directed by them a lot. But it was amazing to work on something that was able to have such a big splash too.

41:25And was the reasoning data set that you created, was that the only data set or is that a relatively small kind of fine tuning-esque data set, but you trained on a much broader data set as well? Yeah. So the idea is we basically didn't have a lot of resources. So we wanted to basically create a fine tuning data set that was small and that was really powerful. So there's a lot of efforts in like sourcing this data set from other different open source data sets, filtering it, and then basically like deduplicating questions, a lot of work in there and checking for data set contamination, and then formatting the questions in certain ways.

42:03And then kind of like the biggest thing is like intuition from the students on which questions really will bring out that strong prior of the model for reasoning. And this selection, ended up in creating one of the best fine-tuning data sets that's very small, so it needs very little resources. You just fine-tune on it one epoch, and basically you get much better results already. So you're bringing out that prior of the model. And now everyone seems to be using this as part of other fine-tuning strategies. So the data set is going to live on for quite a bit, I think. Oh, very cool. It's related into this other prior, pulling the prior out of the model for images, but now this is for language.

42:42Awesome. Well, Nathaniel, it was great to catch up with you and chat about some of your more recent work. Thanks so much for hopping on. Yeah, thank you so much. Yeah, this has been crazy three years and I can't believe I'm here again. It's like looking back, wow, it's really amazing. So thank you so much. This is really a huge, feels like a great opportunity for me. Thank you. Awesome. Care to speculate what we're going to be talking about in three more years? I mean, I really did not expect the generative models to be this good. So I think it'll be insane. That's all I can say. I don't even know where I'm going to end up, what I'll be doing in terms of research, but hopefully it'll be cool.

43:21Awesome. Well, thanks so much. It was great to see you. Thank you, Sam. All right, everyone, that's our show for today. To learn more about today's guest or the topics mentioned in this interview, visit twimla.ai.com. Of course, if you like what you hear on the podcast, please subscribe, rate, and review the show on your favorite podcatcher. Thanks so much for listening, and catch you next time.

From the publisher

Today we’re joined by Nataniel Ruiz, a research scientist at Google. In our conversation with Nataniel, we discuss his recent work around personalization for text-to-image AI models. Specifically, we dig into DreamBooth, an algorithm that enables “subject-driven generation,” that is, the creation of personalized generative models using a small set of user-provided images about a subject. The personalized models can then be used to generate the subject in various contexts using a text prompt. Nataniel gives us a dive deep into the fine-tuning approach used in DreamBooth, the potential reasons behind the algorithm’s effectiveness, the challenges of fine-tuning diffusion models in this way, such as language drift, and how the prior preservation loss technique avoids this setback, as well as the evaluation challenges and metrics used in DreamBooth. We also touched base on his other recent papers including SuTI, StyleDrop, HyperDreamBooth, and lastly, Platypus.

The complete show notes for this episode can be found at twimlai.com/go/648.

More from The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

All 156 episodes
Personalization for Text-to-Image Generative AI with Nataniel Ruiz - #648The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) · 44 min
Listen in VO